Some people chase storms. I chase trees: through a million points of laser noise, into the Finnish Forest Centre's stand records, and out the other side with numbers that survived contact with the real thing. The forest above is synthetic, with all 450 trees placed by a script, so the method can be graded. Then the same method goes to work on a real 36 km² map sheet west of Helsinki and is checked, stand by stand, against the Finnish Forest Centre's forest inventory.
Finding 1The scoreboard lied
Thin the laser points out and you'd expect tree detection to get worse. Recall didn't: it held near 78% all the way down to one point per square metre, then rose to 82% at half a point, while the height error nearly quadrupled. At that density about one cell in eight of the canopy model gets no laser return at all (one in twenty at full density, mostly the lake), and treating those holes as zero breaks crowns into fake peaks, some of which happen to land near real trees.
Slide the density down and watch it happen: orange crosses are detections with no tree under them, orange rings are trees the method missed.
Height error is the honest number here. Recall only looked honest.
Finding 2Resolution beat point count
Building the canopy model on a 2 m grid instead of 1 m cost more trees than any amount of thinning did. At every density tested the coarser grid found fewer, and birch, with its wide flat crowns, ended up lowest every time. (Height error barely moved with the grid; it's the point count that wrecks height.)
Finding 3Perfect ground, biased canopy
The Forest Centre publishes canopy height models of the same sheet from several flights. Comparing 2008 with 2020, bare ground agreed to the centimetre: the median offset was +0.00 m for every pair of years. That told me nothing about the trees.
Real trees slow down as they get tall. So I grouped the change by each pixel's height in a third, independent year (2015) and looked at growth per band. From 3 m up it came out nearly flat, about a quarter to a third of a metre a year, even in stands already over 28 m, where the forest can plausibly add half that. A constant gain regardless of tree size looks like an offset, not biology, and the likeliest reading is that the 2008 flight measured the ground correctly and the canopy short. (The lowest band, mostly open ground in 2015, went slightly backwards, probably harvests between the flights, which this check doesn't remove. The dashed ceilings are my own rough limits, not a published growth table.)
Rankings survive a bias like this. Absolute growth rates don't.
Finding 4It nails height and misses most stems, and that's physics
On the real sheet, the method finds about 16% of the stems the inventory implies are there. That sounds like failure until you remember what a laser from an aircraft sees: the tops of the trees that reach the canopy. The suppressed trees underneath don't make peaks. Recovery climbs as stands mature and crowns separate.
Height is a different story. The mean height of the detected tree tops tracks the inventory's mean height closely, about a metre high, with a correlation of 0.96. Averaging every pixel instead sits below the inventory, because it averages in the gaps. Same raster, opposite sign.
One caution on that correlation. The inventory is the Forest Centre's, and most of it isn't a tape measure in the woods.
And one clean negativeThe plan had already decided
I built a ranking to pick which mature stands to harvest and scored it against the stands the forest plan proposes for cutting. It looked fine, until I counted the pool it was choosing from. Once stands are filtered to development class 04, regeneration-mature, and to the ones my ranking could pick (unrestricted, at least 0.3 ha, at least ten detected trees), nearly every one is already on the cutting list. Here is that pool, one square per stand:
With a base rate that high there's nothing left to beat: development class is the list, and the canopy model adds no lift on top of it. I'm reporting that because it's true, not because it flatters the method.
The Forest Centre's new data model explains why. Its cutting proposals are labelled by origin, and on this sheet they're simulated by the Forest Centre's planning calculation from the same inventory, not marked by a forester walking the stand. A model that proposes cutting for mature stands will propose it for nearly every mature stand.
Where the machines can go
Which stand to cut is decided. Terrain answers a question the plan doesn't: when, and how. A loaded forwarder on saturated ground leaves ruts, tears roots and sends silt into streams, so wet ground is cut frozen. Steep ground needs a winch. From Finland's open elevation model, worked on a 2 m grid, I measured both for every mature stand with a cutting proposal.
The two problems almost never land on the same stand, so they stay two flags with two fixes instead of one blended score. The first version had only the wetness-based season label, so stands that were 89% steep came out "summer-trafficable" with no warning at all. They still drive fine in summer as far as wetness goes; now they also carry an access flag. The thresholds are percentiles of this sheet, a ranking, not a rutting model.
Washington: does slope predict the regulator?
My Finnish thresholds were my own choice. Washington offers a real one: every Forest Practices Application carries the Department of Natural Resources' call on whether it involves potentially unstable slopes. In the Willapa Hills and Cowlitz valley about half are flagged, so a coin flip is the bar.
Move the cutoff and see how far the steepest slope in each application gets you.
The rules define unstable landforms partly by gradient, so this mostly reproduces a decision from one of its own inputs. The interesting part is the disagreements: inner gorges and bedrock hollows a 10 m elevation model can't resolve.
Things that silently corrupt results
- Joining the wrong inventory. Finland's stand records come in three flavours: observed, projected to 2026, projected to 2036. Join a projection and you're comparing a laser to a simulation.
- Stale ground truth. Observations run from 1999 to 2024. A 2020 raster against a 1999 measurement reads as detection error when it's two decades of growth. In the original run, keeping only inventories within six years of the flight moved the height correlation from 0.907 to 0.962. (On the current data most inventories are recent, and the filter drops only 28 stands.)
- Absence isn't permission. The stand file covers private forest only. Nuuksio National Park simply isn't in it, so a missing stand is treated as off-limits, never as open.
- One decision, counted many times. Washington's flag belongs to the application, not the harvest unit. Scoring units counted some decisions seventeen times.
- Nulls that look like findings. Stands too small to measure carried a null label, and "not summer-trafficable" is true for a null. Two phantom stands, one wet, one steep.
- A stand on the edge of the tile. One stand with 463 stems per hectare showed two detected trees, because only 3.7% of it had canopy data. Check coverage before trusting an outlier.
- A file format that changes under you. The Forest Centre moved its downloads and renewed its stand data model. The script failed loudly, which is the good outcome: it named the missing column instead of quietly comparing the wrong things.
How this page stays honest
Everything above is rebuilt from scratch on GitHub from the open data: the synthetic forest from the sample in the repository, then the canopy models, stand inventory and terrain downloaded fresh from the Finnish Forest Centre and National Land Survey. The numbers on this page come out of that run, not from a spreadsheet I typed.
The full write-up, with every wrong turn, is Finding the Trees on the blog. The code, the technical report and the rebuild are on GitHub.
Data: Suomen metsäkeskus / Finnish Forest Centre (canopy height models, Metsävarakuviot stand data), CC BY 4.0. Maanmittauslaitos / National Land Survey of Finland (1 m elevation model via CSC), CC BY 4.0. Washington State Department of Natural Resources (Forest Practices Applications) and USGS 3DEP. Imagery © Esri, Maxar, Earthstar Geographics. The synthetic Nuuksio forest is generated by generate_nuuksio.py.