Any classified map is a hypothesis about what is on the ground. Without an accuracy assessment it is an untested one, and "we used a validated global dataset" is not an assessment of your map in your area. The rules that make validation meaningful are few, well established, and routinely broken in the same way.
The Rule That Gets Broken
Reference data must be independent of training data.
A classifier evaluated on the samples it learned from reports how well it memorised, not how well it generalises. The accuracy will be high and it will mean nothing.
This sounds obvious and it happens constantly, usually not through carelessness but through scarcity: collecting field samples in Indonesian conditions is expensive, there are never enough, and the temptation to use the same points for both purposes is strong.
The discipline is to split the sample before training, and never look at the validation set until the classification is final.
What an Assessment Produces
| Output | What it tells you |
|---|---|
| Overall accuracy | Share of reference points classified correctly |
| Producer's accuracy | Of the real class X, how much was mapped as X |
| User's accuracy | Of what was mapped as X, how much really is X |
| Confusion matrix | Which classes are being mistaken for which |
| Area estimates with confidence intervals | The quantity that actually matters |
The last row is the one that has changed practice. Reporting mapped area as a bare number ignores that the map has known error rates, and those rates can be used to produce a bias-adjusted area estimate with a confidence interval, which is a more honest and more useful product than the raw pixel count.
The confusion matrix is the most diagnostic output. An overall accuracy of 85% is uninformative, whereas knowing that the error is concentrated in one specific area of confusion, such as degraded forest compared to regrowth, tells you exactly where the map is weak and whether that weakness affects your conclusion.
Sampling the Reference Points
Points must be selected by a probability design, not by convenience. Points chosen because a road goes there sample roadsides, and roadsides are systematically different from the landscape.
Stratified random sampling by mapped class is the workhorse: it guarantees enough points in rare classes, which simple random sampling would miss entirely. The stratification then has to be accounted for in the accuracy calculations, which is a step that gets skipped surprisingly often.
What Counts as Ground Truth
Rarely an actual field visit for every point, that is unaffordable at national scale.
The realistic hierarchy: field visits for a subset, very high resolution imagery interpretation for more, and expert interpretation of the best available imagery for the remainder. Each is less reliable than the one before, and the assessment should state which was used for which points.
An assessment built entirely on interpreting the same imagery the classification used is weaker than it appears, because interpreter and classifier can share a systematic misreading.
Why This Is Not Optional Any More
Accuracy assessment has moved from good practice to requirement across both carbon standards and statistical reporting, for the same reason: an area figure without an error estimate cannot be compared across periods or aggregated with confidence.
For anyone producing maps that feed an account or a carbon claim, the assessment is not a quality-control afterthought. It is part of the product, and budgeting for it as an optional extra is how projects end up unable to state how good their central number is.
Move beyond estimates
Verifiers test sampling design, uncertainty and whether a number traces back to the field. TREEO dMRV captures that evidence in real time, in one auditable chain.


