Blog · Earth Observation · October 2, 2026

The Ground Truth Has a Date Too

A land-cover model and its reference map can disagree because they describe different moments, different categories, or different versions of the mapping process. Evaluation needs to establish which question both were supposed to answer.

A field after harvest

A hypothetical evaluator opens a satellite image of a recently harvested field. The model calls it bare ground. The reference dataset calls it cropland. A scoring script records an error. Before changing the model, the evaluator needs to ask what the label means: visible surface on the image date, cultivation during the growing season, or the field's continuing agricultural use?

The disagreement alone cannot settle that question. Under a definition that retains harvested fields within cropland, the model may indeed be wrong. Under a strictly instantaneous surface description, the reference may answer a different question. Neither conclusion follows merely from the colors on the maps.

This essay's argument is that temporal comparability belongs inside the definition of an evaluation task. It cannot be repaired afterward by attaching a year to the accuracy score. The issue differs from the site's satellite forecasting essay: here the difficulty is deciding what already happened, and what the reference evidence actually establishes.

Dates and definitions

FAO's classification definitions distinguish the physical cover observed on the surface from the activities and arrangements through which people use it. Cover and use are related, but they describe different properties. A label named “cropland” therefore needs its product-specific definition before it can become an evaluation target.

We propose recording the image acquisition date, the period the reference describes, and the date on which someone interpreted or revised that reference. These serve different purposes. An annotation made today from an old photograph supplies evidence about the photograph's date. It does not establish today's surface condition. An annual reference may deliberately summarize a period rather than describe any particular day.

For composites, the image side also needs a time window. Calling both files “2021” does not establish that their observations align. Nor does identical vocabulary establish identical meaning. A comparison needs explicit agreement about the spatial unit, the category definitions, and whether it concerns a moment, a season, or an annual summary.

The 2022 Dynamic World paper describes predictions generated from individual Sentinel-2 images, with class probabilities as well as labels. In its validation, the authors ran the model on the reference image dates. They also explain why spatially and temporally varying image quality makes it difficult to supply representative design-based accuracy measures for the entire continuously updating collection. That qualification limits how far a reported evaluation can travel.

Reference labels need maintenance

The WorldCover 2021 validation report provides a concrete example of maintaining reference evidence. Evaluators revisited a random subset of sites and sites flagged as likely to have changed, then updated labels where change was found. The report also addresses reference-image geolocation ambiguity using primary and alternative labels. Its headline accuracy is consequently the result of a specified reference and agreement procedure, rather than comparison with an infallible picture of the planet.

Our inference is that a reference dataset should be treated as a maintained scientific instrument. Revising it can improve an evaluation, but revision also changes what a score means. If a new model is assessed against refreshed labels and an older model against stale ones, the apparent improvement has more than one possible explanation.

A useful comparison would run both models against the same frozen reference release. A separate analysis could examine what changed when the reference was refreshed. Keeping these questions distinct lets researchers improve their evidence without quietly rewriting the history of model performance. This extends the site's discussion of constructed datasets into a practical evaluation choice.

A difference is a finding to investigate

ESA's WorldCover data page, consulted October 2, 2026, explicitly warns that its 2020 and 2021 maps used different algorithm versions, v100 and v200. Differences include both actual land-cover change and effects of the algorithms. Subtracting the maps does not isolate the physical process.

Consider another hypothetical case: a patch switches from grassland to shrubland between releases. Possible explanations include a real transition, a revised category boundary, different observations, or a corrected classification. Calling the patch “changed” is a reasonable description of the files. Calling it ecological change requires evidence about the patch.

The reverse problem matters too. Matching labels can hide a real transition if both outputs are coarse or both make the same mistake. Agreement is useful evidence only relative to the question being asked. A monitoring system built solely around disagreement would also need a way to check places it believes stayed stable.

Temporal mismatch must not become a general defense of poor models. Where definitions and observation periods align and the reference is credible, a wrong prediction remains a model error. An unexplained discrepancy should stay unexplained until further evidence resolves it; it should not be reassigned to “outdated truth” because that improves the score.

Designing a useful comparison

The remote-sensing guidance by Olofsson and colleagues recommends probability sampling, reference evidence with adequate spatial and temporal representation, analysis consistent with the sampling design, and uncertainty reporting. It also calls for assessing potential reference error. Those principles supply the foundation; the workflow below is our proposed application, not a tested new method.

Begin with a sentence specifying the target: for example, a hypothetical project might estimate the area that changed between annual cover categories under one fixed legend. That target determines what observations are needed. A convenient snapshot should not silently redefine the annual question merely because it is easier to label.

Keep an audit sample covering apparent stability as well as apparent change. Examine reference evidence for the relevant periods, with interpreters initially unaware of the model's answer where feasible. Record whether each comparison has adequate temporal coverage, compatible definitions, and sufficient positional confidence. These are reasons for assessing comparability, not additional land-cover classes.

Report performance for adequately supported comparisons alongside the extent and character of unresolved cases. Excluding ambiguous places may make evaluation feasible, but the result then describes that restricted population. A difficult region should not disappear from the report simply because it cannot be scored confidently.

When resources permit, obtain better evidence for unresolved cases using a documented sampling plan. Preserve the original prediction and record why a reference label changed. If a review is triggered by model disagreement, keep that selection visible: a pile of investigated disagreements cannot by itself estimate how common errors are across the whole map.

What the result can say

The practical output is a bounded claim: performance against a named reference release, for a stated period and legend, with unresolved coverage made visible. That is more useful than an accuracy number whose temporal meaning must be guessed.

It also allows different projects to want different things. A seasonal monitoring tool may need to detect short-lived surface states. An annual inventory may intentionally absorb those fluctuations into a stable category. Neither objective supplies the other's evaluation automatically. The reference earns its authority by matching the intended question and documenting the evidence available to answer it.

Sources

Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.


Return to Blog