AI and Earth Observation – The Annotated Data Bottleneck

A research paper published in 2025 states the field’s paradox in one sentence: remote sensing offers a wealth of Earth observation data for ecological studies, but the scarcity of labeled datasets remains a major challenge.

Abundance of Earth observation data on one side, scarcity of labels on the other. The opening article set out that finding; this one examines it, sets out what research has attempted in order to work around it and what remains unresolved. It extends the article on Copernicus.

What the research finds

A peer-reviewed review takes stock of foundation models in the field. It was published in 2025.

It notes that those Earth observation models face several key challenges, including a scarcity of high-quality multimodal datasets, limited capability for multimodal feature extraction, weak cross-task generalization, and the absence of unified evaluation.

That finding deserves attention for two distinct reasons. It comes from a scientific publication rather than a commercial actor, which removes any suspicion of interest. And it places data scarcity at the head of a list whose other three items belong to modelling.

Another publication frames the same tension differently, noting that the tension between large scale and small scale motivates foundation modelling precisely to enable more efficient transfer to tasks with limited labels.

The self-supervised response in Earth observation

This approach works around scarcity without removing it.

It deserves understanding in its results as much as in its limits. Its principle is to train an Earth observation model on unlabelled data, having it solve artificial tasks constructed from the data itself. The model thereby learns general representations, which later work specialises on a precise task with few labelled examples.

That approach has produced genuinely real results and it has considerably reduced the annotation need for initial training. It is today the field’s dominant paradigm.

Three limits nonetheless accompany it, and they are documented.

The first is that it shifts the need rather than removing it. Specialising on a task still requires labelled examples, in smaller quantity but of higher quality.

The second is that evaluation remains entirely dependent on a labelled reference. A model trained without labels must be measured with labels, failing which nothing allows its worth to be judged.

The third is geographic bias, which the next section examines.

Geographic bias in Earth observation corpora

This difficulty is documented precisely by research. It is the most important point of this article.

The work cited above notes that these models are often trained on datasets biased toward areas of high human activity, leaving entire ecological regions underrepresented.

The authors describe their own response by contrast: unlike comparable datasets, theirs uniformly covers the entire landmass, thus capturing all environment types without favoring urban and agricultural areas, or ignoring entire ecoregions.

That formulation by contrast is instructive. It indicates that existing datasets favour urban and agricultural areas, and that they ignore entire ecoregions.

Three consequences follow directly. A model trained on those datasets works better where the data was plentiful, meaning in already well documented regions. Its announced performance therefore reflects an average dominated by those regions. And its application to an underrepresented region is a transfer whose validity nothing guarantees, a limit the opening article flagged and which research confirms.

A benchmark analysis published in 2025 lists the gaps in existing evaluation sets: limited diversity, narrow image resolution ranges, restricted sensor modality coverage and geographic bias.

The inherited annotation trap

This mechanism is widely used and rarely discussed.

One of the field’s reference datasets is described thus in a publication: it consists of 590,326 image patches from a European optical sensor, annotated with multiple land-cover classes, the annotations having been provided by an existing land cover database.

The same source notes that training on that dataset improved accuracy compared with pretraining on a general image corpus, which demonstrates its usefulness.

A methodological observation nonetheless applies here. The annotations come from an existing database rather than from dedicated annotation. The opening article flagged the mechanism: using an existing map as a reference propagates its errors and its nomenclature, and caps achievable accuracy at that map’s own.

Three limits follow concretely. The nomenclature is that of the source database, which answers an administrative purpose rather than a project’s need. The granularity is that of the source database, generally coarser than the image resolution. And the database’s date does not necessarily coincide with the images’, which introduces unavoidable disagreements.

That mechanism in no way invalidates those datasets, whose usefulness is demonstrated. It indicates where their ceiling sits, and why dedicated annotation retains a value that volume does not replace.

The scarcity of Earth observation image-text pairs

This difficulty concerns the current generation of models.

A publication describing a recent dataset states the finding: vision-language models have shown strong performance in computer vision, yet their performance on remote sensing data remains limited due to the lack of large-scale, multi-sensor image-text datasets with diverse textual annotations.

The same source specifies that existing datasets predominantly include aerial imagery in three visible channels, with short or weakly grounded captions, and provide limited diversity in annotation types.

That finding matches exactly what the medical imaging cluster established. Image-text pairs are the current generation’s resource, and they are scarce in specialised domains.

Two precise terms of that citation deserve noting. Existing captions are described as short, which limits the information they carry. And they are described as weakly grounded, meaning poorly tied to a precise image region, which prevents a model from learning where what is described actually sits.

That second limit designates a very precise need: textual descriptions attached to identified regions, rather than global captions. That is annotation of a particular kind, more demanding and markedly less offered.

What these findings imply

Four practical implications emerge for an Earth observation project.

The first is that a foundation model does not remove the need to annotate. It reduces the volume needed for specialisation and it does not reduce the evaluation need, which remains entirely dependent on a reference.

The second is that a corpus’s geographic representativeness conditions what a model will be able to claim, research having documented the bias of existing datasets.

The third is that dedicated annotation retains higher value than inherited annotation, for the nomenclature, granularity and dating reasons set out above.

And the fourth is that grounded annotation, attaching a description to a region, corresponds to a growing and poorly served need.

Where the annotation need is strongest

Five situations concentrate it in Earth observation. They deserve distinguishing from one another.

Underrepresented regions, for which no existing dataset supplies examples and where a generic model performs poorly.

Specific nomenclatures, where a project needs classes existing databases do not distinguish, a frequent situation in agriculture and infrastructure monitoring.

Evaluation sets, whose quality requirement exceeds that of training, a principle the medical imaging cluster established and which holds identically here.

Poorly served modalities, radar and hyperspectral having markedly fewer corpora than optical, a gap the benchmark analysis cited above noted as restricted sensor modality coverage.

And temporal tasks, the constellations article having shown change annotation requires distinct conventions and is rarely encountered.

Those five situations share a decisive common feature. They are not resolved by waiting for a public dataset to appear, since public datasets are built precisely where data and references were already plentiful.

Strategies for reducing annotation cost

Four approaches combine in Earth observation. They allow the volume needed to be limited.

Active annotation consists of selecting the examples to annotate rather than drawing them at random, favouring those on which the model hesitates. It appreciably reduces volume at equal quality.

Automatic pre-annotation consists of having a model produce a first proposal, which the annotator corrects. It accelerates the work and it carries a risk documented in the segmentation cluster: annotators tend to validate what is proposed to them, which calls for a check on non-pre-annotated cases.

Weak annotation consists of using coarse or partial labels, a point rather than an outline for instance, which divides the unit cost at the price of less information.

And temporal propagation consists of carrying an annotation from one date to neighbouring dates, which presupposes verifying the absence of change and returns to the temporal question set out above.

One common caveat applies to all four of those approaches. They reduce the cost of training and not that of evaluation, whose quality requirement forbids shortcuts. A project therefore gains from using them for training and proscribing them for the evaluation reference.

What geospatial annotation demands

The work differs from ordinary image annotation on four points.

Those points were set out in the opening article and deserve developing. Ground knowledge intervenes constantly. Distinguishing two crops, two building types or two causes of disturbance requires a familiarity the image alone does not supply.

Conventions must be fixed finely. A parcel boundary, the inclusion of a hedgerow or the handling of a shadowed area are decisions to document, failing which annotators diverge.

Scale imposes an organisation. A scene covers hundreds of square kilometres, which makes exhaustive annotation impracticable and requires a sampling whose method conditions representativeness.

And seasonality imposes temporal discipline. One place changes appearance across the year, which makes an annotation valid at one date and wrong at another.

Those four requirements explain why this work resembles that described in the medical segmentation cluster more than ordinary image annotation, with the same consequence: the protocol and the case library determine quality more than volume does.

Synthetic Earth observation data and its limits

This alternative response to scarcity is frequently raised in Earth observation.

It consists of artificially generating images together with their annotation, by physical simulation or learned generation, which produces unlimited volume at low marginal cost.

Three uses are relevant there. Augmenting a real corpus for rare classes, a fire or a flood being poorly represented in an ordinary archive. Simulating varied acquisition conditions, viewing angles and illuminations, to test a model’s robustness. And testing a processing chain before real data is available.

Three limits frame it in return. The gap between synthesis and reality, generation reproducing what it has learned rather than what it has never seen, which makes it of little use precisely for the rare cases that motivated its use. The difficulty of reproducing acquisition artefacts, which the Copernicus article showed affect exploitation. And the impossibility of validating with it, an evaluation set having to be real on pain of measuring the fidelity of the simulation rather than the performance of the model.

That third limit is decisive and worth retaining. Synthesis can feed training; it cannot found an evaluation, which leaves intact the need for a real and documented reference.

What distinguishes a serious Earth observation corpus

Five signs allow it to be appraised, and they match those earlier clusters established.

The corpus composition is documented by region, by season and by sensor, information without which no claim of scope is verifiable.

The nomenclature is explicit and tied to a reference framework, with the handling of edge cases.

The method of obtaining ground truth is described, distinguishing field survey, existing database and photo-interpretation, those three sources not having the same reliability.

Inter-annotator disagreement is measured on a sample, which gives a bound on achievable performance.

And the evaluation set is built on areas and periods distinct from training, the only way to estimate a transferability research has shown to be problematic.

How a useful corpus is built

Five steps structure the construction of a corpus that will actually serve, and their order matters.

Define the nomenclature before annotating, tying it to an existing reference framework where possible and documenting the departures where not.

Build a case library of edge cases before production, having an expert settle ambiguous situations, a method the segmentation cluster established and which holds identically here.

Sample according to a documented plan rather than by availability, geographic and seasonal representativeness conditioning what the corpus will allow to be claimed.

Measure inter-annotator disagreement on a sample before launching full production, which gives a bound on achievable performance and reveals conventions that were insufficiently specified.

And build the evaluation set separately, on distinct areas and periods, with greater care than the training set.

Those five steps cost days on a project spanning months, and omitting them produces a corpus whose value can never be established.

What this means commercially

Three consequences emerge for a data provider.

The first is that the need is qualitative rather than volumetric. Large generalist corpora exist and are public; what is missing concerns specificity, geographic, thematic or temporal.

The second is that documentation is part of the value delivered rather than an accessory. A corpus whose composition is not documented supports no claim, which makes it unusable for a serious demonstration.

And the third is that a provider’s geographic position becomes an asset. Research having documented the underrepresentation of entire regions, an actor able to work on those regions holds precisely what is missing, an observation the climate article made regarding measurement networks.

An Earth observation asymmetry that does not correct itself

This observation explains the persistence of the problem.

Public datasets are built where conditions permit: well covered regions, available administrative references, established research teams. That is rational and it produces a cumulative effect.

Three mechanisms sustain it. An existing dataset attracts work, which produces comparable results, which reinforces its status as a reference and diverts effort from uncovered regions. Models trained on it perform better where they learned, which apparently confirms the relevance of those regions. And evaluation happens on those same datasets, which renders failure elsewhere invisible.

That third mechanism is the most problematic. A model can display excellent performance and fail across an entire region with nothing in the published measurements signalling it, a situation the medical imaging cluster documented as the gap between announced and real-world performance.

The consequence is that this asymmetry does not correct itself through better models. It corrects through the deliberate construction of corpora where they are missing, work nobody undertakes by accident and which is precisely the possible contribution of a provider established in an under-documented region.

What a client should ask a provider

Five questions separate an Earth observation corpus that will support a claim from one that will not, and they are worth asking before commissioning rather than after.

How was the ground truth obtained, and from which of the three sources. Field survey, existing database and photo-interpretation carry different reliability, and a provider unable to answer precisely has not thought about it.

What is the geographic and seasonal composition. A single figure for corpus size says nothing; the breakdown says everything about what the resulting model will and will not handle.

What was the inter-annotator agreement, and on what sample. That number bounds achievable performance, and its absence means nobody knows what the ceiling is.

How were ambiguous cases decided, and is that record available. A case library is the artefact that makes a corpus extensible; without it, a later batch will be annotated to different conventions.

And is the evaluation set separate in area and period. If it is not, the reported performance measures memorisation rather than generalisation, and the number cannot be relied on.

Those five questions take ten minutes. A provider who answers them fluently has built corpora before; one who treats them as unusual has not.

Common errors of reading

These misreadings recur often enough that naming them is usually enough to avoid them.

  • Believing a foundation model removes the need to annotate.
  • Neglecting the annotation need for evaluation, which remains entire.
  • Using a public dataset without examining its geographic coverage.
  • Taking an existing database as reference without knowing its accuracy ceiling.
  • Using an administrative nomenclature for a need requiring another.
  • Ignoring the date gap between a reference database and the annotated images.
  • Producing global captions where regional grounding is expected.
  • Annotating without fixing boundary and ambiguous-case conventions.
  • Building an evaluation set on the same areas as training.
  • Assuming an annotation remains valid in any season.
  • Founding an evaluation on synthetic data.
  • Using automatic pre-annotation without a check on non-pre-annotated cases.

Why this is a durable position rather than a transitional one

A closing observation addresses the objection this article most often meets.

The objection runs: models keep improving, so the annotation need will shrink and eventually vanish. It is worth taking seriously, and the evidence does not support it.

Three reasons stand out. Evaluation is irreducible. However a model is trained, establishing what it does requires a reference built independently of it, and no advance in training method changes that.

Specificity resists scale. Larger general corpora improve general performance; they do not supply the classes, the regions or the seasons a particular project needs, and the research cited above documents rather than contradicts this.

And each capability opens a new need. The constellations article showed new measurable variables create new insurable risks; each also creates a new annotation requirement, since a capability nobody has trained on is a capability nobody has validated.

Read together, those three explain why annotation demand has not fallen despite a decade of model progress. What has changed is its shape: less bulk labelling for training, more careful construction for evaluation and specialisation. That is a shift towards higher-skill work rather than away from work altogether, which is a better position for a provider than the alternative.

What to take away

Remote sensing offers a wealth of Earth observation data, and the scarcity of labelled datasets remains a major challenge, a finding a peer-reviewed review places at the head of the difficulties facing the field’s foundation models.

Three readings emerge. Self-supervised learning has reduced the annotation need for initial training without removing it, specialisation and above all evaluation remaining entirely dependent on a labelled reference. Existing datasets are documented as biased toward areas of high human activity, favouring urban and agricultural areas and ignoring entire ecoregions, which makes geographic representativeness a condition of what a model can claim. And annotation inherited from an existing database propagates its nomenclature, its granularity and its date, which caps achievable accuracy and leaves dedicated annotation a value volume does not replace.

For the strategic dimension of this question, the article on European sovereignty examines the continent’s position. For the data resource most of this work concerns, the article on Copernicus describes its access.

To explore delivery arrangements, supported formats and applicable control mechanisms, see our dedicated page on geospatial data processing. And if you need an annotated corpus on a region or a nomenclature public datasets do not cover, let us discuss your project.

Tags

Découvrez nos articles