Why Earth Observation Projects Fail

The fourteen preceding articles in this series on Earth observation described a sector whose resource is massive and free, whose tooling is mature and whose models are available. Despite that, a substantial share of projects never reaches production.

The causes are almost never technological. This article lists them by when they occur, the method the handling errors cluster used and which allows a difficulty to be detected where it arises rather than at delivery. It extends the article on the cost of Earth observation data.

Scoping failures

This first family occurs before any data is handled and it is the most expensive, since it invalidates everything downstream.

The badly posed question leads the Earth observation failures. A project launched on the idea of exploiting satellite imagery rather than on a decision to inform produces a capability nobody uses. The cost article posed the founding question: what decision depends on this data, and within what delay.

The mismatch between phenomenon and resolution is the second cause. Many operational questions concern objects or changes below the resolution freely available, which settles the cost question immediately rather than after six months.

The third is ignorance of actual availability. The opening article showed an announced revisit is an orbital capability rather than a guarantee of observation, and a frequently cloudy area may offer only a fraction of theoretical acquisitions.

The fourth is the absence of any assessment of available ground truth, which the AI article showed to be the field’s limiting factor.

Those four scoping causes are detected in a single meeting, using the four questions the opening article set out, and they eliminate infeasible projects before any commitment.

Data access failures

This second family delays rather than invalidates, without being any less real.

It occurs while obtaining the data. Underestimating administrative delay comes first where non-public data is needed, a difficulty the medical cluster documented and which recurs here for field data held by administrations.

Late discovery of a licence restriction is the second. The market article noted a licence limiting use to internal purposes forbids delivering a derived product, a constraint that invalidates a business model rather than merely raising its cost.

The third is exceeding a quota, the Copernicus article having shown free access comes with limits beyond which commercial conditions apply.

And the fourth is dependence on a single source in a sector where, as the market article noted, not all entrants will survive.

Reference data failures

This is the most frequent family in Earth observation.

It concentrates most of the field’s difficulties and deserves the longest treatment. The absence of ground truth leads the reference failures. A project discovering after processing the imagery that it holds no reference can neither train nor evaluate, and it generally ends with a visual demonstration and no measurement.

Using an existing database as reference is the second cause. The AI article showed this propagates the source database’s nomenclature, granularity and date, and caps achievable accuracy at that database’s own.

The third is an unsuitable nomenclature. A project using classes defined for an administrative purpose discovers they do not distinguish what matters to it, which requires starting again.

The fourth is geographic bias, documented by research and set out in the corresponding article: existing datasets favour areas of high human activity and ignore entire ecoregions.

The fifth is the absence of conventions for ambiguous cases, which produces variability between annotators that nothing can subsequently separate from real variability.

Method failures in Earth observation

This fourth family produces results that look good without being so.

It occurs during development. Evaluating on the same areas as training comes first. It produces a performance deployment does not reproduce, a mechanism the medical imaging cluster quantified in another domain.

Leakage through metadata is the second cause. A model learning to recognise a sensor, a season or a site correlated with the answer produces high performance without having learned what it was supposed to.

The third is estimating area by pixel counting, which the validation article showed good practice explicitly cautions against.

The fourth is using a non-probabilistic reference sample, which the same literature indicates forfeits the basis for any inference.

And the fifth is not measuring disagreement between annotators, which deprives the project of the upper bound on what it can measure.

Integration failures

This fifth family concerns technically successful projects.

It occurs at the point of putting a product into service. Absence of integration into an existing workflow leads. A result delivered in a format or a tool foreign to the user’s practice goes unused whatever its quality, a difficulty the Copernicus article raised regarding delivery into the client’s environment.

The second is the false alarm rate. The security article noted an arrangement flagging too many unfounded events is abandoned by its users, a mechanism the medical imaging cluster documented.

The third is a latency mismatch, a result arriving after the decision had to be taken, which the near real-time article covered.

And the fourth is absence of ownership, an arrangement designed without its end users meeting a resistance its performance does not overcome.

Durability failures

This sixth family is the least anticipated of all.

It occurs after entry into service. Performance drift comes first among durability failures. The cost article recalled that sensors change, the ground changes and processing chains change, which degrades a model with nothing signalling it.

The absence of a maintained reference corpus is the second cause. Without an up-to-date reference, drift is observed without being measurable or correctable.

The third is the disappearance of a data supplier, a real risk in a consolidating sector.

And the fourth is loss of competence, an arrangement built by a person who leaves the organisation becoming unusable if their method was not documented.

That fourth cause connects directly to a principle every cluster in this series established: documentation does not only demonstrate, it enables resumption.

Organisational failures

This family belongs to no technical moment and it is frequent in Earth observation.

Insufficient sponsorship comes first. A project without a sponsor holding a budget and an authority stops at the end of the exploratory phase, whatever the quality of its results.

The gap between the technical team and the operational team is the second cause. An image processing team optimising a metric the user does not recognise produces a result satisfying its authors and nobody else.

The third is disproportionate expectation, an executive having taken from a commercial presentation capabilities the available data does not support, which the opening article flagged regarding the gaps the sector passes over.

And the fourth is a change of priority, an Earth observation project spanning several months and crossing successive budget arbitrations.

Those four organisational causes share a remedy. Delivering a partial but usable result early sustains support, where a project promising a complete result in a year is exposed to every hazard of that period.

The cause running through every family

This observation explains the persistence of these failures in Earth observation.

Most of the causes listed belong to decisions taken before any development: the question posed, the reference available, the nomenclature adopted, the sampling design, the intended user.

Those prior decisions share three properties. They are taken quickly, in a few meetings. They cost little at that moment. And they become very expensive to correct afterwards, since they invalidate the work built on them.

A fourth property largely explains why they are neglected. They produce no visible deliverable, whereas image processing produces a map that can be shown.

That asymmetry between effort and visibility is probably the primary cause of failure in this field, and it is corrected by scoping discipline rather than by technical investment.

What distinguishes an Earth observation project that succeeds

Six characteristics recur in those reaching production.

A precise operational question, framed as a decision rather than as a capability.

Ground truth identified or obtainable, assessed before commitment rather than discovered after.

A pilot on a small sample, which reveals ambiguous cases and establishes a rate, which the cost article recommended.

An evaluation set built separately, on areas and periods distinct from training.

An end user involved from scoping, which pre-empts the integration difficulties set out above.

And documentation produced during, which allows resumption, demonstration and maintenance.

What Earth observation failures actually cost

This economic reading clarifies priorities.

A scoping failure costs the entire project, since nothing produced serves.

A reference failure costs the annotation work redone, the line the cost article showed to be dominant.

A method failure costs the evaluation redone and, more seriously, the credibility of a result already announced.

An integration failure costs the whole development, the product existing without being used.

And a durability failure costs reconstruction, generally more than the initial build since the context has been lost.

That cost gradation leads to a simple conclusion. Prevention effort is most efficiently placed at scoping, where it costs days and avoids months.

The difficulty nobody flags

One cause of failure deserves naming separately because it is structural and rarely raised.

A successful project produces a measurement, and that measurement may contradict a decision already taken or a position already defended.

Three situations show it. A map contradicting a declarative inventory places its sponsor in difficulty, a mechanism the climate article raised regarding gaps between declarations and observations. A rigorous assessment revealing lower performance than announced forces a communication to be corrected. And objective monitoring of a situation may establish a fact some actors preferred left indeterminate.

Those three situations involve no technical error and they cause technically impeccable projects to fail.

One practical consequence follows for a provider. Establishing at scoping who will receive the result and what happens if it is unfavourable avoids a late abandonment, and that question, awkward though it is, is asked more easily at the start than at delivery.

Resuming a troubled Earth observation project

A provider is regularly approached after a first failure rather than at the outset.

Four questions determine what can be salvaged.

Which family does the difficulty belong to. A scoping failure requires starting again; a method failure generally leaves the data and the annotations intact.

Is the existing corpus documented. A corpus whose composition, nomenclature and conventions are recorded can be reused; an undocumented one is re-annotated, which the cost article showed to be the dominant line.

Are the initial decisions revisable. A nomenclature imposed by a client or by a regulatory reference cannot be changed, which bounds the options.

And is the existing evaluation usable. An evaluation set built on the same areas as training informs nothing and must be redone, an operation that conditions any resumption.

Those four questions take days and they produce a diagnosis rather than an impression. A resumption launched without that diagnosis generally reproduces the original difficulty, having failed to identify its nature.

Earth observation failures that are not failures

Not every project that fails to reach production is a failure.

Three situations are results rather than failures.

An exploratory project concluding infeasibility has produced valuable information, provided it produced it quickly and cheaply. That is precisely the function of the scoping described above.

A project whose achieved performance is insufficient for the intended use, but whose measurement is reliable, has established a fact. The difficulty is not the performance but the claim that would have been made without measurement.

And a project interrupted by an external change of priority leaves a documented corpus that will serve later, provided its documentation exists.

Those three share a feature. They differ from a failure through the presence of a measurement and of documentation, which is precisely what the articles in this series recommended producing during rather than after.

One practical consequence follows for an Earth observation project. A rigorous arrangement does not guarantee success; it guarantees that a failure will produce usable information rather than a dead loss, which is a sufficient justification for the effort.

What the prevention discipline costs

This calculation settles the commonest objection to that discipline.

The six preventive measures described here represent a measurable effort.

The four-question scoping takes a meeting. Verifying availability over the target area takes a day. The annotation pilot takes a few days and produces a quantified rate. Building the evaluation set separately adds a share to the annotation budget. Involving an end user takes a few hours spread out. And documentation produced during adds a fraction to production time.

The whole typically represents a few per cent of a project, and it addresses the causes of most of the failures listed.

One objection nonetheless recurs, and it deserves acknowledging rather than dismissing. These measures delay the first visible result, which is hard to defend to a sponsor awaiting a demonstration.

The answer lies in an observation made above. A project delivering an early partial but measured result meets that expectation without abandoning rigour, whereas one delivering an early unmeasured result creates an expectation it will not be able to honour.

Earth observation warning signals

Six indicate a project is starting badly, and they are visible early.

Nobody can name the decision the result will inform.

The ground truth question has not been raised after several technical meetings.

The classes were taken from an existing nomenclature without examining their suitability.

Evaluation is planned on part of the same set as training.

No end user has taken part in the discussions.

And the schedule contains no corpus building phase, which indicates that work has not been identified.

What these failures offer a provider

Three commercial consequences follow from this inventory.

The first is that a scoping engagement has demonstrable value. An interlocutor able to pose the four feasibility questions before commitment spares a client a failure whose cost far exceeds that of the scoping.

The second is that the most frequent causes concern the reference, which is precisely the competence domain of an annotation provider.

And the third is that these failures are little discussed publicly, which makes setting them out differentiating. A provider able to name what makes a project fail demonstrates an experience no capability presentation replaces.

Where a provider can intervene

Not every Earth observation failure listed is one a data provider can prevent, and distinguishing them keeps a commercial promise honest.

Three families sit squarely within an annotation provider’s competence. Reference failures, which are the most frequent and which concern nomenclature, ground truth and conventions. Method failures on the evaluation side, since building a separate evaluation set and measuring annotator disagreement are both annotation work. And part of the durability family, since a maintained reference corpus is what makes drift measurable.

Two families sit partly within it. Scoping, where a provider can pose the feasibility questions and assess ground truth availability without deciding the project’s purpose. And integration, where a provider can deliver into the client’s environment without controlling whether the client’s users adopt it.

And two families sit outside it entirely. Organisational failures, which belong to the client’s internal situation. And the difficulty nobody flags, where a provider can raise the question at scoping and cannot resolve it.

That distribution matters commercially. A provider claiming to prevent every failure listed overstates and invites disappointment; one that names precisely which three it addresses makes a promise it can keep, which is the more durable position.

Common errors of reading

These misreadings recur often enough that naming them is usually enough to avoid them.

  • Launching a project on a capability rather than on a decision to inform.
  • Not verifying the phenomenon is visible at the available resolution.
  • Confusing announced revisit with usable acquisitions over the area.
  • Discovering the absence of ground truth after processing the imagery.
  • Taking an existing database as reference without knowing its ceiling.
  • Evaluating on the same areas as training.
  • Estimating area by pixel counting.
  • Delivering a result in a format foreign to the user’s practice.
  • Not planning a maintained reference corpus after entry into service.
  • Designing an arrangement without involving its end users.
  • Launching a resumption without diagnosing the family of the original difficulty.
  • Promising a complete result in a year rather than delivering an early measured partial one.

Why this list is uncomfortable to publish

A closing observation concerns this Earth observation article itself rather than its subject.

A provider setting out how projects fail is describing risks its own clients will run, some of which its own engagements have run. That is uncomfortable and it is the reason such lists are rare.

Two arguments make the discomfort worth accepting.

The first is that a buyer who has read this list asks better questions, which produces better-scoped projects and fewer disputes. A provider whose clients arrive already knowing that ground truth is the constraint spends its first meeting on method rather than on persuasion.

The second is that these failures happen whether or not they are discussed. Naming them does not create the risk; it moves the conversation from after the failure to before it, which is where it can still change an outcome.

That is the argument for the whole series, compressed. Everything in these fifteen articles concerns things that determine whether a project produces something defensible, and almost all of them cost little when addressed early and a great deal when addressed late.

What to take away

Earth observation projects rarely fail for technological reasons, the resource being massive and free, the tooling mature and the models available.

Three readings emerge. Seven families of cause are distinguishable by when they occur, scoping, access, reference, method, integration, durability and organisation, and the third concentrates most of the field’s difficulties. Most causes belong to decisions taken before any development, which are taken in a few meetings, cost little at that moment and become very expensive to correct, while producing no visible deliverable, an asymmetry that explains why they are neglected. And the cost of a failure grows with how late it surfaces, a scoping failure costing the entire project and a durability failure costing more than the initial build since the context has been lost.

For the forward view of this sector, the article on Earth observation trends examines the developments. For the economics of these trade-offs, the article on the cost of data details the lines.

To explore delivery arrangements, supported formats and applicable control mechanisms, see our dedicated page on geospatial data processing. And if you would like a project scoped before committing to it, let us discuss your project.

Tags

Découvrez nos articles