The question sounds simple and it is almost always answered wrongly, because the data itself is frequently the smallest line. The Copernicus article showed that an operator of the European platform states all functionalities are available free of charge for general users with some predefined quotas, and that commercial conditions apply where a user wants to download or process data on a large scale.
Free access to the raw material, therefore, and a cost that sits everywhere else. This article sets out where. It extends the article on validating Earth observation products.
The six lines of an Earth observation budget
Separating them is the precondition for any realistic estimate, since a project budgeting only the first two will overrun.
Data access, which may be free or paid depending on the regime, and which the market article described.
Processing infrastructure, meaning compute and storage, whose scale depends on volume and on whether processing happens locally or where the data sits.
Annotation, meaning the production of the training and reference corpora this series has described as the field’s scarce input.
Validation, meaning the reference sample and the statistical work the previous article set out.
Development, meaning the models and the processing chain.
And operation, meaning what it costs to keep the arrangement running once built.
One observation runs across those six. Only the first varies with the choice between free and commercial data, which is why that choice matters far less to a total budget than it appears to.
The cost of Earth observation data access
Two regimes coexist and their economics differ structurally.
The open regime costs nothing per scene and it costs in quotas. Beyond a certain processing volume, commercial conditions apply, which turns a free resource into an infrastructure line rather than a data line.
The commercial regime prices by scene, by area subscription, by volume subscription or by tasked acquisition, four models the market article described. Their unit costs differ by an order of magnitude between an archive scene and a guaranteed tasked capture.
Three factors dominate a commercial bill. Resolution, finer imagery costing substantially more. Freshness, an archive scene being worth markedly less than a recent acquisition. And the licence, internal use, redistribution and commercial exploitation not being priced the same way.
That third factor deserves particular attention. The market article noted that a licence restricting use to internal purposes forbids delivering a derived product, a constraint sometimes discovered late and which invalidates a business model rather than merely raising its cost.
The cost of infrastructure
This line has changed nature and the change is worth understanding.
The classical arrangement downloads scenes, stores them and processes them locally, which requires storage measured in terabytes and compute proportionate to the volume.
The arrangement the Copernicus article described processes where the data sits and returns only the result, which the platform documentation illustrated with an example turning a several-gigabyte need into a few kilobytes.
Three consequences follow for a budget. Local storage largely disappears as a line for projects adopting the second arrangement. Compute becomes a variable cost tied to actual use rather than a fixed investment. And the setup time falls, which reduces the human cost of the early phase more than the machine cost.
One caveat applies. Processing quotas exist and are exceeded by production volumes, which means the second arrangement is free for exploration and paid at scale, a threshold worth estimating at scoping rather than discovering.
The most underestimated line
Annotation is regularly absent from an initial budget and it frequently exceeds every other line.
Three reasons explain the omission. It is invisible in a technical architecture, which is where a project is usually designed. It is assumed to be resolvable by a public dataset, an assumption the AI article showed fails for any specific nomenclature or underrepresented region. And its cost is genuinely hard to estimate before the conventions are fixed.
That third reason is real and it has a remedy. A pilot on a small sample establishes the throughput, reveals the ambiguous cases and produces the case library, which converts an unknown into a rate. The segmentation cluster recommended that pilot for quality reasons and it serves budgeting equally.
One observation completes this. A project that discovers the annotation cost after committing to a schedule either cuts the corpus, which caps what the model can do, or overruns. Both outcomes are avoidable by a pilot costing days.
What determines Earth observation annotation cost
Six factors drive it and knowing them allows an estimate before any quotation.
The geometry required. A point costs a fraction of a polygon, and a precise boundary costs several times an approximate one.
The number of classes and their separability, an ambiguous distinction requiring expert arbitration on a substantial share of cases.
The object density, a scene containing hundreds of objects costing far more than one containing a few, at equal area.
The domain expertise required, which the opening article showed is constant in geospatial work and which sets the hourly rate.
The temporal dimension, the constellations article having shown change annotation requires distinct conventions and costs appreciably more.
And the quality requirement, an evaluation corpus needing multiple reading and documented arbitration where a training corpus does not.
Those six multiply rather than add, which is why estimates vary by an order of magnitude across projects that superficially resemble one another.
The cost of Earth observation validation
This line is distinct from annotation and it is smaller in volume and higher in unit cost.
The previous article showed a rigorous assessment requires hundreds to thousands of reference points under a probabilistic sampling design, with per-class figures and confidence intervals.
Three components make it up. Building the reference sample, which is annotation work at a higher quality standard. The statistical analysis, which is a distinct competence from annotation. And where relevant, field collection, which is the most expensive component and the one that cannot be compressed.
One observation matters for the trade-off. The AI article recommended concentrating expert effort on a restricted evaluation set rather than spreading it evenly across a large training corpus, which is a budgetary recommendation as much as a methodological one.
The line nobody budgets
One cost appears after delivery and it is almost never anticipated at proposal stage.
A model degrades. The standard developments article documented why: sensors change, the ground changes, and processing chains change. A product deployed without provision for revalidation will need revalidating regardless.
Three items make up that recurring line. Periodic revalidation on a maintained reference corpus, which the security article showed institutional buyers now require. Corpus extension when the product covers new geography, which the AI article showed does not transfer freely. And harmonisation when sources change, which the constellations article showed happens as generations are replaced.
The consequence is that an Earth observation project has a running cost and not only a build cost, and a proposal presenting only the latter is understating what the client will actually spend.
Orders of magnitude by Earth observation project type
A rough gradation helps at scoping, with the caveat that any figure depends on the six factors above.
An exploratory project on one region, one modality and an existing nomenclature is measured in weeks and its dominant line is human time rather than data or infrastructure.
An operational product on a defined territory is measured in months, with annotation and validation together representing the majority of the effort.
A multi-region product is measured in a year or more, geographic extension requiring corpora that do not transfer, which the AI article documented.
And a product entering an institutional or contractual context adds the documentation and traceability the security and insurance articles described, which is a modest cost when planned and a large one when retrofitted.
How to estimate an Earth observation project
Five steps produce a defensible estimate rather than an optimistic one.
Establish what the data must support, since the decision determines the resolution, the revisit and therefore the data regime.
Verify availability over the target area and period, cloud cover and revisit being able to invalidate a project before any budget is discussed.
Run an annotation pilot on a small sample to establish the throughput and reveal the ambiguous cases.
Size the validation according to the number of classes and the precision sought, following the previous article.
And add the running line, which is the one most proposals omit.
The trade-offs that actually reduce cost
Four levers work and they are not the ones usually pulled.
Narrowing the question. A product answering one question well costs a fraction of one answering several approximately, and it is generally more useful.
Reducing the nomenclature. Each additional class costs annotation, costs validation and reduces per-class accuracy, which the previous article showed.
Using an existing thematic product where one exists, which the Copernicus article recommended and which a check of an hour establishes.
And reusing a corpus across projects, which requires having documented it well enough to be reusable, a condition few corpora satisfy.
That fourth lever is the one with the longest payback. A documented corpus serves a second project at a fraction of the first project’s cost, and an undocumented one serves nothing.
The false economies in Earth observation
Four apparent savings cost more than they save.
Reducing the evaluation corpus, which saves annotation and produces a figure nothing supports, invalidating any claim about the product.
Using an existing map as ground truth, which the AI article showed caps achievable accuracy at that map’s own and imports its nomenclature.
Skipping the pilot, which saves days and produces conventions discovered mid-production, requiring rework the segmentation cluster documented.
And omitting documentation, which saves hours during production and makes the corpus unusable for any later claim, unreusable for a second project, and indefensible in an audit.
Those four share a structure. Each saves a visible cost now and creates an invisible one later, which is the configuration that makes them attractive and wrong.
Comparing two Earth observation quotations
Six questions make two proposals comparable, since headline prices rarely are.
What geometry and what precision, since a point and a precise polygon differ by a large factor.
What quality control, meaning whether a second pass exists and on what proportion.
Whether the evaluation corpus is included or separate, a proposal excluding it appearing cheaper and delivering less.
What documentation accompanies the delivery, composition and provenance being what allows the corpus to be used later.
Who arbitrates ambiguous cases and how those decisions are recorded.
And what happens on the cases that cannot be annotated, an exclusion log being the difference between a known gap and an invisible one.
A proposal answering those six may cost more per unit and less per project, which is the comparison that matters.
The cost of doing nothing
One line rarely enters the comparison and it belongs there.
The alternative to an Earth observation project is generally a manual process: field visits, declarative surveys, expert inspection or simply no monitoring at all.
Three comparisons apply. Against field visits, satellite monitoring is cheaper per unit area and less precise per unit, which makes the trade-off depend on the density of what is being monitored. Against declarative surveys, it is independent rather than reported, which the climate article showed produces different answers. And against no monitoring, it changes what is knowable rather than what is cheaper.
That third comparison is the honest one for many projects. The justification is not that the satellite is cheaper than an existing method, but that no existing method covers the area at the frequency required.
What the Earth observation economic model implies
One structural observation follows from the whole picture.
A sector where the raw material is free and the interpretation is expensive has an unusual cost structure. Entry is cheap and scale is not, which is the reverse of most data industries.
Three consequences hold. Prototyping costs little, which explains the number of entrants the market article noted. Reaching production costs substantially more, which explains why many do not. And the difference between the two is almost entirely annotation, validation and documentation.
That is the practical summary of this cluster expressed as a budget. The gap between a demonstration and a product is not technical, and it is not data. It is the corpus and the evidence, which are the lines a proposal is most tempted to trim.
The Earth observation spending calendar
One timing observation completes the estimate, since cash matters as much as totals.
Data and infrastructure costs are spread and variable, rising with use.
Annotation costs are concentrated at the start, before any result exists, which is the hardest moment to justify a spend.
Validation costs arrive at the end of the build, when the budget is already committed.
And running costs arrive after delivery, when the project is formally closed and nobody owns them.
That distribution explains a common failure. The two lines that determine whether a product can be defended, annotation and validation, fall at the two moments where a budget has least flexibility, which is why they are the ones cut.
Build or buy the corpus
One Earth observation decision recurs and it is usually taken implicitly rather than examined.
A project needing annotated data can build it internally, contract it out, or attempt to assemble it from public sources. The three differ in more than price.
Building internally suits a project with a long horizon and a recurring need, since the domain knowledge accumulates. It costs recruitment, training and management, and it becomes economic only above a certain sustained volume.
Contracting out suits an intermittent or peaked need, which the insurance article showed characterises post-event work. It costs a margin and it buys throughput, method and the ability to stop.
Assembling from public sources costs least and, as the AI article documented, imports a nomenclature, a granularity and a date the project did not choose, capping what the result can support.
One factor decides between the first two more reliably than volume does. If the annotation conventions are stable and the domain is one the organisation will work in for years, internal capability compounds. If the conventions change with each client or the need arrives in bursts, external capacity is the cheaper structure even at a higher unit rate.
Presenting a budget to a client
One practical point closes the operational side, since a correct estimate presented badly still loses the work.
Four framings help a client accept a budget whose largest line is annotation.
State the free data explicitly. A client who knows the imagery costs nothing understands why the remaining lines carry the total, and a proposal that leaves this implicit invites the wrong comparison.
Separate build from run. Presenting both prevents the conversation where a competitor’s build-only figure looks cheaper and the difference surfaces a year later.
Show what each line buys in terms of claims. Validation is not a quality gesture; it is what allows a number to be stated publicly, which the previous article established.
And offer the pilot as a priced first step. It converts an argument about estimates into a measurement, it costs days, and it lets both parties commit to the main phase on evidence rather than on assertion.
That fourth framing resolves most budget disagreements in this field, because the disagreement is usually about uncertainty rather than about price.
Common errors of reading
These misreadings recur often enough that naming them is usually enough to avoid them.
These misreadings recur often enough that naming them is usually enough to avoid them.
- Equating free data with a cheap project.
- Budgeting data and development while omitting annotation and validation.
- Discovering processing quotas after sizing a production run.
- Overlooking a licence restriction that forbids delivering a derived product.
- Skipping the annotation pilot that converts an unknown into a rate.
- Reducing the evaluation corpus to save cost.
- Using an existing map as ground truth without knowing its ceiling.
- Omitting documentation, which makes the corpus unreusable.
- Presenting a build cost without the running cost.
- Comparing quotations on unit price rather than on scope.
- Adding classes without accounting for their annotation and validation cost.
- Assuming a corpus transfers to a new region without extension.
Why the cheapest input is the most expensive line
A closing observation names the inversion this Earth observation article describes.
In most data industries, acquiring the raw material dominates the budget and processing it is comparatively cheap. Earth observation reverses that. The imagery is free or nearly so, and what it costs to know what the imagery shows exceeds everything else.
Three consequences follow for how a project should be argued internally.
A budget dominated by annotation is not a badly designed budget. It is what a correctly designed budget looks like in a sector where the input is free, and a proposal where annotation is a minor line is more likely to be incomplete than efficient.
The lines that can be cut without visible short-term consequence are precisely the ones that determine whether the result can be defended, which makes budget discipline here a matter of resisting a temptation rather than finding savings.
And the asset a project retains at the end is the documented corpus, not the model, since the model will be superseded and the corpus will still describe what was on the ground.
That third point is the one worth carrying into any budget conversation. The spend that looks like a cost is the spend that produces the thing that lasts.
What to take away
Earth observation data can be free while the project is not, since only one of six budget lines varies with the choice between open and commercial data.
Three readings emerge. Annotation is regularly absent from an initial budget and frequently exceeds every other line, an omission a pilot of a few days converts into a measurable rate. The running cost of revalidation, corpus extension and source harmonisation is almost never anticipated, which means a proposal showing only the build cost understates what the client will spend. And the false economies share a structure, each saving a visible cost now and creating an invisible one later, which is exactly why the corpus and the evidence are the lines a budget is most tempted to trim.
For the failures these misjudgements produce, the article on why Earth observation projects fail examines the causes. For the quality requirement that shapes these costs, the article on validating products sets out the method.
To explore delivery arrangements, supported formats and applicable control mechanisms, see our dedicated page on geospatial data processing. And if you need a defensible estimate for an annotation and validation programme, let us discuss your project.