Geospatial Foundation Models

A model trained on millions of satellite images promises to reduce the annotation work. The promise is real, its scale much less so than announced, and what it leaves to be done looks remarkably like what was already costing the most.

This article sets out that contribution. It extends the article on the cost of annotation.

What these models bring

Five benefits are observed in practice on a satellite imagery project.

A representation already suited to strictly vertical views.

Recognition of textures as well as of the most common patterns.

A markedly reduced need for annotated examples to start.

Better performance over the least represented territories.

And appreciably easier transfer from one region to another.

One important practical consequence follows. Those five benefits all bear on the start rather than on final performance, which appreciably shortens the delay before the first useful result without changing the ceiling the project will eventually reach.

What they do not remove the need to do

Five satellite imagery tasks stay whole despite these models.

Writing the reference proper to the client’s exact need.

Settling the limiting cases the model ignores entirely.

Building a genuinely representative evaluation set.

Annotating the rare classes the pretraining had not seen.

And measuring what the result obtained is really worth.

One important observation follows for a satellite imagery project. Those five tasks already constituted on their own the most expensive share of a project, which explains why a gain announced on annotated volume very rarely translates into a proportional fall in budget.

What adaptation demands

Five elements condition a successful adaptation in satellite imagery.

A set of examples annotated to the reference finally targeted.

Deliberate coverage of difficult cases rather than easy ones.

Faithful representation of the territories genuinely concerned.

An evaluation set separated on a geographic basis.

And a measurement established before and then after the adjustment.

One practical consequence follows. The second element decides the real contribution, an adjustment conducted on examples that are too easy improving nothing where the model was already failing, which makes the selection of examples markedly more determining than their number alone.

What the pretraining actually saw

Five characteristics of a satellite imagery model are worth knowing.

The sensors the training images genuinely came from.

The resolutions effectively covered as well as those missing.

The regions of the world effectively represented in the corpus.

The seasons as well as the acquisition conditions present.

And the spectral bands effectively taken into account.

One important practical consequence follows for a satellite imagery project. The third characteristic explains most disappointments, a model trained mainly on a few regions of the world performing badly over a territory with very different building methods and landscapes, which reproduces exactly the defect these models were supposed to correct.

What the measurement must establish

Five comparisons reveal such a satellite imagery model’s true contribution.

The performance obtained with and then without pretraining.

The number of examples needed to reach an equivalent result.

The performance obtained on rare rather than common classes.

The behaviour observed over an unrepresented territory.

And the human time that is genuinely saved.

One important observation follows. The last comparison is the only one that interests a budget and the least often produced, a gain measured in annotated examples never translating mechanically into hours saved.

What the annotation provider becomes

Four shifts affect satellite imagery annotation work.

The raw volume annotated falls markedly on the easy cases.

The share given to difficult cases rises correspondingly in proportion.

Verifying a proposal sometimes comes to replace tracing.

And building the evaluation set gains in importance.

One important practical consequence follows for a satellite imagery project. Those four shifts make the remaining work harder on average, the easy share disappearing first, which raises the unit cost of what remains while reducing the total volume.

What these models change about the corpus

Five shifts affect the very building of a satellite imagery dataset.

The total volume of examples needed falls appreciably.

Selecting the examples becomes more determining than their number.

Difficult cases must be deliberately over-represented in it.

The evaluation set keeps exactly the same requirement.

And rare classes still call for exactly as many examples.

One important practical consequence follows. The last shift limits the gain where it would have served most, a class entirely absent from the pretraining demanding the same annotation effort as before, which concentrates the whole saving on the objects one already knew how to detect.

What the client must check before believing it

Four points separate a serious satellite imagery demonstration from a mere promise.

Is the gain measured on their own reference or else on a public set.

Is the evaluation set geographically separated from the adjustment set.

Is performance broken down by class rather than given globally.

And is human time counted rather than the number of examples alone.

One important observation follows for a satellite imagery project. The first point separates them immediately, a result obtained on a public set saying roughly nothing about what the model will produce on a reference of one’s own, which makes a demonstration run on a few dozen of the client’s tiles far more informative than any published figure.

What throughput becomes with these tools

Four effects modify the working rhythm in satellite imagery.

The time per tile falls markedly on well covered scenes.

The time per difficult case stays unchanged, or even rises.

The proportion of genuinely fast tiles varies greatly with the territory.

And the vigilance demanded rises correspondingly since routine disappears.

One important practical consequence follows for a satellite imagery project. The last effect deserves attention and is often neglected, a day composed solely of difficult cases tiring far faster than a day alternating routine tracing and adjudication, which recalls the findings of the chapter on visual fatigue.

What these models do not change

Five findings from the course stay intact in satellite imagery.

A footprint or continuity convention is still decided with the client.

An omitted object stays undetectable by any control on what was produced.

A set separation left random still distorts the measurement.

A result delivered too late still directs no decision.

And a proof of compliance remains as demandable as before.

One important practical consequence follows. Those five findings bear on scoping and on measurement rather than on execution, which explains why no progress in the models comes to touch them and why a badly scoped project fails in exactly the same way as ten years ago.

What the first trial must establish

Four results justify a trial before any budget commitment.

The performance obtained on the client’s own specific reference.

The number of examples needed to reach an agreed threshold.

The performance obtained on classes the pretraining had never seen.

And the human time observed on a batch comparable to an ordinary one.

One important observation follows for a satellite imagery project. The last result settles the debate more surely than all the others, a direct comparison between two equivalent batches handled with and without the model giving a figure neither the provider nor the vendor can contest, which makes this trial more useful than a long technical discussion.

What the contract must provide for here

Four clauses are worth writing on this kind of satellite imagery project.

The performance threshold from which the result is accepted.

The division of work between the automatic proposal and human rework.

The fate of the annotated corpus should the model change mid-project.

And the agreed measurement of the gain, expressed in time rather than volume.

One important practical consequence follows for a satellite imagery project. The third clause is almost always neglected and becomes expensive, a model change decided along the way voiding part of the adjustment work already paid for, which justifies stating from the outset who bears that expense.

Three errors proper to this subject

Three satellite imagery defects accompany the adoption of these tools.

A budget built on a gain expressed in examples rather than in real hours.

An adjustment conducted on the cases the model already handled well.

And a quality control lightened on the sole grounds that the model is good.

Those three defects are discovered after the budget commitment, they belong to how the gain is measured rather than to the tool itself, and their common point is to confuse a reduction in volume with a reduction in cost.

What this subject shares with pre-annotation

Four satellite imagery findings recur across these two approaches.

The gain bears above all on the cases one already knew how to handle.

A proposal at once directs the gaze towards what it has found.

Doubt is expressed far more rarely than in free annotation.

And quality control becomes far more necessary, not less.

One important practical consequence follows for a satellite imagery project. Those four findings suggest that foundation models extend pre-annotation rather than replacing it, which makes every precaution established in the chapter on quality applicable here.

What the provider brings here

Four contributions distinguish a satellite imagery engagement relying on these tools.

A selection of examples deliberately directed towards difficult cases.

An evaluation set built before any adjustment at all.

A measurement of human time established before and then after.

And a quality control kept in full rather than lightened.

One important practical consequence follows. The first contribution constitutes the provider’s principal value in this framework, a model improving precisely and only where it is shown its own failures, which shifts the competence expected from producing volume towards identifying what is missing.

What the provider must tell the client

Four satellite imagery messages deserve carrying rather than leaving unsaid.

The gain does exist but bears on the start rather than on the ceiling.

The share that disappears is precisely the one that cost least.

Quality control retains rigorously the same usefulness.

And a comparative trial settles matters far better than any argument.

One important observation follows for a satellite imagery project. The second message carries badly and verifies easily, a provider announcing it appearing to slow the adoption of a tool the client often expects much from, whereas they are precisely avoiding the disappointment that will follow a budget built on a badly measured promise.

What the client can ask to see

Four elements are worth asking for before selecting a satellite imagery solution.

A demonstration run on their own tiles rather than on a public set.

Performance detailed class by class rather than a single global figure.

A precise description of what the pretraining actually covered.

And the exact protocol by which the announced gain was measured.

One important practical consequence follows for a satellite imagery project. The last element is rarely refused and yet seldom obtained, a protocol left uncommunicated often signalling a measurement conducted in conditions that will not recur, which suffices to set a solution aside without having to test it.

Approaching a project built on these models

Five questions scope such a satellite imagery project.

What gain is announced and how is it measured. Often in examples.

Is the client’s reference covered. Rarely in full.

Are the rare classes seen. Pretraining ignores them.

Is the evaluation set separated. Otherwise the measurement misleads.

And is human time measured. That is the only budgetary gain.

Those five answers separate a real contribution from a promise. Asking them before committing avoids building a budget on a gain that will not materialise.

The question that frames the use

One question determines what these tools contribute in satellite imagery.

Does the need resemble what the model has already seen.

A close need, built form, vegetation or roads over an ordinary territory, draws a real benefit from pretraining and starts with few examples.

A distant need, a particular domain class or an atypical territory, recovers the requirements of an ordinary project and sometimes more.

That question belongs to the need and not to the tool, it is asked before any budget commitment, and it separates a substantial saving from an expensive disappointment.

Three decisions before relying on such a model

Three decisions commit a satellite imagery project built on these tools.

Verifying that the pretraining covers the territory as well as the sensor targeted.

Building an evaluation set well before any adjustment.

And requiring a measurement of human time rather than of annotated volume alone.

Those three decisions cost a few days, they precede the budget commitment, and their absence leads to funding a project on a gain that will not materialise.

Three checks on a promised gain

Three checks qualify the contribution announced for a satellite imagery model.

The pretraining’s coverage, set against the territory as well as the sensor targeted.

The geographic separation between the evaluation set and the adjustment set.

And the unit the gain is expressed in, examples or else hours.

Those three checks are each asked in one question, they require no competence in machine learning, and their absence indicates a promise whose scope nothing permits assessing.

What this course has established

Four findings run through this course’s fifteen chapters.

A convention left unwritten is always discovered far too late.

A measurement left global masks failure on what really counts.

What is prepared in calm conditions determines what production can reach.

And the difficulty always sits in the judgement rather than in the hand.

One important observation follows for a satellite imagery project. The last finding explains all the others, work whose difficult share belongs to adjudication resisting automation far longer than repetitive work, which makes these fifteen chapters markedly more durable than the tools they mention.

What this chapter concludes about the trade

Three traits characterise the annotation trade as this course has described it.

Value shifts towards scoping as well as measurement rather than towards production.

The rarest competence becomes identifying what the corpus is missing.

And the provider’s role now belongs more to advice than to execution.

One important observation follows for a satellite imagery project. Those three traits strengthen with every advance in the models, which makes the trajectory predictable even where its pace is not, and which invites a provider to invest in judgement rather than in production capacity alone.

What this chapter suggests for what follows

Four developments are taking shape without yet being settled.

A pretraining covering more regions as well as more sensors.

Native handling of the bands lying beyond the visible.

Better performance on rare as well as small objects.

And far more direct integration into annotation tools.

One important observation follows for a satellite imagery project. Those four developments will reduce the automatable share of the work without touching scoping or measurement, which suggests the decisions described across this course will keep their value even as throughputs change.

What this course leaves to those who extend it

Three subjects stay open at the end of these fifteen chapters.

Building references shared between the actors of one same field.

Preserving quality measurements lastingly, well beyond a single project.

And comparing corpora produced to initially different conventions.

One important observation follows for a satellite imagery project. Those three subjects belong to coordination between organisations rather than to technique, which makes them slower to advance than the models themselves and probably more determining for the value of the corpora already built.

Why saying this plainly is the better position

One point about how to talk about these tools belongs at the close.

Understating the gain costs a provider less than overstating it.

Three reasons follow in satellite imagery.

An overstated gain is discovered by the client, not confessed by the provider.

The discovery lands mid-project, when the budget is already committed.

And it recasts every other claim the provider has made as suspect.

One important practical consequence follows for a satellite imagery provider. Setting expectations low and then beating them is the only version of this conversation that ends well, since a client who was promised a modest saving and received it will trust the next estimate, whereas one who was promised a transformation and received a modest saving will scrutinise everything that follows.

What closes this cluster

Three claims survive every chapter of this satellite imagery course.

A convention decided late costs more than one decided badly.

A measurement that hides its failures is worse than no measurement.

And the work that resists automation is the work that requires a decision.

One important observation follows. Those three claims were true before these models existed and will hold after the next generation, which is why a provider is better served by building judgement into their practice than by tracking whichever tool is current.

Common mistakes

These failures recur often enough that naming them is usually enough to avoid them.

  • Building a budget on a gain announced in annotated examples.
  • Adjusting a model on examples that are easy to annotate.
  • Neglecting the rare classes the pretraining ignores.
  • Evaluating without a geographically separated set.
  • Assuming the client’s reference is already covered.
  • Confusing a reduction in volume with a reduction in cost.
  • Omitting the measurement before and after adjustment.
  • Accepting an automatic proposal without control on rare cases.
  • Believing quality control loses importance.
  • Comparing two models with no shared evaluation reference.

What to take away

These models shift the annotation work rather than removing it.

Three readings emerge. The tasks they do not remove the need to do already constituted the most expensive share of a project, which explains why a gain announced on annotated volume rarely translates into a proportional fall in budget. The easy share of the work disappears first, which raises the unit cost of what remains while reducing the total volume and makes an average rate misleading from one generation to the next. And a gain measured in annotated examples does not translate mechanically into hours saved, which makes human time the only measurement that interests a budget.

For a project’s economics, the article on the cost of satellite annotation details the lines. To return to the overview, the guide to satellite imagery annotation sets it out.

To explore delivery arrangements, supported formats and applicable control mechanisms, see our dedicated page on annotation for geospatial. And if you are preparing a geospatial project, let us discuss your need.

Tags

Découvrez nos articles