Geospatial Annotation Quality – Small Objects and Visual Fatigue

An annotator scanning an empty tile for twenty minutes produces nothing visible. That is nonetheless where a geospatial corpus’s quality is decided, since the dominant error is not a bad outline but an object nobody saw.

This article sets out those difficulties. It extends the article on preparing a dataset.

What separates this field’s errors

Five annotation defects occur in satellite imagery.

An object plainly present and not annotated, invisible in the deliverable.

An object annotated that quite simply did not exist in the scene.

A class wrongly assigned between two neighbouring categories.

An imprecise outline traced on a correctly spotted object.

And an inconsistency observed between two adjacent tiles.

One important practical consequence follows. The first defect largely dominates the others in this field, an omitted object leaving absolutely no trace in the delivered data, which makes it undetectable by any control bearing on what has been produced.

What the smallness of objects imposes

Five difficulties accompany a satellite imagery object a few pixels across.

Its presence is guessed at far more than it is observed.

Its outline holds only very few distinct points.

Two operators delimit it differently without either being wrong.

The zoom required reduces the field being scanned accordingly.

And visual fatigue sets in far faster.

One important observation follows for a satellite imagery project. The third difficulty distorts any disagreement measurement, a gap of only two pixels representing a considerable share of a small object, which produces high divergence rates without any operator having made an error.

What prolonged scanning produces

Five satellite imagery effects accompany hours of visual search.

The detection rate declines steadily across the session.

Attention slackens first on the least dense areas.

The threshold for flagging a doubt rises progressively.

Objects near the visibility threshold are the first to be lost.

And the operator themselves does not perceive their own decline.

One practical consequence follows. The last effect forbids relying on self-regulation, a tired annotator judging themselves just as attentive as at the start of their session, which requires organising the work rather than issuing an instruction about vigilance.

What the rarity of objects changes

Four satellite imagery effects accompany a low object density.

One tile in several holds no object at all to annotate.

The search then occupies most of the working time.

Expectation of an object falls as that object fails to appear.

And a control by sample encounters only few positive cases.

One observation follows. The third effect is well documented and rarely taken into account, the probability of missing an object growing with the very rarity of what one is looking for, which makes a prior sorting of tiles more effective than any exhortation to vigilance.

What the control must look for here

Five satellite imagery arrangements detect different defects.

A double annotation conducted on the same tiles.

A rereading specifically targeted at the densest areas.

A reconciliation conducted with an existing reference base.

A verification of junctions carried out after reassembly.

And a comparison of the detection rates obtained between operators.

One important practical consequence follows. The first arrangement is the only one that reveals omissions, a simple rereading being able to flag only what has already been annotated, which makes it blind to this field’s principal defect.

What organising the work permits

Five satellite imagery measures act on fatigue rather than on instruction.

Sessions of strictly limited duration on this kind of task.

An organised alternation between searching and tracing.

A prior sorting of the tiles liable to hold objects.

A regular rotation of the areas between the operators.

And a production rhythm monitored throughout the day.

One important observation follows for a satellite imagery project. The third arrangement produces the clearest gain, an operator working on tiles that are almost all populated staying attentive far longer than one scanning empty expanses.

What double annotation costs and returns

Four contributions justify this control arrangement in satellite imagery.

An estimate of the omission rate, impossible to obtain otherwise.

A spotting of the classes two operators come to separate differently.

A measurement of the variability proper to small objects.

And a flagging of the precise areas where difficulty concentrates.

One important practical consequence follows. Those four contributions are obtained on a fraction of the corpus rather than on the whole of it, only a few percent of tiles annotated twice sufficing to estimate what is missing, which makes this arrangement far less expensive than its reputation suggests.

What the measurement must reflect

Five indicators describe the quality of a satellite imagery corpus.

The omission rate estimated by the double annotation.

The detection gap observed between operators over the same areas.

The evolution of the detection rate observed across a session.

The consistency of the outlines obtained on the smallest objects.

And the number of inconsistencies found at tile junctions.

One important practical consequence follows. The third indicator is easily produced and rarely asked for, a comparison between the start and the end of a session revealing immediately whether the organisation of the work allows for fatigue.

What training changes here

Five gains distinguish a satellite imagery operator trained for this kind of task.

A systematic scanning method rather than a wholly free traverse.

Knowledge of the positions where the object is most often met.

The spotting of indirect clues that come to betray a presence.

The habit of flagging a doubt rather than settling it alone.

And a lucid awareness of their own decline in form.

One important observation follows for a satellite imagery project. The first gain produces the most measurable improvement, a methodical traverse of the tile finding markedly more objects than a mere overall examination, which makes a scanning instruction one of the rare training contributions whose effect shows from the very first day.

What pre-annotation changes about quality

Four effects accompany an automatic proposal made in satellite imagery.

The omission rate falls markedly on well contrasted objects.

The operator then moves from spotting to mere verification.

Objects absent from the proposal are noticed markedly less.

And doubt is expressed far more rarely than in free annotation.

One important practical consequence follows. The third effect deserves attention, an automatic proposal at once directing the gaze towards what it has already found, which reduces omission on the easy cases while worsening the omission that already bore on the difficult ones.

What the omission rate reveals

Five readings follow from an omission measured in satellite imagery.

Its absolute value, very rarely as low as one imagines.

Its distribution by object size, almost always very uneven.

Its evolution observed across a working session.

Its gap observed between operators over the same areas.

And its concentration on certain kinds of visual context.

One important observation follows for a satellite imagery project. The second reading usefully directs the effort, an omission concentrated on the smallest objects calling either for a higher resolution or for an accepted abandonment of that class, whereas a uniform omission designates a problem of method or of organisation.

What the nature of the scene adds

Five satellite imagery contexts make the visual search harder.

A background whose texture strongly resembles that of the object.

A very cluttered scene within which the object is lost.

A cast shadow area that lowers the local contrast.

An image whose overall tone remains barely contrasted.

And an object partly masked by an immediately neighbouring element.

One important observation follows for a satellite imagery project. The first context produces the most systematic omission, an object whose texture exactly matches that of the background disappearing for every operator rather than for one of them alone, which makes it invisible even to a double annotation.

What the working tool changes here

Five interface characteristics bear on the satellite imagery omission rate.

The fluidity of movement within a tile of large size.

The possibility of marking an area as already scanned.

The speed of moving from one zoom level towards another.

Immediate access to neighbouring tiles for edge objects.

And a simple means of flagging a doubt without having to leave the tile.

One important practical consequence follows for a satellite imagery project. The second characteristic acts directly on the dominant defect, an operator unable to know what they have already examined returning to some areas and plainly forgetting others, which produces an omission of purely organisational origin.

What the first batch must establish

Four results justify a trial batch before production.

The omission rate estimated from a restricted double annotation.

The proportion of tiles holding no object across the territory handled.

The exact duration beyond which the detection rate declines.

And the disagreement actually observed on the corpus’s smallest objects.

One important practical consequence follows for a satellite imagery project. The second result decides the organisation to put in place, a high proportion of wholly empty tiles justifying a prior automatic sorting, whereas a densely populated territory makes that arrangement useless and permits concentrating the effort on rereading.

What the corpus must hold to be testable

Five elements make a satellite imagery corpus verifiable.

A fraction of the corpus annotated twice by different operators.

Tiles deliberately covering the most difficult visual contexts.

Contiguous areas permitting the junctions to be controlled.

A sample reconciled with an external reference base.

And the lasting preservation of the disagreement measurements obtained.

One important observation follows for a satellite imagery project. The last element is almost always lost when a project closes, a disagreement rate that has not been recorded becoming impossible to reconstitute afterwards, which deprives the corpus of the only information that would have permitted judging its reliability years later.

What this chapter shares with annotation errors

Four findings recur whatever the field being annotated.

The order of severity of errors remains the exact inverse of their visibility.

A reviewer sharing the operators’ biases confirms the corpus without testing it.

Flagging a doubt grows rarer once it slows production down.

And a disagreement rate not preserved stays irreconstitutable afterwards.

One important observation follows for a satellite imagery project. Those four findings hold outside geospatial work, what the field adds being only the scale of the first, an omission becoming here the dominant defect rather than one category of error among others.

Three errors proper to this field

Three satellite imagery defects belong to visual search itself.

An omission that leaves absolutely no trace in the delivered data.

A drop in vigilance the operator concerned does not perceive themselves.

And a disagreement bearing on small objects, wrongly taken for an error.

Those three defects escape the usual control, they are handled by organisation rather than by rereading, and their common point is to bear on the operator’s gaze rather than on their hand.

What the client should ask a provider

Four questions reveal a serious practice in satellite imagery.

How they precisely estimate what their operators did not see.

What maximum session duration they adopt on this kind of task.

Whether they break their quality measurements down by object size.

And what they actually do with tiles holding no object.

One important observation follows for a satellite imagery project. The first question separates them immediately, a provider answering simply that they reread their work not having understood that rereading stays blind to absences, whereas an answer mentioning a double annotation even partial indicates an understanding of the dominant defect.

What this chapter changes about the client relationship

Three consequences follow from adopting these satellite imagery arrangements.

The quotation carries a line competitors do not have.

The quality report states an omission rate rather than passing over it.

And the conversation bears on what is missing rather than on what was produced.

One important observation follows for a satellite imagery project. The second consequence causes needless worry, a provider stating a measured omission rate appearing less capable than a competitor who measures none, whereas they are the only one of the two who knows what they are delivering.

What the provider brings here

Four contributions distinguish an engagement seriously conducted in satellite imagery.

A double annotation provided for from the outset rather than added afterwards.

A sorting of tiles conducted before even handing them to the operators.

An organisation of the work that genuinely allows for fatigue.

And measurements broken down by object size within the quality report.

One important practical consequence follows. The first contribution shows on the quotation and justifies itself badly without explanation, a double annotation adding only a few percent to the volume billed, which makes it necessary to explain to the client that this extra cost buys the only measurement of what is missing.

Approaching quality on this kind of project

Five questions scope quality on such a satellite imagery project.

Is the omission rate measured. It dominates the other defects.

Does a double annotation exist. Nothing else reveals absences.

Are the tiles sorted. Searching costs more than tracing.

Are sessions limited. Fatigue is not corrected by instruction.

And does disagreement allow for size. Small objects diverge naturally.

Those five answers determine the corpus’s reliability. Asking them before production avoids a control that validates what exists without seeing what is missing.

The question that frames the control

One question determines the arrangement to put in place in satellite imagery.

Does missing nothing matter most, or tracing well.

Missing nothing, an inventory, a count or surveillance, requires a double annotation, a sorting of tiles and short sessions, geometry mattering little.

Tracing well, a surface calculation or a land registry update, requires a written convention and a control of outlines, exhaustiveness mattering less than accuracy.

That question belongs to the corpus’s use and not to the field, it is asked before production, and it separates two control arrangements whose cost and nature bear no comparison.

Three decisions before producing

Three decisions commit the quality of a satellite imagery project.

Providing for a double annotation on a fraction of the corpus, the only way to estimate omissions.

Sorting the tiles before handing them over, searching costing more than tracing.

And limiting session duration, fatigue being corrected by no instruction.

Those three decisions cost one meeting and a few percent of volume, they precede production, and their absence produces a control that validates what exists without seeing what is missing.

Three checks on a quality arrangement

Three checks qualify a quality approach in satellite imagery.

The existence of a double annotation, even on a small fraction of the corpus.

The breakdown of measurements by object size rather than in overall value.

And the monitoring of the detection rhythm across a working session.

Those three checks are each asked in one question, they require no particular tool, and their absence indicates an arrangement blind to this field’s principal defect.

What this chapter teaches

One cross-cutting observation deserves closing this examination.

This field’s principal defect is an absence.

Three findings compose it.

An omitted object leaves absolutely no trace in the delivered data.

A simple rereading can never flag anything but what has already been annotated.

And a tired annotator judges themselves just as attentive as at the start of their session.

That finding extends the preceding chapter, preparation having removed the technical defects while being powerless against those belonging to human attention.

What this chapter leaves to the next

Some projects add constraints quality alone does not cover.

Three questions stay open in satellite imagery.

What a regulated context imposes on the annotation arrangement.

How to trace who accessed which image and when.

And what confidentiality changes about the choice of a provider.

Those three questions belong to defence contexts, which constitute the subject of the following chapter.

Why measuring omissions is worth the discomfort

One point about competing on quality belongs at the close.

Measuring what you missed produces a number nobody enjoys presenting.

Three reasons follow in satellite imagery.

The number is always higher than anyone expected going in.

Competitors who measure nothing appear to have no such problem.

And the client has no independent way to check either claim.

One important practical consequence follows for a satellite imagery provider. Presenting the number anyway is the strongest position available, since a client who has seen one provider quantify their own limits and another simply assert quality has learned which of the two knows what they deliver, and that lesson tends to outlast the discomfort of the first conversation.

Common mistakes

These failures recur often enough that naming them is usually enough to avoid them.

  • Controlling quality without ever measuring the omission rate.
  • Relying on a simple rereading to detect absences.
  • Handing empty tiles to an operator for hours on end.
  • Reading a disagreement on small objects as an error.
  • Treating fatigue with an instruction about vigilance.
  • Measuring performance without breaking it down by object size.
  • Neglecting the comparison of detection rates between operators.
  • Controlling by random sample despite a low density.
  • Ignoring how the rhythm evolves across a session.
  • Checking tiles in isolation without controlling the junctions.

What to take away

Geospatial quality is decided on what is missing rather than on what was produced.

Three readings emerge. An omitted object leaves absolutely no trace in the delivered data, which makes it undetectable by any control bearing on what has been produced and makes double annotation the only arrangement that reveals this defect. A tired annotator judges themselves just as attentive as at the start of their session, which forbids relying on self-regulation and requires acting on the organisation of the work rather than on instruction. And a gap of two pixels represents a considerable share of a small object, which produces high divergence rates without any operator having made an error and makes any disagreement measurement misleading unless it is broken down by size.

For a project’s economics, the article on the cost of satellite annotation details the lines. For regulated contexts, the article on annotation in defence contexts sets it out.

To explore delivery arrangements, supported formats and applicable control mechanisms, see our dedicated page on annotation for geospatial. And if you are preparing a geospatial project, let us discuss your need.

Tags

Découvrez nos articles