Preparing a Satellite Dataset – Resolution, Georeferencing, Tiling

An image archive is not a dataset. Between the two sits a series of technical decisions which, taken lightly, produce a corpus annotated with care and unusable for the very purpose it was meant to serve.

This article sets out those decisions. It extends the article on airborne LiDAR.

What resolution commands

Five consequences follow from the ground pixel size in satellite imagery.

It fixes the minimum size of a genuinely identifiable object.

It decides which annotation geometries stay usable.

It finally conditions the separation of two neighbouring objects.

It bears directly on the number of objects per tile.

And it directly determines the cost of acquiring the data.

One important practical consequence follows. A simple rule guides the scoping, an object having to cover several pixels in its smallest dimension to be delimitable, which permits checking a project’s feasibility before any acquisition.

What resampling changes

Four satellite imagery effects accompany a change of resolution.

An image deliberately degraded loses the removed detail for good.

An upsampled image gains strictly no information at all.

Mixing several resolutions produces a heterogeneous corpus.

And the interpolation method adopted modifies the appearance of outlines.

One important observation follows for a satellite imagery project. The second effect regularly misleads, a merely enlarged image looking finer while containing nothing more, which sometimes leads to annotating objects the sensor had never resolved.

What georeferencing requires

Five verifications precede any production in satellite imagery.

The coordinate system employed along with its projection.

The correction of relief, done by a prior orthorectification.

The effective superimposition of two successive acquisitions.

The consistency between the image and the vector layers already existing.

And positional accuracy set against known control points.

One practical consequence follows. The fourth verification is often neglected, a gap between the image and the client’s own reference making the delivered annotations unusable without a re-registration nobody had thought to budget for.

What tiling imposes

Five decisions structure the cutting up of a satellite imagery scene.

The tile size, which arbitrates between visible context and manageability.

The overlap finally adopted between two adjacent tiles.

The treatment given to objects cut by a boundary.

The counting rule applied to those shared objects.

And the reassembly of the result at the whole scene’s scale.

One observation follows. The first decision is judged on the object rather than on convenience, a tile that is too small depriving the annotator of the context they need to recognise what they are looking at.

What overlap brings

Four satellite imagery benefits follow from a margin between tiles.

A cut object appears in full on the neighbouring tile.

Reassembly is done with no visible discontinuity.

Edge errors are then detected by simple comparison.

And the context remains available right to the boundaries.

One important practical consequence follows. Those four benefits are paid for in annotated volume, too large an overlap multiplying the number of objects seen twice, which requires an explicit trade-off between junction quality and production cost.

What a heterogeneous archive imposes

Five gaps occur within an assembled satellite imagery archive.

Different resolutions depending on each of the sources.

Sensors with markedly distinct colour renderings.

Acquisition dates spread across several years.

Viewing angles that vary greatly from one image to another.

And very uneven correction states from one image to the next.

One important observation follows for a satellite imagery project. The last gap is the most insidious, an archive mixing corrected images and raw images producing appearance variations the annotator will attribute to the terrain rather than to the processing.

What the data format imposes

Five technical characteristics condition the work in satellite imagery.

Radiometric depth, very often greater than that of a photograph.

The number of bands along with their order in the file.

The compression employed along with its possible losses.

The georeferencing metadata actually attached to the file.

And the file’s internal structure, which permits partial access or not.

One important practical consequence follows. The first characteristic surprises teams coming from image annotation, a satellite file coding each band across far more levels than an ordinary screen displays, which requires a visualisation choice that is anything but automatic.

What separating the sets requires

Four rules govern the splitting of a satellite imagery corpus.

The separation must always be geographic rather than random.

Neighbouring tiles necessarily belong to the same set.

The overlap must never cross two distinct sets.

And the same areas taken at different dates stay together.

One important practical consequence follows. Those four rules target a single defect, a training tile neighbouring an evaluation tile producing a measured performance that largely overstates the system’s real capacity.

What selecting the areas decides

Five criteria guide the choice of satellite imagery areas to annotate.

The representativeness of the territories actually targeted by the use.

The presence of the rare classes the project seeks to obtain.

The variety of acquisition conditions genuinely encountered.

The availability of a ground truth over each of those areas.

And the contiguity needed for controlling junctions between tiles.

One important observation follows for a satellite imagery project. The first criterion often conflicts with the fourth, the best documented areas being very rarely the most representative, which requires choosing explicitly between ease of verification and validity of the corpus.

What the archive documentation must record

Six pieces of information accompany each image of a satellite imagery corpus.

The source along with the original sensor.

The date along with the exact hour of acquisition.

The native resolution observed before any processing.

The atmospheric as well as geometric correction state.

The coordinate system finally employed.

And every processing step applied since acquisition.

One important practical consequence follows for a satellite imagery project. The last piece is almost always missing and reconstitutes badly, a corpus whose processing chain is unknown becoming impossible to extend coherently, which limits its lifespan far more surely than its volume does.

What the annotator actually receives

Five elements compose a properly prepared satellite imagery annotation station.

Tiles whose size supplies the necessary context.

A visualisation fixed once and for all and identical for every operator.

Access to the neighbouring tiles for handling edge objects.

The superimposable reference layers, where these already exist.

And a written instruction bearing on the treatment of boundaries.

One important observation follows for a satellite imagery project. The third element is often neglected and costs little, an operator deprived of the neighbouring tile being unable to decide whether a cut object continues or stops, which turns a question of preparation into a permanent source of disagreement.

What this step’s cost covers

Four lines compose the preparation of a satellite imagery archive.

Acquiring or else collecting the source images.

The correction as well as harmonisation processing.

The tiling, the overlap as well as the splitting of the sets.

And the documentation of the whole processing chain.

One important practical consequence follows. Those four lines all precede the first annotation and are very regularly forgotten at costing, a project budgeted on annotation volume alone discovering along the way a preparatory phase that sometimes weighs as much as production itself.

What this step has that is particular

Four traits separate preparation from a satellite imagery project’s other phases.

It precedes everything else and is judged only much later.

Its defects are never visible on one single isolated tile.

It belongs to technique far more than to the field handled.

And its cost is very regularly omitted from initial costings.

One important observation follows for a satellite imagery project. The second trait explains why these decisions are taken badly, a control run by sample of tiles being unable to reveal either a reference gap or a faulty set separation, which lets these defects run through the whole production undetected.

What the first batch must establish

Four results justify a trial batch before full production.

The effective superimposition of the annotations with the client’s layers.

The tile size annotators judge genuinely sufficient in context.

The number of objects met per tile across the areas actually handled.

And the proportion of objects effectively cut by a tile boundary.

One important practical consequence follows for a satellite imagery project. The first result is verified within an hour and avoids the costliest dispute, a simple transfer of a few annotations into the client’s own system showing immediately whether the two line up, a check almost nobody runs before having produced everything.

What this step makes possible afterwards

Four satellite imagery freedoms follow from a well conducted preparation.

The corpus can then be extended without losing any of its consistency.

A new use can directly reuse the same tiles.

A measured performance finally reflects the system’s real capacity.

And the annotations transfer into the client’s tools with no rework.

One important practical consequence follows. The second freedom holds the greatest economic value, a genuinely reusable corpus amortising its cost across several successive projects, which justifies the extra preparation effort far better than any argument about quality.

Three errors proper to this step

Three satellite imagery defects are born before the first annotation.

An insufficient resolution that upsampling has come to make misleading.

A reference gap discovered only at the moment of delivery.

And a random splitting of the sets that comes to flatter the measured performance.

Those three defects are not corrected by reworking the annotations, they require redoing the preparation and then the production, and their common point is to cost almost nothing to avoid and an entire project to repair.

What the client must supply here

Four elements come from the client rather than from the provider.

The coordinate system employed within their own tools.

The reference layers the annotations will have to superimpose on.

The provenance as well as the correction state of the images they supply.

And the later uses they may envisage for the corpus.

One important observation follows for a satellite imagery project. The third element is rarely asked for and often missing, a client having themselves received their images from a third party sometimes not knowing what processing those underwent, which requires documenting what can be documented and explicitly flagging what stays unknown.

What reusing an existing corpus imposes

Five verifications precede the extension of an already assembled corpus.

The resolution of the new images, compared with that of the old images.

The coordinate system, either unchanged or converted with no loss at all.

The tiling convention along with the overlap employed originally.

The class reference along with its possible evolutions.

And the existing splitting of the sets, to be respected rather than redone.

One important practical consequence follows for a satellite imagery project. The last verification protects the initial corpus’s value, an addition split without regard for the already existing separation mixing training areas and evaluation areas, which invalidates any comparison with the performances measured beforehand.

What the volume to handle changes here

Four effects accompany a large satellite imagery archive.

Tiling becomes a wholly automated operation rather than a manual one.

The georeferencing control bears on a sample rather than on the whole.

Storage as well as transfer become lines in their own right.

And rework following a preparation defect costs several days.

One important practical consequence follows. The last effect justifies controlling early rather than broadly, a verification conducted on a few tiles before the full tiling avoiding rework whose cost keeps growing with the volume already produced.

What this step owes to the rest of the course

Four technical decisions recur across the preceding chapters.

Resolution, which already bounded the detection of solar installations.

Tiling, whose junctions already threatened a network’s continuity.

Set separation, which distorted the measurement from the pillar guide onwards.

And harmonising the sources, without which two dates do not compare.

One important observation follows for a satellite imagery project. Those four decisions appeared until now as difficulties proper to each field, whereas they all belong to the same preparatory step, which suggests handling it once and for all rather than rediscovering it project after project.

What the provider brings here

Four contributions distinguish a preparation seriously conducted in satellite imagery.

A superimposition control conducted before production rather than at acceptance.

A tiling dimensioned for the target object rather than for the tool’s convenience.

A spatial separation of the sets applied from the very first tiling.

And full documentation of the processing chain handed over with the corpus.

One important practical consequence follows. The first contribution attracts little notice and avoids this field’s most frequent dispute, a client discovering at the moment of delivery that the annotations do not superimpose on their data rejecting the whole of the work for a cause that has nothing to do with its quality.

Approaching the preparation of a dataset

Five questions scope such a preparation in satellite imagery.

Does the resolution suffice for the target objects. Check before acquiring.

Do the images superimpose on the client’s reference. Otherwise the deliverable is unusable.

What tile size supplies the necessary context. The object decides it.

Is the archive homogeneous in correction. Otherwise the annotator will be misled.

And is the set separation spatial. Otherwise the measurement misleads.

Those five answers determine the very validity of the corpus. Asking them before production avoids work that is impeccable and yet unusable.

The question that frames the preparation

One question determines the whole technical chain in satellite imagery.

Will the corpus serve one single project or several.

One project permits a tiling cut for the target object, a visualisation chosen for it and summary documentation, the corpus disappearing with the model it served.

Several projects require the native resolution preserved, a documented processing chain and a tiling other uses will be able to take up.

That question belongs to the client’s ambition and not to technique, it is asked before acquisition, and it separates a disposable corpus from an asset reused for years.

Three decisions before producing

Three decisions commit the preparation of a satellite imagery dataset.

Verifying that the native resolution permits resolving the target objects.

Checking the superimposition of the images with the client’s reference.

And fixing a geographic separation of the sets before any tiling.

Those three decisions cost only a few hours, they precede the very first tile, and their absence produces annotation work that is impeccable and yet unusable.

Three checks before the first tile

Three checks qualify a prepared satellite imagery archive.

The native resolution, set against the real size of the target objects.

The effective superimposition obtained with the client’s reference layers.

And the set separation rule adopted, spatial rather than random.

Those three checks are each asked in one question, they require no particular tool, and their absence indicates a corpus whose defect will reveal itself after production rather than before it.

What this chapter teaches

One cross-cutting observation deserves closing this examination.

Preparation decides what the annotation will be able to produce.

Three findings compose it.

An enlarged image looks finer while containing nothing more.

A gap with the client’s reference makes the annotations unusable.

And a training tile neighbouring an evaluation tile distorts every measurement.

That finding casts light back on the preceding chapters, several difficulties there having their origin in a technical decision taken before the first annotation.

What this chapter leaves to the next

A well prepared corpus does not stop an annotator missing what they are looking at.

Three questions stay open in satellite imagery.

How to handle objects approaching the limit of the detectable.

What hours of scanning do to an operator’s attention.

And how to control work whose errors are absences.

Those three questions belong to annotation quality, which constitutes the subject of the following chapter.

Why offering preparation separately pays

One commercial point belongs at the close of this chapter.

Preparation is worth quoting as its own line rather than folding into the rate.

Three reasons follow in satellite imagery.

Clients rarely realise this phase exists until it is named.

Folded into a per-tile price it looks like an inflated rate.

And a provider who skips it wins on price and fails on delivery.

One important practical consequence follows for a satellite imagery provider. Quoting it openly turns an apparent price disadvantage into evidence of experience, since a client comparing two quotes where only one mentions georeferencing checks learns something useful about both, and the conversation moves from cost per tile to whether the work will be usable at all.

Common mistakes

These failures recur often enough that naming them is usually enough to avoid them.

  • Annotating objects the resolution did not permit resolving.
  • Upsampling an image in the belief of gaining detail.
  • Discovering a gap with the client’s reference at delivery.
  • Choosing a tile size for convenience rather than for the object.
  • Mixing corrected and uncorrected images within one archive.
  • Splitting the sets randomly rather than by geographic area.
  • Letting an overlap cross two distinct sets.
  • Omitting the treatment rule for objects cut by a tile.
  • Neglecting the positional check against known control points.
  • Assembling an archive without documenting each image’s provenance.

What to take away

Technical preparation decides what the annotation will be able to produce.

Three readings emerge. An upsampled image looks finer while containing no more information, which sometimes leads to annotating objects the sensor had never resolved and to producing a corpus part of which describes noise. A gap between the image and the client’s reference makes the delivered annotations unusable without a re-registration nobody had budgeted for, which makes this verification a prerequisite rather than an acceptance check. And a training tile neighbouring an evaluation tile produces a measured performance that largely overstates the system’s real capacity, which requires a geographic separation of the sets only georeferencing makes possible.

For the quality of the work, the article on small objects and visual fatigue details the approach. For a project’s economics, the article on the cost of satellite annotation sets it out.

To explore delivery arrangements, supported formats and applicable control mechanisms, see our dedicated page on annotation for geospatial. And if you are preparing a satellite dataset, let us discuss your need.

Tags

Découvrez nos articles