Sensor Fusion – Annotating LiDAR and Cameras Together

Combining a range sensor with a camera looks like adding two sources together. That combination in fact introduces a third thing, a geometric relation between them, whose correctness conditions the validity of everything that follows.This article sets out what this configuration demands. It extends the article on choosing between LiDAR and photogrammetry.

What fusion brings

Four benefits justify this 3D annotation configuration.Classification becomes more reliable, an object ambiguous in geometry often being obvious in an image.Distant objects stay identifiable, the image keeping a resolution the cloud has lost.Materials become distinguishable, two surfaces of the same shape separating by their appearance.And verification is easier, a checker finding in the image what they cannot distinguish among the points.One observation follows. The second benefit is the most decisive outdoors, it pushes back the distance beyond which an object becomes unclassifiable and it therefore widens the usable share of every scene.

What calibration establishes

A geometric relation links the two sensors.It must be known before any 3D annotation.Calibration determines where a point measured in space projects into the image.Three elements compose that calibration.The camera’s internal parameters, focal length, optical centre and distortion.The relative position and orientation of the two sensors.And the associated uncertainty, a value rarely supplied and nonetheless necessary.One important practical consequence follows for a 3D annotation project. The third element is the most often absent, a calibration delivered with no margin of error suggesting an exact correspondence physics does not permit.

How a calibration degrades

Four causes degrade a calibration correctly established at the outset.Vibration from the carrier, which progressively displaces the sensors relative to one another.Thermal variation, which expands the mountings differently.Impacts, a single knock sufficing to invalidate a setting.And maintenance work, a sensor removed and refitted not returning to exactly its position.One important observation follows. Those four causes make a calibration perishable rather than definitive, which requires dating each acquisition and attaching it to a given calibration in 3D annotation.

Temporal synchronisation

This second requirement accompanies the first.It is more often neglected in 3D annotation.Both sensors must observe the scene at the same instant.Three situations make this requirement critical in 3D annotation.A moving carrier, a few milliseconds of offset displacing the whole scene.Moving objects, whose position differs between the two observations.And the progressive sweep of a range sensor, whose points from one acquisition do not date from the same instant.One practical consequence follows for a 3D annotation project. The third situation is the least known, a cloud acquired by sweeping on a fast carrier presenting a deformation only trajectory compensation corrects.

What a bad calibration produces

Four symptoms betray a misalignment.They are recognisable visually.An object’s colour spills onto its neighbour, which betrays a lateral offset.The sky colours the tops of objects, the symptom of a vertical offset.The error grows with distance, which indicates an angular error rather than a translation.And the offset varies with position in the image, which betrays badly corrected distortion.One important observation follows. Those four symptoms are distinguishable from one another, which permits a trained annotator to qualify the defect rather than endure it, and to report it precisely enough for it to be corrected.

Verifying a calibration received

Four trials qualify an alignment before committing to a 3D annotation production.Project the cloud onto the image and examine object outlines at short range, where angular error shows little.Repeat the examination at long range, where the same angular error produces a large offset.Find a scene containing a fast-moving object, revealing a synchronisation defect.And compare several acquisitions spaced in time, whose divergence signals a drift.One important practical consequence follows for a 3D annotation project. Those four trials take an hour on a sample, they precede any production, and they avoid annotating a whole corpus with a false projection nobody will notice before delivery.

What fusion changes in the annotation

Four effects appear on the 3D annotation work itself.Throughput rises on ambiguous classes, the image lifting a doubt geometry left open.The gaze moves between two representations, which demands suitable tooling and a habit.A contradiction becomes possible, geometry and image not suggesting the same class.And confidence in the projection becomes a precondition, an annotator not being able to rely on a correspondence they know to be false.One important practical consequence follows for a 3D annotation project. The third situation calls for a written priority rule, failing which each annotator decides according to their preference and the corpus loses its consistency.

The priority rule between sources

This convention is indispensable in 3D annotation.It is almost always absent.Three formulations coexist in 3D annotation.Geometry takes precedence, the image serving only as a class cue, which suits measurement tasks.The image takes precedence for the class and geometry for the position, a division exploiting the best of each source.And the contradiction is flagged by an attribute rather than resolved, which preserves the information for downstream use.One observation follows. The second formulation suits most projects and the third is the most informative, a flagged contradiction locating either a calibration defect or a genuinely ambiguous configuration in 3D annotation.

The multi-camera configurations

A frequent extension complicates fusion and it deserves setting out.Several cameras surround the range sensor to cover its full field.Four difficulties follow in 3D annotation.Each camera demands its own calibration, which multiplies the relations to establish and maintain.The overlap zones receive two possible colours, which calls for a choice rule.Exposure settings differ between cameras, which produces a colour discontinuity on a single object.And an object straddling two fields receives a composite appearance that can mislead an annotator.One important observation follows for a 3D annotation project. The third difficulty is the most visible and the most disorienting, a vehicle crossing the boundary between two fields changing hue without changing nature, which a convention must flag explicitly.

What fusion does not resolve

Four limits persist despite the combination.A zone hidden from both sensors stays without information.The fields of view differ, part of the cloud having no corresponding image.Occlusion differs between the two viewpoints, a point visible to the range sensor potentially being hidden from the camera.And the exposure gap makes some zones unusable in the image while the geometry there is correct.One practical consequence follows. The third limit produces a characteristic artefact, a background point receiving the colour of the object hiding it in the image, a systematic error a convention must provide for.

What the tooling must permit

Five functions condition the work on a fused scene.Simultaneous display of the cloud and the image, with a visual correspondence between the two views.Cross-designation, a point selected in one view lighting up in the other.Switching between coloured cloud and raw cloud, which permits verifying what colour masks.Display of the projection’s confidence level where available.And the ability to annotate from either view according to which is more convenient.One practical consequence follows for a 3D annotation project. The third function is the most decisive for quality, a coloured cloud appearing complete where the underlying geometry is sparse, an illusion a switch lifts immediately.

Quality control on a fusion

Five verifications are added to those of a single-sensor 3D annotation corpus.Outline consistency, the coloured boundary having to coincide with the geometric discontinuity.Alignment stability across the campaign, a drift signalling a degrading calibration.Image coverage of the cloud, the share of points actually holding a colour.The rate of flagged contradictions, whose rise betrays an alignment problem.And verification on moving objects, revealing a synchronisation defect invisible on a static scene.One important observation follows for a 3D annotation project. The last verification is the most effective and the least practised, an object in rapid motion revealing in one frame a temporal offset no fixed scene would show.

The metadata of a fused scene

Six pieces of information accompany a multi-sensor delivery and their absence limits exploitation.The identifier of the calibration used, with its date of establishment.The internal and external parameters themselves, without which the projection does not reproduce.The associated uncertainty, which bounds the confidence given to the correspondence.The timestamp of each source, which permits verifying synchronisation.The share of the cloud covered by an image, which bounds the real contribution.And the priority rule applied in case of contradiction.One observation follows for a 3D annotation project. The first metadata is the one whose absence is paid latest, a corpus mixing acquisitions calibrated differently presenting an uneven projection quality nothing signals.

Radar as a third source

A complementary sensor appears in some configurations and its contribution differs from the other two.Radar measures a distance and a radial velocity with very few points.Three contributions characterise that sensor.Velocity measured directly, information neither of the other two sources supplies without comparing two instants.Robustness to degraded conditions, rain and fog affecting it little.And range, greater than that of many optical sensors.Three limits accompany it.Spatial resolution is very low, which precludes any delimitation.Multiple reflections produce ghost detections.And height is often badly resolved, which prevents distinguishing an obstacle on the ground from a gantry above.One practical consequence follows for a 3D annotation project. Those limits make radar a source of attributes rather than of geometry, its contribution consisting in enriching an object already delimited by the other sensors rather than in delimiting it.

What fusion contributes to control

Three uses of the second sensor belong to control rather than to production in 3D annotation.Rapid review of a batch by displaying the projections, a class defect being immediately visible in colour.Detection of omitted objects, a coherent coloured zone with no annotation drawing attention.And arbitration of disagreements between annotators, the image settling a share of the contested cases.One observation follows for a 3D annotation project. The first use transforms the control of a 3D corpus, a rapid visual review becoming possible where examining a raw cloud demands slow and tiring navigation.

What fusion costs

Four lines are added to a single-sensor 3D annotation project.Establishing and verifying the calibration, a one-off cost per hardware configuration.The projection processing, which associates a colour with each point.Handling the contradiction cases, which slows a share of the decisions.And the additional data volume, images and parameters accompanying the clouds.One practical consequence follows. Those four lines are offset by the throughput gain on ambiguous classes, which makes fusion economically neutral in the short term and favourable as soon as the corpus extends.

What fusion changes in the costing

Four effects modify an estimate and two of them run against intuition.Throughput rises on ambiguous classes, which reduces the time per object in complex scenes.The usable share of each scene widens, which increases the number of annotatable objects for the same acquisition duration.Control becomes faster, visual review by projection replacing slow navigation in the cloud.And the data volume grows, which weighs on transfer and storage rather than on annotation.One practical consequence follows for a 3D annotation project. The second effect is the most often forgotten in a costing, one acquisition producing more annotatable objects where an image accompanies the cloud, which raises the invoice while improving the corpus.

The first batch of a fused project

Four scenes compose a pilot batch testing the configuration as much as the conventions.A static short-range scene, which verifies alignment under the most favourable conditions.A scene containing distant objects, where angular error manifests.A scene with a fast-moving object, which reveals a synchronisation defect.And a scene in difficult lighting, backlight or low light, which measures the image’s real contribution in the conditions where it is least reliable.One observation follows. The fourth scene is the one most often missing, a pilot batch run in good conditions systematically overestimating the camera’s contribution and producing an announced throughput production will not match in 3D annotation.

Approaching a multi-sensor project

Five questions scope a fused configuration in 3D annotation.Is the calibration supplied and documented. Its absence makes the fusion unusable.Is its uncertainty quantified. That value bounds the confidence given to the projection.Are the carrier or the scene moving. A positive answer makes synchronisation critical.What share of the cloud has an image. That answer bounds the real contribution of fusion.And what priority rule applies in case of contradiction. Its absence produces inconsistency between annotators.Those five answers determine the feasibility and the real contribution. Asking them before starting avoids a corpus whose colour misleads rather than informs.

The question that frames the fusion

One question determines whether this configuration actually contributes anything.What proportion of the class decisions does geometry alone leave undecided.A low answer, a closed catalogue or geometrically distinct objects, makes fusion of little value against the constraints it imposes.A high answer, an urban scene or a nomenclature distinguishing materials, makes it decisive.That question is measured on a pilot batch rather than guessed, and it avoids imposing a calibration requirement on a project geometry alone would have served.

Three checks before producing

Three checks suffice to qualify a fused configuration.Project the cloud onto the image and examine outlines at short then at long range, which separates a translation error from an angular one.Find a scene containing a fast object, the only revealer of a synchronisation defect.And measure the share of the cloud actually covered by an image, a value bounding the configuration’s real contribution.Those three checks take an hour on a sample, they precede any production, and they avoid annotating a whole corpus with a projection nobody has verified.

What fusion commits over time

Three consequences extend beyond the initial project.A fused corpus depends on the hardware configuration that produced it, a change of sensor or of mounting making the new acquisitions different from the old.A system trained on fused data presupposes the same fusion in production, which ties the deployment to a particular equipment set.And retaining the calibration parameters becomes indispensable, without them the projection no longer reproduces and the corpus loses its second channel.One observation follows for a 3D annotation project. The second consequence is the most structuring and the least anticipated, a client not being able to deploy on simpler hardware a system trained on a rich configuration.

Three decisions before the first acquisition

Three decisions condition the usability of a fused corpus.Establishing and documenting the calibration with its uncertainty, rather than assuming it supplied by the manufacturer.Fixing the priority rule between sources in case of contradiction, with the attribute that flags it.And planning periodic re-verification of the alignment, whose frequency depends on the carrier and the conditions.Those three decisions cost one meeting and an hour of verification, they precede the first scene, and their absence produces a corpus whose colour misleads with nothing signalling it.

What this chapter teaches

One cross-cutting observation deserves closing this examination.Fusion adds information and a dependency.Three findings compose it.The contribution is real and it bears principally on the class decision, geometry staying the source of position.The relation between sensors is data in its own right, perishable and rarely documented with its uncertainty.And a contradiction between sources is information rather than a defect, provided a convention plans to flag it.That finding matches the one the preceding series established about combined sources: what links them deserves as much attention as what they contain, and that relation is what degrades first.

What to ask a client about their rig

Four questions about the hardware reveal more than any discussion of annotation method.They are asked before quoting.When was the calibration last established, and by whom. A vague answer usually means it dates from installation and has never been rechecked.Is the mounting rigid, and has the rig been transported or serviced since. Both are the commonest causes of the drift described above.How are the two sources timestamped, and against what clock. A shared clock and two independent ones are not the same arrangement.And has anyone looked at a projection on a scene with a moving object. The answer is usually no, and it is the check that matters most.One practical consequence follows for a 3D annotation provider. Those four questions cost nothing to ask, they frequently surface a problem the client did not know they had, and raising it before production is a service rather than an obstacle.

Common mistakes

These failures recur often enough that naming them is usually enough to avoid them.
  • Treating a calibration as definitive rather than perishable.
  • Accepting a calibration without its associated uncertainty.
  • Neglecting synchronisation on a moving carrier.
  • Not writing the priority rule for contradictions.
  • Assuming the whole cloud has a corresponding image.
  • Ignoring the occlusion difference between the two viewpoints.
  • Assigning a foreground colour to a background point.
  • Checking alignment only on static scenes.
  • Omitting to date acquisitions and attach them to a calibration.
  • Relying on colour where the camera’s field does not reach.

What to take away

Fusion adds a geometric relation between sensors, and that relation conditions the validity of everything the combination permits.Three readings emerge. A calibration is perishable, vibration, thermal variation and maintenance progressively degrading it, which requires dating acquisitions and attaching them to a given calibration. Verification on moving objects is the most effective and the least practised, a fast object revealing in one frame a temporal offset no fixed scene would show. And a priority rule between sources must be written, failing which each annotator settles contradictions according to their preference.For the tooling that makes this configuration usable, the article on 3D annotation tools details the options. For the general frame, the guide to 3D annotation lays out the panorama.To explore delivery arrangements, supported formats and applicable control mechanisms, see our dedicated page on 2D and 3D annotation. And if you are preparing a multi-sensor corpus, let us discuss your project.
Tags

Découvrez nos articles