3D Cuboids – Annotating Objects for Autonomous Perception

The cuboid is the most used representation in 3D annotation and the most underestimated. Nine values look simple to fill in, and three of them concentrate most of the disagreement between annotators.This article sets out what that representation demands. It extends the guide to 3D annotation.

What a cuboid describes

Nine values compose this 3D annotation representation.They do not pose the same difficulty.Three centre coordinates, generally the easiest to establish.Three dimensions, length, width and height, at least one of which often rests on an extrapolation.And three angles, only one of which is frequently retained where objects rest on the ground.One important practical consequence follows. The last six values concentrate the difficulty, which makes a convention on dimensions and orientation more decisive than the care given to centring in 3D annotation.

Why this representation dominates

Four reasons explain the adoption of the cuboid in 3D annotation for autonomous perception.It directly supplies the quantities a driving system exploits, position, footprint and heading.It is fast to place, a box being adjusted in a few gestures where a segmentation demands point-by-point work.It permits automatic checks, an aberrant dimension being detectable with no human examination.And it compares easily between systems, volumetric overlap supplying an established metric.One observation follows. The first reason explains this representation’s persistence despite its limits, a planning system working on footprints rather than on exact shapes.

The orientation convention

This decision precedes any production in 3D annotation.Omitting it produces an unusable corpus.An object’s principal axis must be defined by a rule rather than by intuition.Four elements compose that convention.The definition of the front for each class, a vehicle and a pedestrian not calling for the same criterion.The direction of rotation and the angular origin, two incompatible conventions producing opposite values.The treatment of objects with no defined front, a post or a block having no natural orientation.And what to do where the front is not discernible in the data.One practical consequence follows for a 3D annotation project. The last element is the most often omitted, the absence of a rule leading each annotator to guess by intuition and producing a one hundred and eighty degree error on a share of the distant objects.

Dimensions and extrapolation

This difficulty is structural and is handled by convention.An object observed from one side does not supply its depth.Three approaches coexist in 3D annotation.Fitting the box to the observed points alone, which produces a dimension faithful to the measurement and variable with position.Extrapolating to a dimension typical of the class, which produces a stable size and a degree of convention.And fitting to the observed points while marking the extrapolated face with an attribute.One important observation follows. The third approach is the most informative and the least practised, it permits downstream use to distinguish a measurement from an estimate, a distinction neither of the first two preserves.

What the class contributes to the dimension

This property of 3D annotation has no equivalent in images.An object of a given class has bounded and known physical dimensions.Three direct exploitations follow.Extrapolation rests on a typical value rather than on a free estimate.The automatic check flags any dimension outside the expected range.And the annotator has a reference, a box visibly too short being corrected with no external reference.One practical consequence follows for a 3D annotation project. Those three exploitations presuppose a table of dimensional ranges per class, a one-page document that simultaneously improves throughput, consistency and control.

How an annotator places a box

A four-step method produces a consistent cuboid and it is quickly taught.Isolate the cluster of points corresponding to the object, aided by the projected colour where an image is available.Place the ground footprint from a top view, a gesture that simultaneously fixes the position, two dimensions and the orientation.Adjust the height from a side view, the third dimension reading better on a lateral projection.And verify in perspective view, a final check that reveals the points left outside the box.One important practical consequence follows. The second step carries most of the work in 3D annotation, the top view permitting four of the nine values to be fixed in one gesture, which explains the throughput gap between an annotator who uses it and one who stays in perspective view.

The objects that resist the cuboid

Four configurations make this 3D annotation representation unsuitable.Articulated objects, an articulated lorry cornering not being describable by a single box.Non-convex objects, a bicycle with its rider including much empty space.Deformable objects, a tarpaulin or a cable having no stable dimensions.And compact groups, a dense crowd not dividing into reliable individual boxes.One observation follows. The first configuration is handled by an explicit convention, describing the whole with one box or each segment with its own, a decision to be taken before production and not case by case.

The attributes accompanying a cuboid

Five attributes occur and their cost differs strongly.The class, filled once per object and of negligible cost.The occlusion level, which indicates whether the box rests on a partial observation.Confidence in the orientation, which distinguishes an observed value from a guessed one.The motion state, stationary or moving, which presupposes comparing several scenes.And the track identifier, where the scenes form a sequence.One observation follows for a 3D annotation project. The second and third attributes cost one click and transform a corpus, they permit a user to discard badly observed objects during training rather than learning from estimates presented as measurements.

The cuboid in sequence

An object observed across several scenes must keep its identity and a consistent dimension.Three consequences follow.Dimensions must stay constant along a track, a rigid object not changing size.The trajectory must be physically plausible, a position jump signalling an identity error.And orientation must vary continuously, an abrupt reversal signalling a badly settled front-back ambiguity.One important practical consequence follows. The first consequence supplies a very effective check, a dimension varying along a track being physically impossible and automatically detectable in 3D annotation.

The cuboid on the ground and the slope

One implicit assumption deserves setting out since it fails in certain contexts.Most projects retain a single rotation angle, assuming objects rest on a horizontal plane.Three situations invalidate that assumption.Sloping terrain, a rising road tilting every vehicle relative to the sensor frame.Camber, which tilts laterally and distorts the height measurement.And objects not resting on anything, a suspended load or an overhanging sign having no contact with the ground.One practical consequence follows for a 3D annotation project. Those three situations require deciding whether the two additional angles are annotated or ignored, a decision to be taken at scoping since a corpus produced with a single angle is not completed afterwards.

What distance changes

Four effects accompany increasing distance and they compound.The number of points falls, which reduces the information available.The class becomes uncertain, a silhouette of a few points potentially matching several categories.Orientation becomes undeterminable, the front no longer being discernible.And the dimension rests entirely on extrapolation.One important observation follows for a 3D annotation project. Those four effects justify a requirement differentiated by distance band rather than a single requirement, a near object and a distant one not being described with the same certainty.

The classes of a road nomenclature

Four nomenclature decisions recur in this vertical and they are settled at scoping.The level of detail on vehicles, a single class or a distinction between car, van, lorry and two-wheeler.The treatment of the cyclist, one box enclosing the person and their bicycle or two distinct objects.The distinction between moving objects and fixed elements, street furniture and vegetation often warranting separate treatment.And the residual class, indispensable for objects present and not anticipated.One practical consequence follows for a 3D annotation project. The second decision produces the most frequent disagreement, the whole formed by a person and their bicycle changing dimensions according to whether they ride or push it, which calls for a rule rather than a case-by-case judgement.

Quality control on a cuboid corpus

Six 3D annotation verifications exploit this representation’s properties.Dimensional plausibility per class, the most economical and most revealing check.The number of points contained, a nearly empty box signalling an annotation with no support.Position relative to the ground, an object floating or sunk signalling an error.Non-intersection between solid objects.Dimensional constancy along a track.And angular continuity, which detects orientation reversals.One practical consequence follows. Those six checks are programmed once and run on every batch, which makes them the best ratio between setup effort and defects detected.

Measuring disagreement on cuboids

One operation quantifies what can be promised and it is conducted on a few scenes.Two trained annotators handle the same scenes and their boxes are compared.Four usable pieces of information follow in 3D annotation.Agreement on the presence of objects, which measures reciprocal omissions by distance band.The centre position error, generally small and rarely problematic.The dimensional error, whose magnitude depends directly on the extrapolation convention adopted.And the orientation disagreement rate, which includes complete inversions.One important practical consequence follows. The fourth indicator must be reported separately from modest angular errors, a one hundred and eighty degree inversion and a five degree error arising from different causes and calling for different corrections in 3D annotation.

What this work costs

Four factors determine the load per object in 3D annotation.The point density on the object, a sparse cloud slowing the decision as much as a dense one eases the adjustment.The requirement on orientation, a tight tolerance demanding several viewpoints.The number of attributes requested, occlusion and state adding to the geometry.And temporal continuity where the scenes form a sequence.One observation follows. The first factor produces a counter-intuitive variation, a distant object costing more than a near one while carrying less information, which argues for a numeric distance threshold.

What the downstream system expects

Four uses exploit a cuboid and their requirements differ markedly.Obstacle avoidance, which needs a footprint and is satisfied by an approximate orientation.Trajectory prediction, which rests on heading and requires a precise orientation.Manoeuvre planning, which needs accurate rather than cautious dimensions.And conformity assessment, which requires knowing which values are measured and which are estimated.One observation follows for a 3D annotation project. Those four uses call for different tolerances on the same nine values, which makes a single requirement applied across a corpus either too loose for the second use or too expensive for the first.

What pre-annotation contributes here

Three observations situate the contribution of assistance to this 3D annotation task.Proposing boxes on common classes works well, those objects being abundant in public corpora.Fine adjustment of the faces stays human, a model producing a plausible box rather than a correct one.And orientation is the weak point, a model reproducing the front-back ambiguities rather than resolving them.One practical consequence follows for a 3D annotation project. Those three observations lead to a clear division, assistance proposes presence and class, the human validates orientation and adjusts the extrapolated faces.

The first batch of a cuboid project

Four scenes compose a pilot batch that genuinely tests the conventions.An ordinary traffic scene, which supplies the reference throughput and the mean object count.A scene containing objects at the distance limit adopted, which verifies that the threshold is well placed.A scene with an articulated object or a cyclist, which tests the most fragile nomenclature conventions.And a scene where several objects are partially hidden, which tests the extrapolation rule.One observation follows. Those four scenes are handled in a morning, they supply the real throughput and the orientation disagreement, and their absence explains most of the gaps between a costing and a production in 3D annotation.

The conventions to write first

Five rules cover most of the disagreement on this representation.The definition of the front for each class, expressed by an observable criterion.What to do where the front is not discernible, marking rather than supposition.The extrapolation rule for unobserved faces, with the attribute that signals it.The treatment of articulated objects and of person-plus-machine combinations.And the minimum number of points below which an object is not annotated.One practical consequence follows for a 3D annotation project. Those five rules fit on one page, they are written in an hour from the pilot batch, and they remove most of the rework a corpus otherwise undergoes during production.

Approaching a cuboid project

Five questions scope this type of 3D annotation project.What definition of the front for each class. That answer conditions the corpus’s consistency.Are unobserved faces extrapolated. That answer determines the comparability of dimensions.What maximum distance is adopted. That answer bounds the volume and the cost.Do the scenes form a sequence. A positive answer adds checks and requirements.And does a table of dimensional ranges exist. Its absence deprives the project of the most effective check.Those five answers determine the load and the consistency of the result. Asking them before starting avoids a corpus whose dimensions are not comparable with one another.

The question that frames a cuboid project

One question determines the tolerance to adopt and it concerns the use.Does the downstream system need to know where the object is going or only where it is.A presence-oriented answer, avoidance and counting, tolerates an approximate orientation and places the project in the most economical regime.A motion-oriented answer, trajectory prediction, requires a precise orientation and angular continuity across sequences.That question is asked in one sentence, it determines which quantity the requirement bears on, and it avoids applying to a whole corpus a tolerance only part of the uses justified.

What the corpus will be worth afterwards

Four properties determine whether a cuboid corpus stays usable beyond its initial project.Explicit declaration of the axis convention and the direction of rotation, without which no combination with another corpus is possible.The presence of the occlusion and confidence attributes, which permit filtering rather than reworking everything.Retention of the source clouds, an annotation separated from its data losing all value.And the table of dimensional ranges used, which documents the extrapolations practised.One observation follows. Those four properties cost little at the time of building and cannot be added afterwards, which places them in the same category as the conventions set out above.

Three checks before delivering

Three checks suffice to rule out most defects in a batch of cuboids.Run the table of dimensional ranges, which flags the boxes too small on the unobserved side.Count the points contained in each box, an empty box signalling an annotation with no support.And find the orientation reversals along the tracks, which betray a badly settled front-back ambiguity.Those three checks run on a delivered export, they need no external reference, and they cover what distinguishes a usable cuboid corpus from a visually acceptable one.

What this chapter teaches

One cross-cutting observation deserves closing this examination.The cuboid looks simple and concentrates the difficulty in three of its nine values.Three findings compose it.Orientation is the most fragile quantity, it has no equivalent in image annotation and its ambiguity produces one hundred and eighty degree errors invisible to the eye.Dimensions rest partly on an extrapolation, which makes them a convention disguised as a measurement.And dimensional ranges per class turn those two weaknesses into checks, physics bounding what an annotation can assert.That finding matches the one the preceding series established about structural constraints: where data carries a regularity independent of the annotated content, it serves verification as much as production.

The number to put in a contract

One clause recurs in cuboid projects and it is usually written wrongly.Clients ask for an accuracy figure on dimensions, and a single number cannot honestly be given.Three reasons make that so.The accuracy achievable depends on distance, a near object being measured and a far one estimated.It depends on the extrapolation convention, since a box fitted to observed points and one extended to a typical size are not measuring the same thing.And it is bounded by the disagreement between trained annotators, which is itself distance-dependent.One practical alternative exists for a 3D annotation provider. Commit to a figure per distance band, measured on the pilot batch and stated alongside the extrapolation convention, which is verifiable and which a single global number is not.

Common mistakes

These failures recur often enough that naming them is usually enough to avoid them.
  • Not defining the front of each class before production.
  • Omitting the rule applying where orientation is not discernible.
  • Extrapolating unobserved faces without marking it with an attribute.
  • Using two incompatible angular conventions in one corpus.
  • Describing an articulated object with a single box and no explicit convention.
  • Neglecting the table of dimensional ranges per class.
  • Requiring the same orientation precision at every distance.
  • Omitting the dimensional constancy check along a track.
  • Accepting a model proposal without verifying the orientation.
  • Annotating beyond a distance where the class is no longer determinable.

What to take away

The cuboid describes an object by nine values of which six concentrate the difficulty, the dimensions and the orientation.Three readings emerge. The rule applying where the front is not discernible is the most often omitted convention, its absence producing one hundred and eighty degree errors on a share of the distant objects. Marking extrapolated faces with an attribute permits downstream use to distinguish a measurement from an estimate, a distinction no other approach preserves. And the table of dimensional ranges per class fits on one page and simultaneously improves throughput, consistency and control.For the formulation that describes shape rather than footprint, the article on point cloud segmentation sets out its conditions. For the general frame, the guide to 3D annotation lays out the panorama.To explore delivery arrangements, supported formats and applicable control mechanisms, see our dedicated page on 2D and 3D annotation. And if you need cuboids annotated on point clouds, let us discuss your project.
Tags

Découvrez nos articles