The third dimension gives quality control something image annotation does not: physical quantities. A box too short, a floating object or two solids interpenetrating are measurable impossibilities rather than judgements.
This article sets out that arrangement. It extends the article on the cost of 3D annotation.
What metric quantities bring to control
Four properties distinguish controlling a 3D annotation corpus from controlling an image corpus. Dimensions are physical, which permits comparing them to known ranges. Gravity applies, which makes an implausible position detectable. The impenetrability of solids supplies a constraint nothing legitimately violates. And the number of points contained measures an annotation’s real support. One important practical consequence follows. Those four properties make automatic control markedly more powerful in 3D annotation than in image annotation, where no physical constraint bounds what a box can assert.The defects specific to 3D
Six failure modes do not exist in image annotation. The undersized box, fitted to the visible points alone of a partially observed object. Orientation inversion, a one hundred and eighty degree error invisible to the eye. The empty box, placed on a zone with no supporting points. The floating or sunken object, whose vertical position is inconsistent with the ground. Intersection between solids, two objects occupying the same volume. And propagation spillover in segmentation, one class invading its neighbour. One important observation follows. Those six defects are all automatically detectable, which distinguishes this field from video where the characteristic defects required examination across a sequence.The automatic checks
Seven verifications exploit a cloud’s physical quantities. They run with no intervention. Dimensional plausibility per class, a comparison against a known range. The number of points contained, a low threshold signalling an annotation with no support. Distance to the ground, which detects floating and sunken objects. Non-intersection between solid objects, which nothing legitimately violates. Orientation consistency with the trajectory where scenes form a sequence. Dimensional constancy along a track. And detection of isolated islands in segmentation. One practical consequence follows for a 3D annotation project. Those seven checks are programmed once and apply to every batch, which makes them the best ratio between setup effort and defects detected in the whole arrangement.What the automatic checks do not see
Four 3D annotation defects escape any programmed verification. Omission, an object never annotated producing no anomaly. Systematic class error, a consistent confusion between two categories passing every formal check. Erroneous orientation without inversion, an error of a few degrees staying plausible. And misapplication of a convention, an annotator faithfully applying a rule they misunderstood. One important observation follows. Those four defects require human control, which makes automatic and human two complementary arrangements rather than substitutable ones.Detecting omissions
This defect is the gravest and the only one no automatic check reveals directly. Three methods permit detecting it in 3D annotation. Exhaustive re-annotation of a short scene, a reliable and expensive method whose result bounds the omission rate. Comparison with an existing model, whose unannotated detections flag candidates. And the search for unattributed point clusters, a coherent group of points outside any annotation signalling a forgotten object. One important practical consequence follows. The third method is specific to this field and automatable, geometry permitting a structure with no label to be spotted where in image annotation nothing distinguishes an empty zone from a forgotten one.The comparison metrics
Four indicators measure the gap between two annotations. They account for different things. Volumetric overlap, the three-dimensional equivalent of box overlap. The centre position error, generally small and rarely problematic. The dimensional error per axis, whose magnitude depends on the extrapolation convention. And the angular error, reported separately from complete inversions. One observation follows for a 3D annotation project. The first indicator aggregates the other three and masks their origin, which makes it convenient for comparing systems and insufficient for diagnosing a corpus.How to sample
Four principles govern control sampling in 3D annotation. Target the distance bands rather than drawing uniformly, difficulty varying strongly with range. Include the zones of high object density, where intersections and confusions concentrate. Retain the scenes containing problematic materials, glazing and reflective surfaces. And keep a constant sample between batches, the only way to measure drift. One practical consequence follows. The first principle is the most effective, a uniform sample being dominated by near and easy objects while most of the disagreement sits at longer ranges.Control by visual review
A human method complements the programmed checks and it is underused. The principle is to traverse an annotated scene from a top view rather than in perspective. Four 3D annotation defects become immediately visible there. Badly oriented boxes, whose divergence from the lane axis leaps out. Omitted objects, a cluster of points with no enclosure being spotted instantly. Overlaps, two footprints superimposed on the ground. And inconsistent alignments, a line of vehicles one of which departs from the common trajectory. One important observation follows for a 3D annotation project. This view exploits a regularity perspective masks, the objects of a road scene aligning on common directions, which makes an anomaly visible with no individual examination.Measuring disagreement
This operation bounds what can be promised in 3D annotation. It is conducted on a few scenes. Two trained annotators handle the same scenes and the results are compared. Four usable pieces of information follow. Agreement on the presence of objects, by distance band. The dimensional error, whose magnitude reveals the clarity of the extrapolation convention. The orientation inversion rate, reported separately. And the list of contested configurations, the raw material of the reference. One important practical consequence follows. Those four values must be reported by distance band, a global figure masking that performance is excellent at short range and mediocre beyond.The indicators to track per batch
Five values are read automatically. Their evolution reveals drift. The mean number of objects per scene, whose fall signals omissions. The distribution of dimensions per class, whose shift signals a drifting convention. The rate of automatic alerts raised, a direct indicator of upstream quality. The rate of objects marked uncertain, to be interpreted in both directions. And the rejection rate at validation, which measures upstream quality rather than final quality. One observation follows for a 3D annotation project. The second indicator is specific to this field and highly revealing, a mean dimension decreasing across batches betraying annotators fitting increasingly to the visible points alone.Establishing the dimensional range table
One operation conditions the most effective check and it takes half a day. Four sources feed that table. Manufacturer documentation where the objects are identified products. Measurement on real objects from the corpus, a safer method since it incorporates the extrapolation convention adopted. Public corpora from the same domain, which supply an order of magnitude to verify. And observation of the extreme values in the first batch, which reveals the cases the first three sources did not anticipate. One important practical consequence follows for a 3D annotation project. The second source must take precedence over the first, a range established on manufacturer documentation flagging as aberrant annotations that conform to the project’s convention, which produces an unusable flood of alerts.The correction cycle
Four stages organise the improvement of a 3D annotation corpus, and they repeat between batches. Running the automatic checks on the delivered batch. Visual review of a sample stratified by distance. Classing each defect as an individual error or a convention defect. And writing the missing rule before the next batch. One important practical consequence follows for a 3D annotation project. The third stage distinguishes a correction from an improvement, and its distribution constitutes a maturity indicator, a high share of convention defects signalling an incomplete reference rather than an insufficient team.Framing a tenable requirement
A three-step method avoids an unachievable commitment. Separate the requirement by distance band, a near object and a distant one not being described with the same certainty. Distinguish the error families, an orientation inversion and a dimensional error not carrying the same consequence for the use. And compare each requirement with the disagreement measured between annotators on the same dimension. One important observation follows. The first step is specific to this field, a single requirement applied across a corpus being either untenable at range or needlessly loose up close.Controlling a segmentation
Five verifications apply where the task is point-wise classification. The distribution of classes per scene, an unusual proportion signalling an omission or a spillover. The rate of unclassified points, which must stay bounded and declared. Vertical consistency, ground points located above a structure signalling an error. Isolated islands, the characteristic signature of a badly tuned propagation. And comparison between annotators restricted to the boundary zones. One observation follows for a 3D annotation project. The last verification is the only representative one, global agreement on a scene being dominated by the large easy surfaces and able to coexist with massive disagreement where the point of the corpus lies.The evaluation corpus in 3D
One distinction deserves stating since this corpus does not follow the same rules as the training one. Four differences characterise it in 3D annotation. The volume needed is far smaller, a few scenes sufficing to measure. Exhaustiveness becomes imperative, an omitted object counting as a false detection by the system evaluated. Stratification by distance must be deliberate, a set dominated by near objects measuring a performance with no meaning. And stability over time conditions any comparison between two successive versions. One important practical consequence follows for a 3D annotation project. The second difference makes this corpus disproportionately expensive per scene, exhaustiveness forbidding the distance threshold and the point threshold that lighten the training corpus.What the report must contain
Six elements compose a useful control report in 3D annotation. The sampling method, with the distance stratification used. The results of the automatic checks, alerts raised and alerts confirmed. The disagreement measured between annotators, reported by distance band. The evolution of the indicators between batches. The list of configurations for which no stable convention could be established. And the dimensional range table used, which documents the extrapolations practised. One observation follows. The fifth element is the rarest and the most useful, since it turns an invisible limit into a declared one, a principle every preceding series established.What control costs
Four lines compose this 3D annotation control. Their distribution is surprising. Setting up the automatic checks, a one-off cost of a few hours. Running them, near nil and repeatable across every batch. Human verification of the alerts, proportional to the number of real defects. And double annotation of a sample, whose rate determines the surcharge. One practical consequence follows. This line is proportionally cheaper in 3D annotation than in image annotation at equal object volume, physical quantities supplying verifiable constraints images do not offer.The minimal arrangement
Five measures address most of the difficulties and their total cost stays modest. The dimensional range table per class, established on real objects and wired into the automatic control. The seven checks programmed before the first batch and run on every subsequent one. Top-view review of a sample stratified by distance, at every batch. Measuring disagreement between two annotators, conducted once at project start. And tracking the five indicators per batch, a simple table revealing drift. One observation follows for a 3D annotation project. Those five measures take a small share of production time, they require no proprietary tooling, and they turn a quality claim into a verifiable fact.Controlling the first batch
Four verifications apply to the first batch and to it alone. The feasibility of the conventions, a frequent uncovered case requiring a piece of writing before continuing. The throughput observed per distance band, which replaces the estimate and serves as the basis for costing. Disagreement between annotators, which bounds what can be promised contractually. And the validity of the dimensional range table, whose values are verified on real objects rather than on manufacturer documentation. One observation follows for a 3D annotation project. Those four verifications measure the arrangement’s viability rather than the batch’s quality, a distinction justifying treating the first batch differently from the following ones.Demonstrating quality to a client
Four demonstrations convince more reliably than a rate. Top-view review of an annotated scene, which shows in seconds an overall coherence a figure leaves abstract. The list of automatic alerts raised and handled, which demonstrates an arrangement rather than a result. The performance table by distance band, which honestly situates what is achieved and what is not. And the dimensional range table used, whose existence demonstrates a method indicators do not establish. One observation follows for a 3D annotation provider. The third support is the most effective commercially, a client understanding immediately that a provider declaring their limits by distance has measured their work, where a single figure suggests they have not.Approaching the control of a 3D corpus
Five questions scope a control arrangement in 3D annotation. Does a dimensional range table exist. Its absence deprives the project of its most effective check. Are the automatic checks planned from the first batch. Adding them late lets the first batches through. Is the sample stratified by distance. A uniform draw wastes the effort. Has disagreement been measured by distance band. Its value bounds what can be promised. And is a constant sample kept between batches. Its absence makes drift undetectable. Those five answers determine the arrangement and its cost. Asking them before production avoids observing after the fact a defect the first batches already contained.Three checks to set up first
Three arrangements address the majority of defects and they are built in a day. The dimensional plausibility check, which detects undersized boxes, this field’s commonest defect. Counting the points contained, which spots annotations with no support. And top-view review of a sample, which reveals aberrant orientations and omitted objects. Those three run on a delivered export, they need no external reference, and they cover what distinguishes a usable 3D corpus from a visually acceptable one. A corpus that passes all three is not necessarily correct, but a corpus that fails any of them is certainly not.The question that frames the arrangement
One question determines the requirement to adopt and it concerns the use. Out to what distance does the downstream system actually act. A short answer concentrates the requirement on a zone where the data is dense and disagreement low, which makes a high performance achievable. A long answer requires declaring a degraded requirement beyond a threshold, failing which the contract bears on a result sparsity forbids. That question is asked in one sentence, it belongs to the use and not to the technique, and it avoids this field’s commonest dispute, a client measuring on distant objects a performance promised across the whole.What this chapter teaches
One cross-cutting observation deserves closing this examination. Physics supplies here what convention must supply elsewhere. Three findings compose it. Dimensional, gravitational and impenetrability constraints bound what an annotation can assert, which makes a large share of defects detectable with no reference. Difficulty varies strongly with distance, which requires stratifying sampling, measurement and requirement. And the defects physics does not bound, omissions and systematic confusions, remain entire and require human control. That finding matches the one the preceding series established about controls exploiting a structural constraint: where the data carries a regularity independent of the annotated content, it serves verification as much as production.Why this field is easier to be honest in
One observation about commercial posture belongs here, since 3D differs from the other domains in this respect. Declaring limits is easier when the limits are physical rather than a matter of judgement. Three properties of 3D annotation produce that. Sparsity at range is a property of the sensor, not of the provider, so a degraded requirement beyond a threshold reads as a technical fact rather than as a hedge. The dimensional check produces evidence a client can run themselves, which makes a quality claim testable rather than asserted. And disagreement between trained annotators can be measured and shown, which situates the promise against what anyone could achieve. One practical consequence follows for a 3D annotation provider. Those three properties let a provider be more candid here than elsewhere without losing ground, since candour backed by physics reads as competence, while the same candour on a subjective judgement reads as weakness.Common mistakes
These failures recur often enough that naming them is usually enough to avoid them.- Omitting the dimensional range table per class.
- Sampling uniformly rather than by distance band.
- Reporting a single requirement for every distance.
- Aggregating inversions and angular errors in one indicator.
- Relying on volumetric overlap alone to diagnose a corpus.
- Adding the automatic checks after the first batches.
- Neglecting to track the distribution of dimensions per class.
- Expecting the automatic checks to detect omissions.
- Reading a fall in the uncertain rate as progress.
- Delivering a report without the values by distance band.