The preceding verticals share three difficulties their particular contexts inflect differently. This chapter treats them in their own right, because they bound what a three-dimensional corpus can assert independently of the domain.It extends the article on 3D annotation in construction.
The three families of difficulty
Three distinct causes produce the failures of 3D annotation.Sparsity, where the object is present and too sparsely sampled to be qualified.Occlusion, where the object is masked and absent from the data rather than degraded.And geometric ambiguity, where the object is correctly sampled and indistinguishable from another.One important practical consequence follows. Those three families call for different responses, the first a threshold, the second a convention and the third complementary information, which makes their confusion expensive in 3D annotation.What sparsity produces
This property governs the feasibility of a 3D annotation project.It can be quantified.The number of points on an object falls with the square of the distance.Four consequences follow in 3D annotation.The class becomes uncertain, a few points carrying no usable signature.The dimension rests on an extrapolation rather than on a measurement.Orientation becomes undeterminable, the front no longer being discernible.And disagreement between annotators grows strongly, on the class as on the dimensions.One important observation follows. Those four effects appear progressively rather than at a sharp threshold, which makes the choice of a maximum distance conventional and requires founding it on a measurement rather than on intuition.Fixing the sparsity threshold
This method turns an intuition into a measured value.The principle is to measure agreement between annotators by distance band.Three pieces of information follow for a 3D annotation project.The distance beyond which agreement on the class collapses.The distance beyond which orientation ceases to be determinable.And the point count corresponding to those two thresholds, a value transferable to other sensors.One important practical consequence follows for a 3D annotation project. The third is the most useful, a threshold expressed as a point count transferring from one sensor to another where a threshold in metres holds only for the equipment that produced it.What occlusion produces in 3D
One difference from image annotation deserves emphasis.What lies behind an obstacle produces no point, the information being absent rather than degraded.Four consequences follow.A partially observed object is dimensioned by extrapolation rather than by measurement.An entirely masked object is not annotatable, effort changing nothing.The observed face determines which dimensions are measured and which are inferred.And the extrapolation convention becomes the source of part of the dimensions delivered.One important observation follows. That absolute absence is more favourable than a progressive degradation, it can be detected and declared where degraded data blends with valid data in 3D annotation.The degrees of occlusion
Four situations are distinguishable in 3D annotation.They do not call for the same treatment.The object partly visible on several faces, dimensionable with little extrapolation.The object visible on one face only, whose depth is entirely conventional.The object visible through a few scattered points, whose class itself becomes uncertain.And the object entirely masked, absent from the data and therefore not annotatable at all.One practical consequence follows for a 3D annotation project. Those four situations justify a four-level occlusion attribute rather than a binary indicator, the second and third calling for very different treatment downstream.Geometric ambiguity
This third difficulty persists despite sufficient density.Two objects of identical geometry are not distinguished by shape.Four configurations illustrate it in 3D annotation.Two different materials forming the same flat surface.Two vehicles of similar size observed partially.A post and a pruned tree of comparable diameter.And two objects in contact presenting no measurable discontinuity between them.One important observation follows. Those four configurations do not resolve through additional effort, they call for complementary information, colour, intensity or context, or for a class declaring the indecision.The surfaces that deceive the sensor
Four materials produce false data rather than absent data.Glazing, which lets the signal through and makes the background appear in place of the surface.Polished metal, which reflects and produces points at positions that do not exist.Water, whose behaviour combines reflection and absorption according to the angle of incidence.And dark matt surfaces, whose weak return produces a selective absence rather than a uniform one.One important practical consequence follows for a 3D annotation project. The second material is the most dangerous, a reflected point being geometrically valid and physically false, which makes it undetectable by the ordinary plausibility checks.Detecting false points
One operation distinguishes absent data from false data and it can be programmed.Four signatures betray a reflected or aberrant point in 3D annotation.A position behind an opaque surface already measured, physically impossible.Isolation, a point with no neighbour within a given radius belonging to no real surface.Symmetry about a reflecting plane, the characteristic signature of a specular reflection.And an intensity abnormally high or low relative to the immediate neighbourhood.One important practical consequence follows for a 3D annotation project. The first signature is the most reliable and the least used, a point situated behind a measured wall being unable to exist, which supplies an automatic check neither distance nor isolation replaces.What acquisition can correct
Four upstream decisions reduce these 3D annotation difficulties.They escape the annotation work itself.Multiplying viewpoints, which reduces occlusion without changing the sensor.Bringing the sensor closer, which increases density on the objects of interest.Choosing the period, which changes penetration in vegetated environments.And adding a modality, image or intensity, which lifts a large share of the geometric ambiguities.One observation follows. The first decision is the most effective and the least expensive, a second station or a second pass removing part of the occlusion rather than making it more legible.What multiplying viewpoints brings
One correction deserves separate treatment since it is the most effective and the least used.Observing a scene from several positions removes part of the occlusion rather than making it legible.Four benefits follow for a 3D annotation project.Faces masked from one position become measured from another.Density increases on the objects seen several times.False points from reflections are detected by inconsistency between stations.And extrapolation recedes, part of the dimensions becoming measured rather than inferred.Three costs accompany it.Acquisition time, each additional station lengthening the outing.Registration between stations, which introduces its own error source.And the data volume, which grows in direct proportion to the number of stations.One practical consequence follows for a 3D annotation project. Those three costs are borne once at acquisition while the benefits carry across the whole annotation, which makes this trade-off almost always favourable and yet rarely examined.What the convention must provide for
Six convention decisions cover most of these difficulties. They are written before production.The minimum number of points below which an object is not annotated.The maximum distance covered, derived from the preceding threshold.The extrapolation rule for unobserved faces, with the attribute that flags it.The levels of the occlusion attribute.The treatment of aberrant points from reflective surfaces.And the class receiving objects whose nature stays undecidable even after careful examination.One practical consequence follows for a 3D annotation project. The fifth decision is the most often omitted, a reflected point being treated sometimes as an object and sometimes as noise according to the annotator, which produces an inconsistency nothing detects.What the annotator must be able to do
Five tooling capabilities condition the correct handling of these configurations in 3D annotation.Consulting the number of points inside a selection, which objectifies an impression of sparsity.Switching between raw cloud and coloured cloud, which reveals what colour masks.Displaying the sensor’s position, which explains an occlusion and guides the extrapolation.Marking an uncertain decision without interrupting the work.And flagging a suspect point without having to settle its nature on the spot.One important observation follows. The third capability is specific to this field and rarely available, knowing where the scene was observed from immediately explaining which face of an object is measured and which is inferred.The most difficult configurations
Five situations compound several of the difficulties set out and they deserve identifying at scoping.The glazed scene, where absence, reflection and apparent background combine.Dense bulk, where mutual occlusion and the absence of a geometric boundary compound.Long range outdoors, where sparsity and class ambiguity combine.Vegetation cover, where partial occlusion produces sparse and misleading data rather than absent data.And the metallic environment, where reflections produce phantom structures that look coherent.One observation follows for a 3D annotation project. The last configuration is the most insidious, a reflection on a large flat surface producing a geometrically plausible copy of a whole section of scene, which only a check for positions behind a measured wall reveals.What these difficulties cost
Four effects on the 3D annotation load accompany these configurations.Throughput falls sharply, a sparse object requiring examination where a dense one is handled in one gesture.Disagreement rises, which requires more frequent double annotation.Arbitration becomes necessary, an edge case escalating to a lead.And the undecidable share consumes time without producing usable data.One important practical consequence follows. Those four effects concentrate on the most distant and least useful objects, which makes the distance threshold both a saving lever and a quality lever.Measuring the undecidable share
One operation quantifies that share and it bounds any contractual requirement.Two annotators handle the same scenes and their disagreements are classed by cause.Four usable pieces of information follow.The share attributable to sparsity, which falls if the distance threshold is tightened.The share attributable to occlusion, which depends on the acquisition geometry.The share attributable to ambiguity, which calls for complementary information.And the share attributable to an insufficient convention, the only category correctable in writing.One important observation follows for a 3D annotation project. That classification directs the effort, the first three categories belonging to acquisition and the fourth to documentation, which avoids seeking in production a correction that lies elsewhere.Framing a tenable requirement
A three-step method avoids an unachievable commitment on these configurations.Separate the requirement by distance band, a near object and a distant one not being described with the same certainty.Distinguish measured quantities from extrapolated ones, a requirement not bearing on both in the same way.And declare the undecidable share explicitly, which bounds what a system can reach independently of its quality.One important observation follows for a 3D annotation project. The second step is specific to this field, a dimensional tolerance applied indiscriminately to a measured face and an extrapolated one committing to a precision the convention alone determines.The first batch and the difficult cases
Four scenes compose a pilot batch that genuinely tests these difficulties.A scene of ordinary density, which supplies the reference throughput.A scene containing objects at the sparsity limit, which verifies that the threshold is well placed.A scene with mutual occlusions, which tests the extrapolation convention and the occlusion levels.And a scene containing a reflective surface, which tests the treatment of false points.One observation follows. The second scene is the most informative, it supplies both the collapsed throughput on sparse objects and the disagreement between annotators at that distance, two values that together determine whether the threshold should be tightened in 3D annotation.Explaining these limits to a client
Four framings make this conversation productive rather than defensive.Show the configuration rather than describe it, an object reduced to a few points being worth any explanation.Distinguish what belongs to acquisition, to convention and to the irreducible, which allocates responsibility without evading it.Propose the upstream correction where one exists, an additional viewpoint being a constructive answer.And quantify the share concerned, a difficulty affecting a small fraction of the corpus not justifying abandoning the project.One observation follows for a 3D annotation provider. The second framing is the most useful here, the split between acquisition and annotation being sharper than in video, which permits designating a precise correction rather than observing a general limit.Three checks to put in place
Three verifications address most of these difficulties and they are programmed in a day.Counting points per annotated object, which objectifies sparsity and flags annotations with no support.Detecting points situated behind a measured surface, the only reliable signature of a reflection.And checking the consistency of the occlusion attribute, an object declared complete and yet missing a face constituting a contradiction.Those three checks run on a delivered export, they exploit the geometry rather than an external reference, and they cover what distinguishes usable data from data that merely appears usable. None of them requires knowing what the scene contains.What these limits imply for the corpus
Three consequences extend beyond the initial project and deserve stating at scoping.A corpus built with one sparsity threshold does not compare with a corpus built with another, the object population differing.A system trained on extrapolated faces learns a convention as much as a geometry, which ties its performance to the rule used.And the data-free zones must be declared rather than left implicit, failing which an absence reads as an absence of objects.One observation follows for a 3D annotation project. The third consequence is the most damaging and the easiest to avoid, a coverage mask delivered with the corpus distinguishing what was observed and judged empty from what was not observed at all.Approaching these difficulties in a project
Five questions scope these 3D annotation difficulties. They are asked before production.What share of the objects sits beyond the measured sparsity threshold. That answer sizes the principal difficulty.How many viewpoints cover the scene. That answer determines the residual occlusion.Is complementary information available. Its presence lifts part of the ambiguities.Which problematic materials does the scene contain. That answer predicts the false points.And is an undecidable class provided in the nomenclature. Its absence turns a limit into noise.Those five answers determine the conventions to write and the tenable requirement. Asking them before starting avoids committing to a precision the data does not carry.The question that precedes the others
One question determines whether these difficulties are blocking or marginal.What share of the objects of interest sits beyond the measured sparsity threshold.A low answer places these difficulties at the margin, a declared threshold and an undecidable class sufficing to absorb them.A high answer places them at the centre, which justifies an additional station, a denser sensor or a revision of the scope.That question is answered by counting objects per distance band on a few scenes, it precedes any discussion of method, and it avoids sizing a whole project on a difficulty affecting only a fraction of it.What this chapter teaches
One cross-cutting observation deserves closing this examination.These difficulties fall into three categories of which only one belongs to annotation.Three findings compose it.Part belongs to acquisition, density, viewpoints and modality being decided upstream or not at all.Part belongs to convention, thresholds and extrapolation requiring a written rule rather than additional effort.And part is irreducible, two objects of the same geometry with no complementary information not separating.That finding matches the one the video cluster established about occlusion and blur: distinguishing an error from a limit avoids spending effort correcting what cannot be corrected, and that distinction belongs at diagnosis rather than after several attempts.What to check on data before quoting
Four measurements qualify a dataset for these 3D annotation difficulties and they take an hour.The distribution of point counts per object across distance bands, which reveals where the sparsity threshold should sit.The number of viewpoints per scene, which predicts the residual occlusion.The proportion of reflective or transparent surfaces visible in a sample, which predicts the false points.And the presence of sensor position in the metadata, which determines whether extrapolation can be guided or only guessed.One practical consequence follows for a 3D annotation provider. The fourth measurement is the one nobody asks for and the one most often missing, its absence forcing annotators to infer which face was observed from the point distribution itself, which is possible and slow.Common mistakes
These failures recur often enough that naming them is usually enough to avoid them.- Fixing a distance threshold by intuition rather than by measurement.
- Expressing a threshold in metres rather than as a point count.
- Using a binary occlusion attribute rather than a multi-level one.
- Extrapolating an unobserved face without flagging it.
- Treating a reflected point sometimes as an object and sometimes as noise.
- Neglecting the undecidable class.
- Expecting additional effort to lift a geometric ambiguity.
- Ignoring the multiplication of viewpoints as an occlusion correction.
- Conflating sparsity and occlusion in one attribute.
- Requiring the same precision at every distance.
