Five developments are currently reshaping three-dimensional perception. None removes the annotation work, and the most significant of them increases it rather than reducing it.
This article examines them. It extends the article on automatic pre-annotation in 3D.
The falling cost of sensors
This hardware development dominates the others by its effects on 3D annotation. Range sensors are becoming affordable for equipment that carried none. Three consequences follow for 3D annotation. The number of projects rises markedly, uses that made do with images moving to measurement. Cheap sensors produce sparser and noisier clouds than professional equipment. And the diversity of hardware grows strongly, which makes corpora less transferable from one set of equipment to another. One important observation follows. The second consequence runs against intuition, cheaper hardware producing more difficult data rather than more easy data.Solid-state sensors
This technical development changes the geometry of acquisitions. Sensors with no moving part are progressively replacing rotating mechanisms. Four consequences follow for a 3D annotation project. The field of view becomes partial, which requires several sensors for full coverage. The point distribution ceases to be regular, some patterns concentrating density at the centre. Sweep deformation disappears on fast carriers, which simplifies synchronisation. And the characteristic artefacts change in nature, which invalidates part of the experience acquired. One practical consequence follows. The second consequence is the most destabilising in annotation, a sparsity threshold expressed as a distance ceasing to hold uniformly where density depends on position in the field.Neural scene representations
This algorithmic development attracts attention. Its effect on 3D annotation is indirect. Methods reconstruct a continuous scene from multiple views. Three observations situate these methods in 3D annotation. They produce a visually convincing reconstruction from images alone. The geometry obtained is estimated and not measured, which deprives it of absolute scale. And the quality depends strongly on the number of views and their distribution around the scene. One important observation follows for a 3D annotation project. The second reproduces exactly the distinction between measurement and estimation set out about photogrammetry, which makes these methods visually appealing and unfit for metrological uses.Generalist segmentation models
This development bears directly on 3D annotation tooling. Models trained on vast corpora segment without being specialised. Four observations situate these models in 3D annotation. They work well on common and well-sampled objects. They do not know the domain categories, which must be associated with them. They inherit the conventions of their training corpora. And their performance falls on sparse clouds, exactly like that of specialised models. One practical consequence follows. The third observation requires a prior check, a generalist model applying its own extrapolation convention producing systematically offset dimensions.Sensors on light carriers
A logistical development widens the field of possible acquisitions. Light carriers now lift sensors that required a vehicle. Four consequences follow for 3D annotation. Inaccessible sites become surveyable, which opens entire verticals to measurement. Multiplying viewpoints becomes economical, which reduces residual occlusion. The clouds produced are sparser, the available payload limiting the sensor carried. And the trajectory becomes a further error source, registration resting on a less precise localisation than on the ground. One observation follows. The second consequence is the most favourable and the least exploited, occlusion being the difficulty acquisition corrects best and that no algorithmic progress compensates.Synthetic corpora
This development promises to reduce the need for 3D annotation. Its balance is mixed. A simulated scene supplies perfect and free ground truth. Three contributions characterise the synthetic corpus. Rare configurations are produced at will and without waiting. The ground truth is exact, including on faces no sensor would observe. And the volume is unlimited for a mere computing cost. Three limits accompany it in 3D annotation. The sensor’s real artefacts simulate badly, reflections and selective absences notably. The distribution of configurations reflects the designer’s imagination. And the gap between the synthetic and the real remains a source of failure that is hard to measure. One observation follows. Those three limits make the synthetic corpus a training supplement rather than a substitute, and they forbid its use as an evaluation corpus in 3D annotation.What none of these developments changes
Five factors stay stable and they determine the difficulty of 3D annotation work. Sparsity at range, a physical property technology pushes back without removing. Occlusion, which depends on the geometry and not on the sensor. Ambiguity between objects of the same shape, which only complementary information lifts. The need for an extrapolation convention, as soon as a face is not observed. And the absence of a shared convention on orientation and on the frame used. One important practical consequence follows. The last factor is the most striking, none of the five developments examined producing a common convention, which leaves that question exactly where it was.The regulatory requirements
A non-technical development changes what a corpus must carry. The frameworks applicable to artificial intelligence systems require data traceability. Four requirements follow for any 3D annotation corpus. Data provenance, site, date and sensor, must be documented. The annotation conventions must be written and retained. The quality control arrangement must be described and its results retained. And the limits of the domain of use must be declared explicitly in the documentation. One important observation follows. Those four requirements cover exactly the practices the preceding chapters recommended for operational reasons, which makes compliance a by-product of properly conducted work rather than an added burden.What these developments imply for a corpus
Four practical consequences follow for a 3D annotation project. The sensor metadata becomes decisive, hardware diversity increasing. The sparsity threshold must be expressed as a point count rather than as a distance, density ceasing to be uniform. The extrapolation convention must be declared, generalist models imposing one by default. And the evaluation corpus must stay entirely real, synthesis not measuring what counts. One observation follows. Those four consequences reinforce requirements already present rather than creating new ones, which suggests a corpus correctly documented today will pass through these developments without becoming unusable.What makes a corpus durable
Five properties determine whether a 3D annotation corpus will survive the developments examined. Retention of the source clouds at their original resolution, which permits re-exploitation with other conventions. Organisation by sensor and by campaign, which permits extracting a homogeneous subset. Explicit declaration of the frame, the orientation and the sparsity threshold. Marking of extrapolated faces, which distinguishes measurement from convention within the corpus itself. And retention of the coverage mask, which preserves the distinction between an empty zone and an unobserved one. One practical consequence follows. Those five properties cost little at the time of building and cannot be added later, which makes them the best protection against an obsolescence whose form nobody knows.What is not developing
Three expectations recur in discussions of this field and none is being met. That better sensors will remove the need to declare a convention, which no increase in density achieves since the unobserved face stays unobserved. That models will settle orientation on partly observed objects, which no training achieves since the information is absent rather than hard to extract. And that a standard will emerge from practice, which has not happened across the years these tools have existed. One observation follows for a 3D annotation project. The third expectation is the most costly to hold, since waiting for a standard postpones the local declarations that would make a corpus usable in the meantime.What could really change
Three developments would have a structural effect on 3D annotation. None is assured. A shared convention on the frame and orientation, which would make corpora combinable. An exchange format carrying the sensor and calibration metadata, which would remove the loss of information at conversion. And a reliable method of estimating the real dimension on an unobserved face, which would remove conventional extrapolation. One practical consequence follows for a 3D annotation project. The third is the least likely, information absent from the data being producible only by an assumption, which returns it to a convention under another name.What these developments change for the trade
Four shifts affect the daily practice of 3D annotation. The geometric share of the work falls, preprocessing and fitting automating further. The conventional share rises, each new source requiring a declaration of what distinguishes it. Control shifts towards detecting systematic biases rather than isolated errors. And the scoping competence progressively takes precedence over the execution competence. One practical consequence follows. The fourth shift matches what the preceding chapters established, a provider’s value lying in their capacity to ask the right questions before production as much as in producing fast.What distinguishes this field from the others
Four differences separate 3D annotation from image annotation and from video annotation. The data carries physical quantities, which supplies control constraints neither image nor video offers. The absence of information is absolute rather than progressive, which makes it declarable instead of blending with valid data. Difficulty varies with distance continuously and measurably, which permits stratifying requirement and sampling. And part of the dimensions delivered belongs to a convention rather than to an observation, which has no equivalent elsewhere. One observation follows. Those four differences work overall in favour of verifiability, which makes a properly documented 3D corpus more defensible than an image corpus of equivalent quality.What these developments do not replace
Four 3D annotation decisions stay human and none of the developments examined touches them. The choice of geometry, which depends on the downstream use and not on the technique available. The choice of maximum distance, a trade-off between cost and coverage nothing automates. The framing of the extrapolation convention, which is a decision and not a measurement. And the declaration of limits, which presupposes knowing precisely what the corpus does not contain. One observation follows for a 3D annotation project. Those four decisions are taken in one meeting and determine the result more than the choice of sensor does, which explains why they constitute a provider’s point of value rather than their production capacity.How to follow these developments
Four practices permit a 3D annotation team to stay current without chasing every announcement. Testing a new sensor on a known scene rather than on a manufacturer’s demonstration. Checking the dimensional convention of any model before using it, whatever its reputation. Keeping an unassisted share in the batches, which measures the real effect of each new tool. And distinguishing an improvement in the data from an improvement in its appearance, only the first reducing the work. One practical consequence follows for a 3D annotation project. The last practice is the most discriminating, several current developments producing data that is more appealing on screen without reducing the undecidable share that determines the cost.What the market will ask for
Four demands are emerging and they orient what a 3D annotation provider must be able to do. Handling multi-sensor corpora, hardware diversity increasing within one client. Taking over existing corpora, under-documented and needing requalification before use. Producing the traceability documentation the regulatory frameworks require. And scoping advice, distinct from production and bearing on geometry, maximum distance and conventions. One observation follows. The second demand is the least visible and the most frequent, a client almost always holding earlier surveys nobody can say what they permit, a question four verifications answer.What a client should ask for today
Five requirements protect a 3D annotation corpus against the developments examined. Delivery of the source clouds at original resolution, in addition to the lightened version used for annotation. Documentation of the frame, the orientation and the sparsity threshold used. The attribute clearly distinguishing measured faces from extrapolated ones. The sensor and campaign metadata for each scene. And the list of configurations not covered, which bounds the domain of use. One practical consequence follows. Those five requirements are stated in one clause of a specification, they barely raise the cost for a provider who works properly, and they rule out those who do not.Approaching these developments in a project
Five questions scope them in a 3D annotation project. They are asked at the moment of choosing. Does the sensor under consideration produce a uniform density. A negative answer changes how the threshold is expressed. Will a generalist model be used. Its convention must be checked beforehand. Is a synthetic share planned. It must stay outside the evaluation corpus. Is the sensor metadata recorded. Its absence will limit reuse. And must the corpus stay usable after a change of hardware. That answer imposes an organisation by sensor from the outset. Those five answers determine the corpus’s robustness over time. Asking them at the moment of choosing avoids a corpus tied to a hardware configuration that will not last.The question that runs through these developments
One question permits evaluating any announced novelty. Does this development reduce the undecidable share or displace it. A real reduction comes from acquisition, additional viewpoints, increased density or an added modality, and it durably diminishes the human work. A displacement comes from processing, a reconstruction or a proposal making legible what stays undetermined, and it transfers the decision without removing it. That question is asked in front of every announcement, it requires no technical expertise, and it separates the advances that reduce a budget from those that improve a demonstration.What this cluster has established
Five findings run through this whole examination of 3D annotation. Orientation is the quantity with no equivalent in image annotation, and its one hundred and eighty degree ambiguity constitutes its characteristic defect. Physical quantities make automatic control more powerful than elsewhere, dimensions, gravity and impenetrability supplying verifiable constraints. Distance is a coverage axis in its own right, a near object and the same object at range constituting two distinct examples. The distinction between measurement and estimation runs through every chapter, from the choice of sensor to the extrapolation of unobserved faces. And the absence of a shared convention on frame and orientation remains the principal obstacle to combining corpora. One observation follows. Those five findings belong to declaration rather than to technique, which makes them accessible to any project independently of its budget.Three decisions that outlast the technology
Three decisions protect a 3D annotation corpus whatever development comes. Keeping the source clouds at original resolution, the only decision that permits re-exploitation with future conventions. Organising the corpus by sensor and by campaign, which permits extracting a homogeneous subset when the equipment changes. And declaring explicitly what the corpus does not cover, which bounds its domain of use instead of leaving a user to discover it through a failure. Those three decisions depend on no technology, they cost little at the time of building, and they constitute the only protection available against an obsolescence whose form nobody knows.What this chapter teaches
One cross-cutting observation deserves closing this cluster. Technical developments displace the difficulties without removing them. Three findings compose it. Cheaper sensors produce more difficult data rather than more easy data, which increases the need for annotation instead of reducing it. The methods that estimate a geometry without measuring it reproduce an old distinction in a new form. And none of these developments produces a shared convention, which leaves the principal obstacle to combining corpora exactly where it was. That finding matches the one the video cluster established about its own trends, and it holds for this whole examination: what limits a corpus lies in what is declared rather than in what is measured, and that limit is lifted by no hardware progress.What stays true across the cluster
One closing observation situates the whole of this examination. Everything set out across these fifteen chapters reduces to a single asymmetry. The geometry can be measured, and the measurement is verifiable by anyone holding the data. The convention cannot be measured, and it is knowable only if someone writes it down. Three consequences follow for a 3D annotation project. A corpus is judged less on the care of its geometry than on the completeness of its declarations, since the first is checkable and the second is not. A provider’s principal contribution lies in the decisions taken before production, since those decisions are what the data cannot supply. And the defects that survive delivery are almost always conventional rather than geometric, since the physical constraints catch the rest. One practical consequence follows. That asymmetry explains why every chapter in this cluster ends on a declaration rather than on a technique, and why the cheapest improvement available to most projects is writing down what was already being done.Common mistakes
These failures recur often enough that naming them is usually enough to avoid them.- Assuming a cheaper sensor produces easier data.
- Expressing a sparsity threshold as a distance on a non-uniform density sensor.
- Using a generalist model without checking its dimensional convention.
- Treating a neural reconstruction as a measurement.
- Using a synthetic corpus as an evaluation set.
- Neglecting the sensor metadata in a heterogeneous fleet.
- Expecting a hardware development to remove extrapolation.
- Transferring artefact experience from one sensor type to another.
- Assuming a shared convention will emerge on its own.
- Building a corpus tied to a single hardware configuration.