Predictive maintenance by vision differs from industrial inspection through an inversion of its relationship with time. On a production line, an image decides a part’s fate at the moment it is taken. On a structure or a piece of equipment, an image decides nothing on its own: it takes on meaning compared against the previous inspection, and it is the change between the two that drives the maintenance decision.
That difference has a direct annotation consequence. An infrastructure inspection corpus is not a collection of independent images but a set of time series, where every observation must be tied to a precise point on the structure and to a date. Without that link, you obtain a system able to say a crack exists and unable to say whether it is growing.
This article covers what is specific to this vertical: the defects sought, acquisition constraints by drone or fixed camera, the question of longitudinal tracking and what all of it implies for the corpus. It extends the complete guide to defect detection.
What separates predictive maintenance from production defect detection
Four structural differences separate the two contexts, and each feeds through to corpus design.
The object does not move past
In production, parts come to the sensor. In structure inspection, the sensor goes to the object, which inverts every constraint: position is not repeatable, lighting is whatever the day provides, distance varies, and the same area photographed twice six months apart is never captured from the same angle.
Rarity is temporal rather than statistical
A production defect is rare because it affects a small fraction of parts. Structural deterioration is rare because it takes years to appear. A corpus built from a single campaign therefore contains no information about change, and it is precisely that information which carries value.
The decision is not binary
Quality control accepts or rejects. Structure inspection produces a condition rating, generally on a multi-level scale, which triggers heightened monitoring, scheduled repair or urgent intervention. The system must therefore output a gradation, which presupposes that annotation captures it.
The stake is safety
A non-conforming part shipped costs money. A missed structural crack costs something else. That asymmetry pushes the whole system setting towards recall rather than precision, with the consequences that carries for verification volume.
The defects sought
The vocabulary of structural defect detection is well settled, notably for concrete. The field’s reference dataset fixes the categories: a background class and five defect classes, crack, spallation, exposed reinforcement bar, efflorescence and corrosion.
That taxonomy deserves comment, because it mixes three kinds of phenomena. Cracking and spalling are direct mechanical damage. Exposed reinforcement is an advanced consequence of spalling, therefore a stage rather than a distinct category. Efflorescence and corrosion staining are indications, meaning visible signs of a process whose seat is internal.
That distinction between damage and indication is decisive for annotation. A rust stain on a face does not signal a surface defect but reinforcement corrosion located behind it, whose real extent is invisible. The protocol must therefore state whether the observable sign or the supposed phenomenon is being annotated, and the answer is always the sign, otherwise the annotator is being asked for an inference they are in no position to make.
Beyond concrete
Other structure families have their own categories. On steel structures, corrosion breaks down into pitting, general corrosion, lamellar corrosion and section loss, where grading matters more than detection. On networks and pipelines, leaks, deformation and coating damage are added. On rotating equipment and wind turbine blades, leading edge erosion, delamination and skin cracks are the main targets.
One cross-cutting category deserves mention: vegetation and encroachment. On power lines, civil structures and rights of way, vegetation encroachment is a failure cause in its own right, and its detection falls to the same inspection arrangement without being a defect of the structure.
Acquisition constraints
This is the factor that most sharply distinguishes this defect detection vertical, and it largely determines what the corpus can contain.
The drone
The aircraft removes the risk of working at height and covers large surfaces quickly, but it introduces considerable variability. Distance to the object varies within a single flight, so effective resolution does too. Viewing angle depends on the trajectory. Stability governs sharpness. And natural lighting changes through a mission as much as between campaigns.
The annotation consequence is that the same crack can appear very differently across images, and that apparent size says nothing about real size without distance information. A useful corpus must therefore retain flight parameters, altitude, estimated distance, orientation, otherwise no quantification is possible.
The fixed camera
At the opposite end from the drone, a fixed station monitoring one piece of equipment offers excellent repeatability, which makes temporal comparison immediate. Its limits are coverage, restricted to the field of view, and progressive fouling of the optics in an industrial environment. In that configuration longitudinal tracking is the natural mode and the corpus builds itself.
Complementary modalities
Structure inspection frequently mobilises several complementary sensors. Thermography reveals delamination and water ingress invisible in the visible spectrum. Lidar provides geometry and allows an observation to be tied to a three-dimensional position. Other non-destructive methods complete the picture without involving image annotation.
That multiplicity raises an annotation question of its own: observations from different modalities must be tied to the same point on the structure, which presupposes a common reference frame. Without one, you obtain parallel corpora that cannot be cross-referenced.
Longitudinal tracking, the core of predictive maintenance
This is the specificity that justifies the word predictive, and it is also what most projects fail to put in place.
The registration problem
Comparing two inspections presupposes establishing that two images show the same area. On a fixed camera, that is given. On a drone flight, it is a real geometric problem: two campaigns follow different trajectories, and nothing guarantees a crack photographed in March appears in the field of a September image.
Three approaches occur in practice. The repeatable flight, where the trajectory is recorded and replayed identically, which simplifies everything but presupposes flight conditions permitting automation. Registration onto a three-dimensional model, where each image is projected onto a digital twin of the structure, offering the most robust link at the cost of heavy infrastructure. And manual identification by an operator, which remains common and introduces its own errors.
For annotation, the consequence is that registration must be established beforehand, not afterwards. A corpus of unlocalised images does not retrospectively become a tracking corpus.
Annotating change
Tracking introduces an annotation task absent elsewhere: comparing two states. It breaks into three distinct questions. Is this the same defect or a new one? Has it progressed, and by how much? Has a repair intervened in between, which would explain a disappearance without real improvement?
That last question is often forgotten and distorts the analysis. A corpus that does not record maintenance interventions contains inexplicable changes, and a model trained on it learns impossible transitions.
Quantification
Progression is only measurable if observations are quantified in a physical unit rather than in pixels. That presupposes a known scale, obtained through measured distance, a reference target in the field or geometric registration. A structural defect detection project that neglects this produces series of observations that are not comparable with each other, which nullifies the point of tracking.
Volume and sampling
A drone inspection campaign produces a considerable number of images, the vast majority of which show no defect. The ratio of collected to informative images is even less favourable than in production.
Three strategies reduce that volume to a workable defect detection corpus. Pre-filtering by a general-purpose model, which discards background images, sky or vegetation, at no risk since an error there is cheap. Selection by zone of interest, annotating first the critical structural elements identified by the engineer rather than the whole structure. And selection on prior history, targeting zones that showed an observation in a previous campaign, which feeds tracking directly.
That last strategy deserves priority in a maintenance-oriented defect detection project, since it produces exactly the data that carries value, namely series. Annotating a campaign uniformly produces a bulky corpus most of which will never be used.
Site specificity
One documented limitation deserves knowing before relying on public defect detection corpora.
A study on Taiwanese structures notes that the field’s reference dataset addresses defects under relatively controlled conditions but lacks the specificity to assess structures in regions where extreme environmental conditions, high humidity, coastal exposure, typhoons and seismic activity, accelerate deterioration in ways it does not represent.
That observation holds beyond the case cited. The appearance of deterioration depends on climate, material composition, structure age and local construction practice. A model trained on structures from one region transfers poorly elsewhere, and a project should plan for local adaptation rather than hoping for generalisation.
What annotation must produce in structural defect detection
Three levels of information coexist and must be distinguished in the protocol.
The first is detection and localisation of the defect within the image, a conventional defect detection task. The second is its qualification, category and severity, the latter resting on measurable criteria rather than an appraisal. The third is its registration, to an identified structural element and a position, which is not image annotation but contextual information without which the first two lose much of their value.
That third level explains why infrastructure inspection corpora cost more than they appear to. The work consists not only of delineating cracks but of building a base of localised observations, which demands as much documentary rigour as visual competence.
Profile and expertise
Role allocation follows the general logic of defect detection, with one inflection.
Detecting and delineating visible damage, cracks, spalling, corrosion staining, belongs to the perceptual register and delegates to trained annotators. Severity rating and structural interpretation belong to the inspection engineer or technician, since they entail a risk judgement.
The inflection is that this rating is often standardised. Structure inspection standards define condition scales and associated criteria, and the annotation ontology benefits from aligning with them rather than creating its own. That eases acceptance by inspection teams and allows system outputs to be compared against earlier inspections.
Industrial equipment wear
Alongside structures, a second family falls under predictive maintenance: production equipment itself.
The targets differ substantially from civil structures. Here one monitors cutting tool wear, belt and chain condition, fouling of exchangers and filters, leaks at fittings and seals, surface condition of rotating parts. The tempo differs too: these deteriorations are measured in weeks or months rather than years, which makes longitudinal tracking far quicker to build.
This family has a decisive advantage for a project: the equipment is accessible, the camera can be fixed, and the change cycle is short. A usable tracking corpus therefore builds in a few months where infrastructure takes years. It is often the best entry point for an organisation wanting to establish the value of a vision approach before extending it to structures.
The rhythm of campaigns
One organisational constraint specific to this defect detection vertical deserves anticipating.
Structure inspections are periodic, sometimes annual or multi-year, which means a tracking corpus builds over several years. A project therefore cannot wait for long series before starting, and must be designed to enrich progressively.
Two practical consequences follow. First, archives of past inspections, often kept as photographs and reports, have considerable value and are frequently underused: they are the only access to prior history. Their retrospective annotation is a project in itself, with its own difficulties, images of uneven quality, missing metadata, approximate localisation.
Second, corpus design must anticipate from the first campaign what later comparison will require. A campaign that neglects precise localisation will never be comparable to subsequent ones, and that loss is permanent.
What the system cannot see
Honest defect detection scoping requires stating the limits, and they matter on this vertical.
Visual inspection reports only on the surface. Corrosion under coating, internal section loss, incipient fatigue, sealing or anchorage defects remain invisible and fall to specific non-destructive methods. A vision-based defect detection system therefore does not replace technical inspection, it covers part of it and directs the rest.
That complementary position is in fact a stronger commercial argument than a promise of exhaustiveness. Vision brings systematic coverage, repeatability and a record, where human inspection brings judgement and access to internal methods. A project presenting the system as a prioritisation tool, tasked with saying where to look rather than concluding, obtains markedly better buy-in from inspection teams.
Quality control of a structural defect detection corpus
The control arrangement follows the general logic, with two inflections imposed by the domain.
The first is stratification by acquisition condition. A corpus mixing images taken at varying distances, angles and lighting must be controlled separately per stratum, otherwise the aggregate measure says nothing.
The second concerns temporal consistency. On a tracking programme, the costliest error is not misdelineating a crack but delineating it differently from one campaign to the next, which produces an apparent progression with no physical reality. Control must therefore include a specific check: having the same person reannotate images from an earlier campaign, and verifying the result stays consistent with the original annotation.
The economic value of tracking
One final point deserves stating, because it determines what a client will pay for.
Detecting a defect in a single image has limited value: an experienced inspector sees it too, and often better. The value of a vision system in maintenance comes from elsewhere, and breaks into three contributions. Systematic coverage, where every zone is examined every campaign with no fatigue effect and no implicit prioritisation. Repeatability, which makes two campaigns comparable where two different inspectors would not be. And the record, which supports justifying a maintenance decision and documenting a structure’s condition over time.
That breakdown shapes the scoping. A project sold on detection performance compares unfavourably against an inspector; a project sold on building a quantified and comparable history answers a need human inspection does not cover. That second framing is also what justifies investing in registration and quantification, without which none of the three contributions exists.
Access, safety and the annotation window
One practical constraint shapes these defect detection projects more than its mundanity suggests: getting to the structure at all.
Inspection campaigns depend on weather windows, flight authorisations, traffic closures and site availability. A bridge inspection may require lane closures scheduled months ahead; an offshore structure may be reachable only in a narrow seasonal window. The consequence is that acquisition opportunities are scarce and non-repeatable, which changes the risk profile of the whole project.
Two implications follow for annotation planning. First, the acquisition protocol must be validated before the campaign rather than after, since a campaign that produces unusable images cannot simply be rerun next week. A short preliminary flight, reviewed by whoever will annotate, catches framing and resolution problems while they are still fixable. Second, everything worth capturing should be captured while access exists, even beyond the immediate scope, since storage is cheap and a second visit is not. Selecting what to annotate can happen later; collecting what was never photographed cannot.
The most common mistakes
These failures recur often enough across inspection projects that naming them is usually enough to avoid them.
- Building a corpus of independent images where the value lies in time series.
- Not retaining flight parameters, which rules out physical quantification.
- Asking the annotator to infer an internal phenomenon from a surface sign.
- Neglecting registration to a structural element, information not reconstructible afterwards.
- Not recording maintenance interventions, which produces inexplicable changes.
- Quantifying in pixels rather than in a physical unit.
- Assuming a public corpus built in another climatic context transfers.
- Creating an in-house severity scale where an inspection standard exists.
- Ignoring photographic archives of past inspections, the only source of prior history.
- Not checking annotation consistency between successive campaigns.
Working with existing inspection practice
Every structure of consequence is already inspected, usually under a regulated regime with defined periodicity, qualified inspectors and standardised reporting. A vision project enters an established practice rather than an empty field.
That has three practical implications. The ontology should map onto the existing condition rating rather than beside it, so that system outputs can be read against decades of prior reports. The inspection reports themselves are a data source, since they record what was observed where and when, and often reference photographs. And the inspectors are the domain reference for the project, which makes their involvement in protocol design a requirement rather than a courtesy.
There is also a professional dimension worth handling carefully. Inspection is a regulated responsibility carried by qualified individuals, and a system presented as replacing their judgement will meet resistance that no technical argument dissolves. Presented as extending their reach, covering surfaces they cannot easily access and building the quantified history their reports cannot, the same system is welcomed. That framing is not merely diplomatic: it is also the accurate description of what the technology does.
What to take away
Predictive maintenance by vision applies defect detection principles to an object that does not move past, under uncontrolled conditions, and on a timescale measured in years. What gives it value is not detecting a defect in an image, which other verticals handle better, but the ability to establish change.
Three decisions structure a successful project. Build the registration, structural element, position and date, from the first campaign, since it cannot be reconstructed and governs every later comparison. Annotate the observable sign rather than the supposed phenomenon, and reserve structural interpretation for the inspection technician. And quantify in a physical unit, which presupposes a known scale and alone makes progression measurable.
For approaches, ontology and annotation tasks, the complete guide to defect detection sets the frame. For the question of tuning a system whose errors do not cost the same in each direction, particularly acute where safety is at stake, the article on false positives and false negatives covers metrics and thresholds.
To explore delivery arrangements, supported formats and applicable control mechanisms, see our dedicated page on annotation for industry. And if you are preparing a structure inspection corpus and want registration and tracking scoped before campaigns start, let us discuss your project.