Why Defect Detection Needs Trade-Trained Annotators

One figure sums up the problem industrial annotation poses. On a tool wear segmentation task, expert annotators achieved a mean intersection over union of 0.8153 while novices scored significantly lower. Same task, same protocol, same tool: the gap comes down to some knowing what they are looking at and others not.

The same source reports something more troubling still. A welding defect detection study showed that annotator habits create systematic bias: left boundary annotations were consistently less accurate than right boundaries, causing measurable performance degradation as the effect accumulated. A bias invisible in agreement metrics, and present in the corpus nonetheless.

This article covers the competence industrial defect detection actually requires: what it consists of, why inter-annotator agreement does not measure it, and how to organise exchanges between the annotation team and the client’s technical staff. It extends the complete guide to defect detection.

What defect detection actually demands

The question is not whether experts or non-experts are needed in defect detection, but exactly what knowledge the task requires. It breaks into four distinct registers.

The visual register

The first register consists of recognising a scratch, a missing area or an inclusion in a given image. It is the most obvious and least problematic: it is acquired through training, verified by sampling and improves quickly in a regular annotator. It requires no process knowledge.

The normality register

The second consists of knowing what is usual on this specific product. A machining mark, a variation in gloss, a handling trace are normal to someone who knows the part and suspicious to someone encountering it for the first time. This is where the gap between a product-trained annotator and a generalist becomes most visible, and it does not close by reading the protocol alone.

The consequence register

The third consists of understanding why a distinction exists in the ontology. Two defects of similar appearance can have opposite consequences, one compromising function and the other being purely cosmetic. An annotator unaware of that difference will decide borderline cases arbitrarily, whereas one who knows it will decide in the direction of the consequence, which is precisely what the client expects.

The decision register

The fourth consists of ruling on conformity against a quality standard. This register does not belong to the annotator but to the client’s quality function, and trying to delegate it produces a corpus whose validity nobody owns.

That four-register breakdown settles the false debate. The first two delegate to trained annotators. The third is acquired through organised knowledge transfer, which is possible but must be budgeted. The fourth stays with the client.

Why inter-annotator agreement is not enough

The usual reflex in defect detection is to measure agreement between annotators and conclude that the corpus is sound. That reasoning has a documented flaw worth knowing.

An analysis of expert annotation states it bluntly: standard inter-annotator agreement checks measure consistency, not correctness, and a group of generalists can achieve a Cohen’s Kappa of 0.85 on a systematically wrong label schema.

The underlying mechanism is explained in the same source: faced with content they do not master, non-specialist annotators apply the best inference available, meaning pattern matching and the path of least ambiguity. That inference is directional, so a crowd of non-specialists converges on the same wrong interpretation, their shared knowledge anchors producing the same predictable gap.

A second source confirms and refines the point: high inter-annotator agreement can sometimes mask low-quality or superficial judgements, particularly when annotators share common biases or lack subject matter expertise. It draws a clear methodological recommendation: compare individual judgements against expert-provided labels rather than against other annotators, since consensus does not always equate to correctness.

The practical consequence for a defect detection project is direct and immediate. High agreement on difficult classes should prompt verification rather than satisfaction, and the control arrangement must include comparison against an expert reference rather than only a measure of internal consistency.

What training actually transfers

The literature converges on a point that should reassure defect detection projects: the novice-expert gap is not irreducible, it narrows through a structured arrangement.

A synthesis on inter-annotator agreement puts it this way: novice annotators typically introduce higher variance, but consistency improves through structured feedback, curated examples and repeated exposure. The same source notes that in knowledge-intensive domains, experts are often essential to apply specialised concepts consistently and to resolve ambiguities non-experts would misinterpret.

Three arrangements genuinely transfer domain knowledge to a defect detection team, and they compound.

The commented case library comes first and underpins the other two. It is not a plate of examples but a set of cases where each decision is explained by its reason: why this mark is normal on this part, why a similar appearance belongs to another class, what physically distinguishes the two. That library is built during the pilot batch and grows with every arbitration.

The collective calibration session comes next. Having the whole team annotate the same batch, then discussing the discrepancies collectively with a quality referent present, transmits in half a day what a document does not transmit in ten pages. On tasks with a strong tacit component it is irreplaceable, and it repeats periodically.

The site visit, finally, is very substantially underestimated. Seeing the part, handling it, understanding the gesture that produces the defect changes how an annotator reads an image. Where a visit is impossible, a video call from the station with an operator, or simply photographs of real parts alongside the inspection images, reproduces part of the effect.

The cost of knowledge transfer

That transfer has a real cost, and stating it honestly is better than burying it in a unit price.

It comprises three components. The quality referent’s time for scoping and arbitration, the annotation team’s time in training and calibration, and building the case library, which is a deliverable in its own right. On a defect detection project with a rich ontology, that set represents a significant share of the initial budget, and it concentrates entirely at the start.

That concentration explains an important economic characteristic: the unit cost of early batches is structurally higher than in steady state. A quotation showing a constant unit price across the whole project therefore describes reality poorly, and a client comparing offers on that figure alone is comparing different things.

It also explains why small projects come out proportionally expensive. Knowledge transfer does not scale down with volume: training a team on a product costs roughly the same for two thousand images as for twenty thousand, which argues for grouping needs rather than fragmenting them.

The same competence does not travel

One point deserves stating plainly, because buyers and providers alike routinely misunderstand it.

Domain knowledge acquired on one defect detection project does not transfer to the next. An annotator experienced on rolled steel defects is not immediately competent on weld defects, and one accustomed to electronic boards is not competent on wood. What transfers is methodological: sweeping discipline, protocol rigour, the ability to recognise that a case exceeds their scope and must be escalated.

That distinction has a rarely stated contractual implication. A team stable on one client accumulates real value, and its turnover is expensive: every new arrival requires the transfer to be redone. An industrial project therefore benefits from favouring team continuity, which is a more relevant provider selection criterion than available headcount.

Annotators from the trade

An intermediate route deserves consideration in defect detection staffing: recruiting annotators with prior industrial experience, quality inspectors, production operators, maintenance technicians.

The advantage is real and measurable on both the normality and consequence registers. Someone who has held an inspection post knows what an ordinary part looks like, understands what a defect implies downstream and spontaneously asks the quality referent the right questions. Knowledge transfer is faster and goes further.

Two reservations accompany that observation and should be weighed. The first is that experience acquired on one process is not neutral: someone accustomed to a particular quality standard may implicitly apply their own thresholds rather than the protocol’s, producing a subtle bias that is hard to detect. Initial calibration must therefore pay specific attention to this. The second is that such experience costs money and is not justified on every task, which returns to the register breakdown rather than a general principle.

Where the trained annotator outperforms the expert

A point worth making explicitly, because the article’s argument runs one way and the reality runs both.

On repetitive tasks with a stable protocol, a full-time trained annotator is more consistent than an expert working intermittently between other duties. Consistency and expertise are distinct qualities, and on well-framed work the first matters more. An expert who annotates for two hours a week never builds the sweeping discipline, never internalises the boundary conventions and never develops the reflex of escalating rather than deciding. Their individual judgements may be better; their corpus is worse.

This is why the productive arrangement is almost never expert-only. It is a trained team producing under a protocol the expert defined, with the expert’s time reserved for the decisions that genuinely need it. Projects that assign everything to experts stall for lack of availability, and projects that assign everything to generalists drift for lack of grounding. The economics and the quality both favour the middle.

What expertise does not solve

A counterpoint is needed, because the expertise argument can be pushed too far and become a sales pitch rather than an analysis.

First, expertise does not compensate for an absent protocol. An expert without written instructions applies their own judgement, which differs from another expert’s, and the resulting corpus is consistent with nobody. The literature documents this: variability persists between experts, including on high-stakes tasks.

Nor does expertise compensate for faulty acquisition. If the defect is not visible in the chosen optical setup, no level of competence will make it appear, and the annotator will merely invent decisions on an absent signal.

Finally, expertise is not required everywhere or on every task. On frankly perceptual tasks, spotting contrasted objects, counting, assessing image quality, a trained annotator reaches the same quality as an expert at far lower cost. Demanding expertise on those tasks means paying for a competence that goes unused.

Iterations with technical teams

Domain knowledge does not circulate spontaneously in a defect detection project. It circulates if meetings exist, with a defined format and frequency.

Initial scoping

A session with the quality manager and an experienced operator together, before any annotation, to build the ontology from the existing standard rather than alongside it. This is where the most structuring question is settled: which distinctions in the quality vocabulary are visually grounded and which are not.

The periodic arbitration meeting

A short regular meeting where accumulated contested cases are settled. Three simple rules make it effective: cases are prepared in advance as a plate, every decision is justified in one sentence, and the decision is redistributed to the team then folded into the case library. Without that last step, the same question will resurface in a month.

The end-of-batch review

A short and often omitted format deserves establishing: a review at each batch delivery presenting the broken-down metrics, the escalated cases and any protocol changes during the period. It takes half an hour, it gives the client visibility no written report replaces, and it turns an execution service into technical collaboration. It has a useful side effect: it forces the annotation team to articulate what it has understood of the product, which quickly exposes persistent misunderstandings.

The loop with the machine learning team

The most instructive feedback in any defect detection project comes from jointly examining the errors of the first trained model. They regularly point to protocol ambiguity rather than algorithmic weakness: a poorly defined class, a missing boundary convention, a category the corpus does not cover. Establishing that joint review after the first training run is among the project’s highest-return meetings.

Production feedback

Once the system is deployed, disagreements between machine and operator constitute the richest information stream. Routing them back to the annotation team, rather than letting them dissipate on the shop floor, turns passive operation into an improvement loop.

Measuring competence in defect detection rather than assuming it

Three arrangements verify that transfer has actually happened, and they beat trust.

Qualification before production puts each candidate annotator in front of a set of cases whose truth was established by the quality referent, against an explicit pass threshold. It avoids injecting learning-phase annotations into the corpus and provides an objective measure, reusable to document contributor competence.

Continuous insertion of reference cases into the normal production flow measures individual drift. On a repetitive task that drift is real and progressive, and it stays under alert thresholds as long as nobody looks for it.

Periodic comparison against an expert reference, finally, directly addresses the internal-agreement flaw noted above. Having the quality referent settle a sample periodically, and comparing the team’s annotations against that reference rather than against each other, is the only way to detect a shared bias.

What the client must supply

Knowledge transfer presupposes someone supplying it, and that contribution cannot be bought externally.

Four concrete inputs are expected from the industrial client on a defect detection project. An identified and available quality referent, with an agreed time allocation rather than goodwill. The existing quality standard, however imperfect, which is the ontology’s starting point. Access to real parts or their photographs, which beats any written description. And availability for arbitration, whose frequency is agreed at kickoff.

The absence of any one of those four is the best predictor of a defect detection project that drifts. A provider agreeing to start without an identified referent does a short-term favour and stores up trouble: contested cases accumulate, nobody settles them, and the team ends up deciding questions that are not theirs to decide.

The resulting organisation of a defect detection chain

The effective chain in industrial defect detection has four levels, each defined by the decision register it handles.

The product-trained annotator handles the visual and normality registers, which is most of the volume. The reviewer, recruited from the strongest after several months, verifies by sampling and above all triages what must be escalated. The client’s quality referent settles cases in the consequence register and validates the protocol. The project manager keeps the case library alive, which looks administrative and in fact determines corpus consistency over time.

This four-level organisation has a favourable economic consequence. The quality referent’s time, a scarce and expensive resource, concentrates on definition and arbitration rather than production, which makes a corpus of tens of thousands of images sustainable without committing a rare competence full time.

Building the team over time

One consequence of everything above is rarely planned for in defect detection projects: the annotation team a project needs at month twelve does not exist at month one, and cannot be recruited into existence.

The reviewer who triages effectively acquired that judgement through months of production on this product. The annotator who recognises an unusual but normal appearance saw hundreds of ordinary ones first. Neither capability can be bought ready-made, because both are specific to the client’s material and process.

The practical implication is a ramp built into the plan rather than discovered. Early batches carry a heavier control ratio, because the team is still learning and the protocol is still moving. That ratio relaxes as the case library fills and reviewers emerge from the group. Compressing the ramp to hit a delivery date is the single most common cause of industrial annotation projects that finish on schedule and fail on quality, and it is entirely avoidable at planning stage.

The most common mistakes

These failures recur often enough across industrial projects that naming them is usually enough to avoid them.

  • Concluding a corpus is sound from inter-annotator agreement alone.
  • Not periodically comparing annotations against an external expert reference.
  • Assuming competence acquired on one material transfers to another.
  • Treating training as a kickoff meeting rather than a continuous arrangement.
  • Neglecting transfer of normality knowledge, the hardest to document.
  • Demanding costly expertise on frankly perceptual tasks.
  • Handing down arbitrations without justifying them or folding them into the case library.
  • Accepting high team turnover without costing the retransfer.
  • Expecting expertise to compensate for an absent protocol or faulty acquisition.
  • Not establishing a joint review of the first trained model’s errors.

One final remark on the vocabulary used in commercial offers. The terms expert, specialist and trade-trained circulate without definition and cover very different realities. The useful question to ask a provider is not whether they employ experts, but how they organise product knowledge transfer, how they verify it has happened, and what they expect from the client to make that transfer possible. The answers to those three distinguish established practice from a sales argument.

What to take away

The competence industrial defect detection requires is not a qualification but transferable knowledge, provided the transfer is organised. Knowing what is normal on a product and understanding why a distinction exists are acquired through a commented case library, calibration sessions and real contact with the material.

Three decisions structure a chain that holds. Break the task down by decision register, and require expertise only where the register demands it. Control by comparison against an expert reference rather than internal agreement alone, since a consensus can be unanimously wrong. And establish the meetings with the client’s technical teams, from initial scoping through to production feedback, since domain knowledge only circulates if a format exists to circulate it.

For approaches, ontology and annotation tasks, the complete guide to defect detection sets the frame. For the task where this competence is most heavily called on, fine defects whose delineation is decided pixel by pixel, the article on instance segmentation for fine defects covers precision requirements and boundary conventions.

To explore delivery arrangements, supported formats and applicable control mechanisms, see our dedicated page on annotation for industry. And if you are preparing an inspection project and want to organise domain knowledge transfer to the annotation team, let us discuss your project.

Tags

Découvrez nos articles