Contents
- The governing idea
- Key terminology, defined against the inspection problem
- Principles
- Case study: SECOM wafer screening
- Quick reference
1. The governing idea
A classifier does not make a decision. It produces a probability, and someone has to choose where to draw the line.
That line — the decision threshold — determines whether the system catches defects or floods the line with false alarms, and therefore whether the project pays back. Choosing it requires knowing what an escape costs and what an unnecessary rejection costs.
Those are business figures. The threshold is an economic decision presented in technical clothing, and it belongs to whoever owns the cost of both failure modes.
Software defaults it to 0.5. That number carries no meaning whatsoever.
2. Key terminology, defined against the inspection problem
The four outcomes
Every classification prediction lands in one of four cells. Their names are unhelpful; their consequences are not.
| Model says good | Model says defective | |
|---|---|---|
| Actually good | True Negative (TN): Good part passes. No cost. | False Positive (FP): Good part scrapped or re-inspected. Costs material, handling, and capacity. Also called a false alarm. |
| Actually defective | False Negative (FN): Defect ships to the customer. Warranty, complaint, recall, reputation. Also called an escape or miss. | True Positive (TP): Defect caught. Costs the inspection and disposal. |
Never accept a confusion matrix without this translation. The four numbers are meaningless until each is attached to a business consequence, and those consequences typically differ by two or three orders of magnitude.
Performance measures
Accuracy — proportion of all predictions that were correct: (TP + TN) / total. Business meaning: almost none, when defects are rare. On a process with 3% defects, predicting "good" for everything achieves 97% accuracy while catching nothing. Actively misleading on imbalanced data.
Recall (also sensitivity, true positive rate, detection rate) — of all defects that existed, the fraction caught: TP / (TP + FN). Business meaning: the detection rate. 1 − recall is the escape rate — the proportion of defects reaching the customer. Usually the metric that matters most.
Precision (also positive predictive value) — of everything flagged, the fraction genuinely defective: TP / (TP + FP). Business meaning: 1 − precision is the false scrap rate. Also determines operator experience: precision of 0.07 means fourteen false alarms for every real defect.
F1 score — harmonic mean of precision and recall. Business meaning: limited. F1 embeds the assumption that a false negative and a false positive cost the same, which is almost never true in manufacturing. Acceptable for model exploration; strictly inferior to expected cost for the decision that matters.
Specificity (true negative rate) — of all good parts, the fraction correctly passed: TN / (TN + FP).
Curves and summary statistics
ROC curve — true positive rate against false positive rate, across all thresholds. Caution: flattering on imbalanced data. The false positive rate has a very large denominator (all good parts), so hundreds of false alarms barely move it. A model can show a respectable ROC and be operationally useless.
AUC / ROC-AUC — area under the ROC curve. 0.5 is random, 1.0 is perfect. Business meaning: essentially nothing operational. Do not accept AUC as evidence that a system will work on a line.
Precision-Recall curve — precision against recall, across all thresholds. Baseline is the defect rate itself. Business meaning: the achievable trade-off between escape rate and false scrap rate. On imbalanced problems, demand this rather than ROC.
Average Precision (AP) — area under the PR curve. The imbalanced-data counterpart to AUC.
Decision and policy terms
Decision threshold — the probability above which a part is called defective. Business meaning: this is the inspection policy. Lowering it means flagging on weaker evidence: stricter inspection, more alarms, fewer escapes. Raising it means more lenient inspection.
Terminology caution: "lenient threshold" and "lenient inspection" mean opposite things. Say strict inspection or lenient inspection and avoid describing the threshold itself that way.
Cost matrix — the monetary cost assigned to each of the four outcomes. The input that converts a statistical model into a business decision.
Alarm rate — the fraction of parts flagged: (TP + FP) / total. Business meaning: the inspection load created. Directly constrained by capacity.
Alarm fatigue — the well-documented deterioration of operator attention when a system produces predominantly false alarms. Business meaning: reduces effective recall toward zero while measured recall stays unchanged. A system optimal on paper and ignored on the line is worse than none, because it also consumes the organisation's appetite for the next attempt.
Class imbalance — a large disparity between class frequencies. Typical of manufacturing defect data, where 1–10% defect rates are normal.
Base rate — the underlying frequency of the positive class. The reference point for judging whether precision is meaningful: precision of 0.25 against a base rate of 0.066 represents a 3.8× lift.
Lift — how much better than random a model performs. precision / base_rate. Robust to cost assumptions, which makes it useful when costs are contested.
Handling imbalance
Stratified splitting (StratifiedKFold) — cross-validation that preserves the class ratio in every fold. Mandatory for classification. Without it, a fold may contain almost no defects and produce a meaningless score.
Class weights — instructing the model to treat minority-class errors as more costly during training.
SMOTE (Synthetic Minority Over-sampling Technique) — generating synthetic minority examples to balance the training set. Must be applied inside the cross-validation fold only; resampling before splitting places synthetic copies of test samples into training, which is straightforward leakage.
Calibration — whether predicted probabilities correspond to real-world frequencies. Both class weights and SMOTE change calibration, which shifts the optimal threshold. A threshold tuned on an unbalanced model does not transfer to a rebalanced one.
3. Principles
3.1 Accuracy is dangerous on imbalanced data
The demonstration, from a single dataset with three different targets and identical code:
| Target | Positive rate | Accuracy | Recall | Precision |
|---|---|---|---|---|
| K_Scatch | 20% | 97.6% | 0.92 | 0.96 |
| Dirtiness | 2.8% | 97.8% | 0.40 | 0.71 |
The dirtiness model has the higher accuracy and misses 60% of the defects it exists to find. The do-nothing baseline on that target achieves 97.2%, so the model's entire contribution is 0.6 percentage points.
Accuracy ranked these almost exactly backwards. Always compare against the majority-class baseline; on rare events it is nearly as high as anything achievable.
3.2 The threshold is where the money is
Software defaults to 0.5 for no reason but convention. The correct threshold comes from the cost matrix, and moving it requires no better model, no additional data, and no new sensors.
Method:
- Obtain out-of-fold predicted probabilities
- Sweep the threshold across its full range
- Compute the confusion matrix at each point
- Apply the cost matrix to obtain expected cost
- Select the minimum, subject to operational constraints
The cost-versus-threshold curve is usually asymmetric — steep on one side, flat on the other. That asymmetry indicates which way to err when cost estimates are uncertain, which they always are.
3.3 Always run sensitivity on the cost assumptions
Cost figures are estimates. Test whether the recommendation survives being wrong about them.
Two possible outcomes, both valuable:
- Stable across a wide range → say so explicitly. "The recommendation is insensitive to escape cost above £2,500; precise costing is unnecessary." This is the sentence that stops a business case stalling while someone attempts to price reputational damage.
- Swings wildly → the deliverable is no longer a threshold. It is a request that someone establish the true cost. A different and equally useful conversation.
3.4 An optimum on a boundary is a diagnostic, not an answer
When an optimiser runs to the edge of its search space, the objective function is missing a constraint that exists in reality.
Three instances occurred in a single analysis:
| Symptom | Missing constraint |
|---|---|
| Threshold pinned at the bottom of the sweep range | Search range too narrow |
| Threshold at 0.001, recall 1.00 — flag everything | No inspection capacity limit |
| Model "wins" at trivially low inspection cost | Constraints omitted from the comparison loop |
The cure is always the same: identify the real-world limit the model does not know about and encode it. This recurs in optimisation contexts generally — a formulation optimiser proposing a physically impossible recipe has the same cause.
Add to any evaluation checklist: "Where did the optimum land, and is it on a boundary?"
3.5 Cost optimisation must include operational constraints
An unconstrained cost model assumes unlimited inspection capacity and infinitely patient operators. Neither exists.
Two constraints belong in any inspection optimisation:
Alarm rate cap — the fraction of output that can actually be inspected. A policy demanding 100% manual review is not expensive; it is impossible.
Precision floor — the minimum ratio of real to false alarms that operators will sustain. Below it, alarm fatigue destroys effective recall while measured recall is unaffected.
If no policy satisfies both, that is a finding: the model is not good enough to deploy as an automated gate. It may still work as a triage aid.
3.6 Always compute the do-nothing and do-everything anchors
Three numbers are required before any inspection recommendation:
no_inspection = n_defects * COST_FN / n * 1000
inspect_all = (n_defects*COST_TP + n_good*COST_FP) / n * 1000
best_model = <constrained optimum>
Two lines of code, and they convert a plausible success into an honest answer. Most analyses never compute them, optimise the threshold, report the best number, and never discover that the simple option wins.
3.7 Rare events plus modest lift rarely justify a gate
The arithmetic is unforgiving. With a low base rate, a model must be extraordinarily good before flagging becomes cheaper than checking. A 3–4× lift over random is genuine signal and still insufficient to gate on.
Reframe rather than abandon. "Should the model replace inspection?" often fails where "if we can only inspect 5%, which 5%?" succeeds — and the triage benefit is independent of escape-cost assumptions, making it the more robust half of the case.
3.8 Identify which error mode dominates the total cost
A model that inspects few parts has a cost dominated by missed defects, not by inspection. In that regime:
- Improving precision reduces false alarms that cost almost nothing → near-zero value
- Improving recall is the only change that moves the total
- Hyperparameter tuning addresses neither at the required scale
This determines where investment should go, and it is a more useful conclusion than any accuracy figure. Most model reports recommend "further work". A useful one specifies which work and why the obvious alternative is pointless.
3.9 Identifiers must be dropped before modelling
Row IDs, sample numbers, and filenames carry no physical information, but tree models will split on them. If a file was assembled by concatenating groups, ID ranges encode the target and the model achieves excellent scores by memorising row numbers.
Verify:
df.groupby(df[target_cols].idxmax(axis=1))["id"].agg(["min","max","count"])
Cleanly separated ranges confirm the leak. Drop anything encoding where a record came from rather than what it measures.
Note that removing such a column makes performance worse. That is the correct outcome: the lower number is the honest one.
3.10 Change one thing at a time
Modifying two aspects of a pipeline simultaneously makes their individual contributions unrecoverable. The same discipline as design of experiments, and easily lost in code where changes are cheap.
4. Case study: SECOM wafer screening
Situation
A semiconductor line tests 100% of wafer output at end of line. The question: could a classifier trained on in-process sensor data predict failures well enough to reduce that burden?
Data. 1,567 wafers, 590 in-process sensors, pass/fail outcome. 104 failures — a 6.6% defect rate.
Approach
Sensors with >50% missing values and zero-variance sensors were removed, leaving 446 features; remaining gaps filled with column medians. A random forest classifier was evaluated by stratified 5-fold cross-validation, with all figures computed out-of-fold.
Rather than accepting the default threshold, expected cost per 1,000 wafers was computed across the full threshold range under a defined cost matrix, subject to an alarm rate cap of 15% and a precision floor of 25%.
Cost matrix: escape £5,000 (estimated); false alarm £40; correct flag £15; correct pass £0.
What went wrong first — and why it mattered
Three optimisation runs produced boundary-pinned results before the analysis was sound:
- Threshold pinned at the low end of the sweep — search range too narrow
- Threshold 0.001, recall 1.00 — no capacity constraint, so "flag everything" won
- Model apparently winning at £40 inspection cost — constraints omitted from the comparison
Each was diagnosed the same way: an optimum on a boundary means a real-world limit is missing from the objective. Only after both constraints were properly applied did the analysis produce a defensible answer.
Results
| Policy | Cost per 1,000 wafers |
|---|---|
| No inspection | ~£332,000 |
| 100% inspection | £38,341 |
| Best feasible model policy | £272,856 |
Inspection itself removes roughly £294,000 per 1,000 wafers of escape cost. The model captures a small fraction of that.
At the feasible optimum (threshold 0.234): recall 0.18, precision 0.25, alarm rate 4.85%. The model catches 19 of 104 defects and misses 85.
Crossover analysis
| Cost per inspection | Model | 100% inspection | Winner |
|---|---|---|---|
| £40 | £272,856 | £38,341 | Inspect all |
| £200 | £278,676 | £187,722 | Inspect all |
| ~£290 | ~£281,000 | ~£281,000 | Crossover |
| £400 | £285,951 | £374,448 | Model |
| £1,600 | £327,128 | £1,494,805 | Model |
The structural finding
As inspection cost rises fortyfold, model cost rises only 20% — because at a 5% alarm rate it inspects fewer than 80 wafers. The model's cost floor is set by the 85 defects it misses, not by inspection.
Consequently: improving precision has near-zero value; only recall moves the number; and tuning cannot close a gap of this size.
Conclusion
Do not deploy as an automated gate. At current costs, 100% inspection is roughly seven times cheaper than any constrained model policy.
Deploy as triage only where capacity is genuinely limited. At a 5% inspection budget, random selection finds 6.6% of defects; model-guided selection finds 18% — a 2.8× improvement per inspection performed, independent of escape-cost assumptions.
Direct investment at recall — better sensors, better features — not at tuning.
Transferable screening rule
Inspection ML earns its place when inspection cost per part exceeds roughly 5–6% of escape cost, or when capacity prevents 100% coverage.Cheap, fast, non-destructive inspection → inspect everything; a model adds littleSlow, destructive, or capacity-limited inspection → a model may payRare defects with modest model lift → verify the arithmetic before committing
Applied to machine vision: automated visual inspection is cheap per part, placing it in "inspect everything" territory. A business case for vision should therefore not rest on displacing inspection labour. The defensible arguments are capacity (100% manual coverage is impossible at line speed), consistency (human inspection varies by inspector, shift and hour), coverage (destructive tests cannot be applied universally), and feedback latency (detecting drift in minutes rather than at end of shift).
Limitations
- Anonymised sensors — no physical interpretation possible, so domain-informed feature engineering could not be attempted. On real process data this would be the first avenue and could change the recall conclusion.
- Imputation outside the CV pipeline — mildly optimistic; should be re-run before acting.
- No temporal split — timestamps exist and semiconductor processes drift. Random CV may overstate performance on future wafers.
- No replicates — label noise cannot be estimated, so the achievable recall ceiling is unknown.
- Single model family — only a random forest was evaluated.
What this case study demonstrates
The analysis recommended against the application it was built to support. That conclusion took two days and prevented a pilot that would not have paid back.
It also produced something more durable than a model: a reusable screening rule that can be applied to any future inspection proposal before development begins.
An honest negative delivered in month one is worth considerably more than a flattering pilot that collapses in month eighteen.
5. Quick reference
Before accepting any classification result
- [ ] Majority-class baseline computed — is accuracy beating it meaningfully?
- [ ] Confusion matrix translated into business consequences
- [ ] Precision-recall curve requested, not ROC, if classes are imbalanced
- [ ] Stratified cross-validation used
- [ ] Recall expressed as an escape rate; precision as a false scrap rate
- [ ] Identifiers dropped before modelling
- [ ] Any resampling applied inside the CV fold only
Before recommending an operating point
- [ ] Cost matrix defined, with the source of each figure stated
- [ ] Threshold swept across its full range
- [ ] Sensitivity run on the least certain cost
- [ ] Alarm rate constrained to available inspection capacity
- [ ] Precision floor set to avoid alarm fatigue
- [ ] Do-nothing and do-everything anchors computed
- [ ] Optimum checked — is it on a boundary?
- [ ] Dominant cost mode identified, to direct investment
- [ ] Recommendation stated as an inspection policy, not a number