Regression algorithms are tools used to predict continuous numerical values, like as prices, temperatures, or sales numbers, by learning the mathematical relationship between features and targets.
A model that appears to work is easy to build. Most of the effort in credible machine learning goes into establishing whether the appearance is real. The person best placed to catch a flattering result is usually not the person who produced it, which is why the audit checklist at the end of this document matters more than any of the code.
Every technique in this post exists to stop you from fooling yourself.
1. A single train/test split is one draw from a distribution
Splitting once and reporting the result is reporting a sample of size one.
Running the same model on the same data with twenty different random seeds produces twenty different scores. The spread is frequently 20–40% of the mean on small datasets. Nothing changed except which rows happened to land in the test set.
Consequences:
- Any single split result is unreproducible by anyone who chooses a different seed
- If the analyst tried several splits before reporting one, it is likely the most flattering
- Differences between models measured on one split are usually noise
A single split is acceptable only for a quick sanity check, never for a reported result.
2. Cross-validation, and always report the spread
K-fold cross-validation divides the data into k parts, then trains k separate models, each tested on the part it never saw. Every row is used for testing exactly once; no model is ever tested on data it trained on.
cv = KFold(n_splits=5, shuffle=True, random_state=42)
scores = -cross_val_score(model, X, y, cv=cv,
scoring="neg_mean_absolute_error")
print(scores) # all folds, always
print(scores.mean(), scores.std())
(Scikit-learn negates error metrics because its convention is that higher is better. The leading minus sign restores normal, positive error.)
Report mean ± standard deviation. Never the mean alone.
The standard deviation is the more informative half:
- It measures how stable performance is across different training sets
- It defines the significance threshold for model comparison. If the fold-to-fold std is 0.002, a model beating another by 0.001 has not beaten it.
- A single anomalous fold reveals structure the mean conceals, a cluster of unusual samples, or a region of the design space the model cannot handle
Declaring a winner from a gap smaller than the fold-to-fold spread is among the most common errors in model comparison.
3. The split strategy must match the data's structure
Random splitting assumes rows are independent. In manufacturing they frequently are not.
The failure mechanism: ten parts made from one batch of raw material share that batch's characteristics. Random splitting puts seven in training and three in test. The model learns the batch's quirk and is rewarded for recognising it, which looks like understanding the process but is partly memorisation. When a new batch arrives, performance collapses.
| Split | Groups by | Use when |
|---|---|---|
Random (KFold) |
Nothing | Rows genuinely independent |
Grouped (GroupKFold) |
Batch, machine, operator, campaign | Rows share a hidden common cause |
| Temporal | Time | Predicting the future; drift possible |
Random splits on temporal data let the model learn from the future. This single error produces a large share of pilots that look excellent in the notebook and fail on the line.
The question to ask of any dataset: what do rows share that is not in the feature table?
4. Pipelines make leakage structurally impossible
Any preprocessing step that learns from data, scaling, imputation, feature selection, target encoding, must be fitted inside each cross-validation fold, using only that fold's training data.
# Wrong: scaler sees the entire dataset before splitting
X_scaled = StandardScaler().fit_transform(X)
scores = cross_val_score(Ridge(), X_scaled, y, cv=cv)
# Right: scaler refits within every fold
pipe = make_pipeline(StandardScaler(), Ridge())
scores = cross_val_score(pipe, X, y, cv=cv)
With scaling the optimism is small, only the mean and standard deviation leak. With imputation or feature selection it can be severe.
The value of a pipeline is that it removes the possibility of the mistake, not that it is tidier. Once preprocessing is inside a pipeline, this class of error cannot occur.
5. Leakage that no tool can detect
The dangerous form of leakage involves legitimate columns, correct code, and properly executed cross-validation.
The test: at the exact moment the prediction is needed, will this value exist?
A model predicting a property to decide what to manufacture has access to design variables, the things being chosen. It does not have access to anything measured on the finished article, even though those columns sit in the same table and are strongly predictive.
Adding a measured intermediate improves every metric and destroys the model's purpose: obtaining that input requires making the sample, at which point the property could simply be measured.
No automated check catches this. The columns are valid, the code is correct, the validation is sound. Only someone who understands when each value becomes available can ask the question, which makes it a domain expert's responsibility, not a data scientist's.
Related, and easier to spot: columns recorded downstream of the outcome, rework hours, final inspection code, disposition. These produce near-perfect scores. When a result looks too good, go and find the column.
6. Hyperparameter tuning buys less than you think
Grid search over a reasonable parameter space typically improves performance by a few percent.
For comparison, in the same problem:
- Physics-informed feature engineering: often tens of percent
- Correct choice of operating point: can determine whether the project pays back at all
- Fixing a leakage or split error: changes whether the result is real
Keep this ratio in mind when someone requests three weeks for tuning. Tuning is the last few percent, applied after the framing is right. It is rarely where a project is won.
7. Compositional data: predictions are unique, coefficients are not
When ingredient fractions sum to a constant, the variables are perfectly collinear, knowing all but one determines the last. One component must be dropped as a reference.
Dropping different components gives identical predictions and completely different coefficients.
This is not a bug; it reflects a real ambiguity. "Increasing component A by 1% raises the property by X" is meaningless without stating what decreases to make room. Drop the filler and the coefficient means "A replacing filler". Drop a different component and it means something else entirely: different physics, different number, same model.
Practical rules:
- Choose as the dropped reference the component you would actually adjust to balance a recipe, usually the filler or balance
- State the reference explicitly wherever coefficients appear
- The same ambiguity applies to SHAP values and is considerably less obvious there
- Watch for implicit closure: if a component is missing from the table, the remaining fractions will not sum to a constant and the problem appears absent when it is not
8. Report numbers to the precision you have
If fold-to-fold standard deviation is 14% of the mean, reporting "26.83698% improvement" implies precision that does not exist. Report 27%.
Spurious decimal places are a reliable signal that uncertainty has not been considered.
The Evaluation Checklist
Eight questions to ask of any model presented to you.
Each question demands a consequence, not a metric. Consequences are much harder to hand-wave than numbers, which is what makes these difficult to deflect.
1. What is the baseline, and by how much do we beat it?
Good answer: "Baseline is predicting the mean for every sample, MAE 0.0194. The model achieves 0.0142. A 27% reduction in error."
Evasive answer: "R² is 0.83." That is not a baseline comparison.
Why it matters: R² and accuracy are easily impressive. Improvement over the dumbest possible approach is interpretable by a business audience and far harder to inflate. For classification, the baseline is predicting the majority class, which on a 93/7 split yields 93% accuracy and zero value.
2. How was the data split: randomly, or by time, batch, or machine?
Good answer: "Grouped 5-fold by extraction campaign. Panels within a campaign share a raw material batch, so a random split would let batch characteristics leak across folds."
Evasive answer: "Standard 80/20 split." No mention of what rows might share.
Why it matters: the most common cause of a pilot that succeeds in the notebook and fails on the line. Ask what rows share that is not in the feature table.
3. Show me the variance across folds, not the mean.
Good answer: "Folds: [0.0114, 0.0166, 0.0124, 0.0160, 0.0148]. Mean 0.0142, std 0.0020 or 14% of the mean. Differences between models smaller than ~0.002 are not meaningful."
Evasive answer: a single mean, or "cross-validated" with no numbers.
Why it matters: the spread sets the significance threshold for every comparison that follows. It also exposes anomalous folds, which indicate structure the mean conceals.
4. Which features would not exist at prediction time?
Good answer: "Excluded density and extraction yield, both measured on the finished panel, so neither exists at design time. Including density improves MAE by 18% and makes the model unusable for its purpose."
Evasive answer: "We checked for leakage." Usually means they checked for target-derived columns, not for timing.
Why it matters: no tool detects this. It requires knowing when each value becomes available, which is domain knowledge. This question is the domain expert's to ask.
5. What error, in physical units, and is it inside measurement uncertainty?
Good answer: "MAE 0.0142 NRC. Reporting resolution is 0.05, so error is well inside it. True repeatability is unknown, no replicates exist. Recommend 5–10 repeat measurements next campaign to establish the noise floor."
Evasive answer: "The model is 95% accurate."
Why it matters: physical units let a domain expert judge usefulness immediately. The measurement-noise comparison is what tells you when to stop, if model error is already below test repeatability, further improvement is chasing noise, and the path forward is better measurement rather than better algorithms.
Note the distinction: reporting conventions (rounding to 0.05) are an upper bound on what is meaningful, not a measured uncertainty. Only replicates establish the real figure.
6. What happens when the model is wrong, and who notices?
Good answer: "Failure costs a wasted lab trial, roughly £X and Y days, no customer impact. Asymmetric: over-prediction is self-correcting, since the formulation gets made and measured. Under-prediction is silent, a good formulation never gets made and is never disproven. Detection: every proposal is manufactured and measured, giving predicted-vs-actual on each new sample; running error is tracked against the cross-validated baseline. Owner: [named person], reviewed each campaign."
Evasive answer: "We'll monitor performance."
Why it matters: three things must be present, the cost of an error, the asymmetry between error directions, and a mechanism rather than a hope. "Someone will notice" is not a detection plan. Silent failures are the dangerous ones: in an optimisation loop, systematic under-prediction in a region quietly deletes that region from the search and leaves no evidence.
Being explicit that stakes are low is equally valuable, it prevents over-engineering governance for a screening tool.
7. What is the gap between your best two models, relative to the fold-to-fold spread?
Good answer: "Gradient boosting beats Ridge by 0.0008; fold std is 0.0020. The difference is not significant. We selected Ridge because it is simpler, faster to retrain, and its coefficients are interpretable to the formulation team."
Evasive answer: "Gradient boosting performed best," with no spread quoted.
Why it matters: model comparison tables invite false precision. When differences fall inside noise, the tie-breaker should be an engineering criterion, interpretability, training cost, maintainability, ease of explanation to the people who must act on it, not a third decimal place.
8. What is the range of validity, and what happens outside it?
Good answer: "The model interpolates within the observed design space: seaweed 0.25–0.33, temperature 70–95 °C, and so on. Outside those ranges it extrapolates and should not be trusted. Optimisation is bounded to the observed envelope, and any proposal near a boundary is flagged for review."
Evasive answer: "It generalises well."
Why it matters: the highest-value formulations frequently sit outside the region already explored, so this is exactly where an optimiser will try to go. Tree-based models are particularly hazardous here, they predict a constant beyond the training range, so an extrapolated prediction looks confident and is meaningless.
Optional ninth, for anything approaching deployment
Can you reproduce this result from a clean checkout, and who owns the model?
Good answer: "Yes, data version tagged in DVC, environment pinned, seeds fixed, run logged in MLflow. [Named person] owns it, with a budget line for retraining."
Evasive answer: "The notebook is on my laptop."
Why it matters: models do not fail so much as get orphaned. An unowned model drifts, is quietly ignored within a year, and burns the organisation's appetite for the next attempt.
The pattern underneath the checklist
Every question converts a number into a consequence:
- 0.0142 → and differences below 0.002 are not real
- The error is small → and we cannot say how small until we have replicates
- Wrong predictions waste trials → asymmetrically, and here is who checks
That conversion is the difference between reporting a result and owning one. It is also why these questions are hard to deflect: a metric can be quoted, but a consequence has to be reasoned about.