1. Classification predicts categories rather than continuous quantities. Regression asks “what strength?”, while classification can ask “PASS or FAIL?”
2. The class definition is an engineering decision. Our 110 kPa threshold created the PASS/FAIL labels. In a real project, that threshold needs a justified specification or decision basis.
3. Always inspect class balance. High accuracy can be meaningless when one class dominates the dataset.
4. Classification also needs a baseline. A most_frequent classifier showed what could be achieved without learning useful relationships.
5. Accuracy alone does not describe classifier quality. The baseline and Logistic Regression both achieved 57.1% accuracy but made very different mistakes.
6. The confusion matrix tells you what kind of mistakes the model makes. TP, TN, FP and FN translate model performance into concrete engineering outcomes.
7. False positives and false negatives usually have different consequences. For R&D screening, testing one poor formulation may be cheap compared with permanently overlooking a valuable formulation.
8. Choose metrics based on those consequences. Because false negatives were particularly undesirable in our screening scenario, recall became especially important.
9. Precision answers a different question from recall. Recall asks whether we found the promising formulations; precision asks how many of our recommended formulations were actually promising.
10. Perfect recall can still describe a useless model. Predicting PASS for everything gives 100% recall but saves no experiments.
11. Specificity measures the opposite side of the screening problem. It asks how successfully we identify and eliminate actual FAIL cases.
12. F1 combines precision and recall, but it does not understand business consequences. A numerical metric cannot determine whether missing an opportunity is ten times more expensive than running an unnecessary experiment.
13. Classifiers often produce scores or probabilities before producing categories. PASS/FAIL is created by applying a decision threshold to that score.
14. A 0.50 threshold is a default, not a law. The operating threshold should reflect the decision the model supports.
15. Lowering a threshold usually increases recall at the expense of more false positives. Raising it usually does the reverse—but only if the predicted scores actually separate the classes.
16. Thresholds only change predictions when they cross an observation's score. That is why thresholds 0.30–0.60 initially produced exactly the same results.
17. Threshold tuning cannot repair fundamentally incorrect ranking. Two FAIL specimens received the highest PASS probabilities in our test set; changing the threshold could not distinguish them from genuinely strong candidates.
18. Model selection should favour simplicity when predictive performance is equivalent. Logistic Regression and Random Forest produced identical results, so Logistic Regression was the more sensible candidate for continued experimentation.
19. Prediction and decision-making are separate layers. The model provides a score; the engineering policy determines what action to take with that score.
20. Explicit error costs can turn ML metrics into a decision model. We represented the R&D consequence as:
Decision cost = FP × cost(FP) + FN × cost(FN)
21. Cost-based evaluation can reveal value that accuracy hides. At an illustrative 0.25 threshold, Logistic Regression retained all promising candidates while eliminating one unnecessary experiment, outperforming the “test everything” baseline under all three cost scenarios examined.
22. Threshold selection must also be validated out-of-sample. Choosing the best threshold from the final test set is another form of evaluation overfitting.
And the central principle:
The best classifier is not necessarily the one that predicts the most cases correctly. It is the one whose pattern of errors best supports the real engineering decision at an acceptable cost and risk.
That completes Week 6 very nicely. We are now ready for Week 7, where we'll make the validation framework considerably stronger with cross-validation, repeated evaluation, and more rigorous model/threshold selection.
Classification Glossary
| Term | Definition |
|---|---|
| Classification | A machine-learning task where the target is a category, such as PASS/FAIL, rather than a continuous number. |
| Positive class | The class defined as the outcome of interest. In our exercise, PASS = 1. |
| Negative class | The opposite class. In our exercise, FAIL = 0. |
| True Positive (TP) | The model predicts PASS and the specimen actually passes. |
| True Negative (TN) | The model predicts FAIL and the specimen actually fails. |
| False Positive (FP) | The model predicts PASS but the specimen actually fails. In R&D screening, this can lead to unnecessary testing. |
| False Negative (FN) | The model predicts FAIL but the specimen actually passes. In R&D screening, this can mean discarding a potentially valuable formulation. |
| Confusion matrix | A table showing the counts of TP, TN, FP and FN. It reveals the types of mistakes the classifier makes. |
| Accuracy | Fraction of all predictions that are correct: (TP + TN) / total. Useful, but can be misleading when classes are imbalanced or error costs differ. |
| Precision | Of everything predicted as PASS, the proportion that actually passed: TP / (TP + FP). High precision means few false positives. |
| Recall | Of everything that actually passed, the proportion correctly identified: TP / (TP + FN). High recall means few false negatives. Also called sensitivity. |
| Specificity | Of everything that actually failed, the proportion correctly identified as FAIL: TN / (TN + FP). |
| F1 score | A combined measure of precision and recall. Useful when both false positives and false negatives matter. |
| Class balance | The relative number of observations belonging to each class. Strong imbalance can make accuracy misleading. |
| Baseline classifier | A deliberately simple reference model, such as always predicting the most common class. A useful model should justify itself against this baseline. |
| Predicted probability | A model score between 0 and 1 representing how strongly the model associates an observation with the positive class. It should not automatically be treated as a perfectly calibrated real-world probability. |
| Decision threshold | The cutoff used to convert a predicted probability into a class. For example, probability ≥ 0.50 → PASS. |
| Threshold tuning | Adjusting the decision threshold to balance false positives and false negatives according to the engineering objective. |
| Decision cost | A way to translate classification errors into economic or engineering consequences, e.g. FP × cost(FP) + FN × cost(FN). |
| Logistic Regression | A relatively simple classification model that estimates the likelihood of belonging to a class and is often useful as an interpretable benchmark. |
| Random Forest Classifier | An ensemble of decision trees that can capture nonlinear relationships and interactions between variables. |
| Stratification | Splitting data while trying to preserve the proportion of classes in each subset. |
| Group-aware splitting | Splitting data so related observations, such as specimens from the same formulation or batch, do not appear in both training and evaluation sets. |
| Validation set | Data used during model or threshold selection. It should be separate from the final test set. |
| Test set | Unseen data reserved for the final evaluation of a model and decision rule. It should not be repeatedly used for tuning. |
A compact way to remember the four confusion-matrix outcomes is:
TP: good material correctly kept
TN: poor material correctly rejected
FP: poor material unnecessarily kept
FN: good material incorrectly rejected
And for your R&D screening use case, the most important pair is:
Recall = “How many promising formulations did we find?”
Precision = “How many of the formulations we chose to test were actually promising?”