Precision, recall, and the accuracy trap: evaluate classifiers with the right denominator
Work through a complete confusion matrix, see how prevalence changes precision, and choose thresholds using error costs rather than a headline accuracy score.
Reviewed October 1, 2026. Every numerical example is synthetic and calculated from the stated counts. None represents a deployed product's performance. The Python example uses scikit-learn metric functions.

Suppose a classifier is 98% accurate. That sounds strong until you learn that only 2% of cases belong to the class you want to detect. A system that predicts “negative” for everyone achieves that accuracy while finding none of the positive cases.
The problem is not that accuracy is mathematically wrong. It answers a particular question: what fraction of all decisions were correct? When one class dominates or error costs differ, that question may not match the decision you need to make.
First define the positive class
“Positive” does not mean good. It is the event selected for measurement: a spam message, a defective item, or a document relevant to a query. Reversing that choice changes the interpretation of precision and recall.
Google's classification tutorial defines the confusion-matrix outcomes. A true positive is an actual positive flagged positive. A false positive is an actual negative flagged positive. A false negative is a missed positive. A true negative is a negative correctly left unflagged.
For this example, “positive” means a message requiring manual review. Our synthetic test set contains 1,000 messages, of which 20 actually require review. The system flags 48: 16 correctly and 32 unnecessarily.
| Actual class | Predicted positive | Predicted negative | Total |
|---|---|---|---|
| Positive | 16 true positives (TP) | 4 false negatives (FN) | 20 |
| Negative | 32 false positives (FP) | 948 true negatives (TN) | 980 |
| Total | 48 | 952 | 1,000 |
Read each metric as a question
| Metric | Question / formula | Example result |
|---|---|---|
| Accuracy | How many decisions were correct? (TP + TN) / N | 964 / 1,000 = 96.4% |
| Precision | Of the flagged cases, how many were positive? TP / (TP + FP) | 16 / 48 = 33.3% |
| Recall | Of all actual positives, how many did we find? TP / (TP + FN) | 16 / 20 = 80.0% |
| Specificity | Of all actual negatives, how many stayed negative? TN / (TN + FP) | 948 / 980 = 96.7% |
| F1 | Harmonic mean of precision and recall: 2TP / (2TP + FP + FN) | 32 / 68 = 47.1% |
| Balanced accuracy | Average of recall and specificity in the binary case. | (0.8 + 948 / 980) / 2 = 88.4% |
The same system can have high accuracy, useful recall, and poor precision. It catches most cases requiring review, but two-thirds of the review queue are false alarms. Whether that trade-off is acceptable depends on review capacity and the consequences of a miss.
An always-negative baseline reaches 98% accuracy here, compared with the model's 96.4%, but has zero recall. This does not prove the model should be deployed; it proves that accuracy alone cannot select the right system.
Precision depends on how common the event is
Christopher K. I. Williams's analysis explains the relationship between class prevalence and precision. Even if a classifier's true-positive and false-positive rates remain fixed, its precision changes when the class mixture changes.
Consider a separate synthetic system with 80% recall and a 2% false-positive rate. In two populations of 10,000 cases, assume those rates stay unchanged:
| Positive prevalence | True positives | False positives | Precision |
|---|---|---|---|
| 10%: 1,000 positives | 800 | 2% of 9,000 = 180 | 800 / 980 = 81.6% |
| 1%: 100 positives | 80 | 2% of 9,900 = 198 | 80 / 278 = 28.8% |
The false-positive rate is identical, but the rare-event population contains far fewer true positives to compete with false alarms. A precision score from an artificially balanced test set can therefore misrepresent a real review queue. In practice, rates can also change because the population itself changes, so the fixed-rate assumption must not be taken for granted.
A threshold turns scores into actions
A classifier often produces a score or probability estimate before assigning a label. Scikit-learn's threshold guide separates estimating that score from choosing an action. A default threshold is a convention, not a statement about your error costs.
Produce a score→Threshold
Choose a cutoff→Decision
Flag or leave→Outcome
Measure the cost
For a fixed set of scores, lowering the threshold includes more cases, so recall cannot decrease. Precision need not change monotonically: the added cases may be relatively better or worse. Inspect the actual precision–recall curve instead of treating a trade-off slogan as a guarantee.
Choose the threshold using held-out validation or an appropriate cross-validation procedure. Do not tune it on the final test set and then report that same test as untouched evidence. If one missed case is much more costly than a review, evaluate a high-recall operating point; if the review queue has a hard limit, measure precision and recall within that capacity.
F1 is useful, but not a universal objective
Scikit-learn's metric guide documents F1, balanced accuracy, and different averaging schemes. F1 summarizes positive-class precision and recall, but omits true negatives. It also does not encode the real cost of a false alarm or missed event.
For multiclass tasks, macro averaging gives each class equal weight, while weighted averaging uses class support. A strong weighted score can coexist with a weak minority class. Always show the per-class results and counts before choosing an aggregate that matches the question.
Ranking quality is not the same as operating-point quality
ROC curves compare true-positive and false-positive rates across thresholds. Precision–recall curves compare positive prediction quality and coverage. Saito and Rehmsmeier's study demonstrates why PR plots can be more informative for strongly imbalanced binary problems.
That is not a rule that ROC is always useless. Use the curve that reveals the relevant behavior, and examine the region where your system will actually operate. A high area under a curve does not guarantee that the chosen threshold meets a review budget.
Also label the summary precisely. Average precision and trapezoidal PR area are not interchangeable calculations. Do not compare two numbers called “PR AUC” until you know how each was computed.
Check probability calibration separately
A model can rank cases well while assigning misleading probabilities. Calibration documentation describes comparing predicted probabilities with observed positive frequencies. If cases assigned about 0.8 are positive much less often, the score should not be interpreted as a reliable 80% probability.
Calibration data must be appropriately separated from training. Reliability diagrams, probability-sensitive scores, and enough cases per bin help assess this dimension. Calibration is particularly important when decisions use estimated probabilities rather than just a ranking.
Reproduce the first table in Python
from sklearn.metrics import (
accuracy_score, balanced_accuracy_score, confusion_matrix,
f1_score, precision_score, recall_score,
)
# Synthetic labels: 16 TP, 4 FN, 32 FP, and 948 TN.
y_true = [1] * 20 + [0] * 980
y_pred = [1] * 16 + [0] * 4 + [1] * 32 + [0] * 948
print(confusion_matrix(y_true, y_pred, labels=[0, 1]))
metrics = {
"accuracy": accuracy_score(y_true, y_pred),
"precision": precision_score(y_true, y_pred, zero_division=0),
"recall": recall_score(y_true, y_pred, zero_division=0),
"f1": f1_score(y_true, y_pred, zero_division=0),
"balanced_accuracy": balanced_accuracy_score(y_true, y_pred),
}
for name, value in metrics.items():
print(f"{name}: {value:.4f}")
Expected output is [[948, 32], [4, 16]], then 0.9640, 0.3333, 0.8000, 0.4706, 0.8837. Note the label order: scikit-learn puts actual classes on rows and predictions on columns. zero_division=0 is an explicit reporting convention when a denominator is zero; explain such conventions in real reports.
What a responsible evaluation report includes
- The positive-class definition, test population, date, and class counts.
- The split strategy and which dataset selected the threshold.
- A confusion matrix, baseline, per-class metrics, and operating threshold.
- Review volume, error consequences, and relevant subgroup performance.
- Uncertainty or limitations, especially with few positive cases.
With only 20 positives, one additional caught case changes recall by five percentage points. Do not present small changes as strong evidence without enough data. A good metric report makes the denominator, assumptions, and decision visible—not merely the best-looking percentage.
Sources checked October 1, 2026. Counts and numerical comparisons are original synthetic examples.
- Google Machine Learning Crash Course: Accuracy, precision, and recall
- scikit-learn: Metrics and scoring
- scikit-learn: Tuning the decision threshold
- scikit-learn: Probability calibration
- Saito & Rehmsmeier (2015): Precision–recall plots for imbalanced datasets, PLOS ONE
- Williams (2021): The Effect of Class Imbalance on Precision–Recall Curves