NOETRION

Browse by topic

← All articles
STATISTICS · 10 MIN READ

Precision, recall, and the accuracy trap: evaluate classifiers with the right denominator

Work through a complete confusion matrix, see how prevalence changes precision, and choose thresholds using error costs rather than a headline accuracy score.

Reviewed October 1, 2026. Every numerical example is synthetic and calculated from the stated counts. None represents a deployed product's performance. The Python example uses scikit-learn metric functions.

Conceptual illustration of a classifier sorting mixed items with a magnifier and a review tray.
Finding more relevant cases and making fewer false alarms are different goals. Exact counts appear in the tables, not in this conceptual illustration.

Suppose a classifier is 98% accurate. That sounds strong until you learn that only 2% of cases belong to the class you want to detect. A system that predicts “negative” for everyone achieves that accuracy while finding none of the positive cases.

The problem is not that accuracy is mathematically wrong. It answers a particular question: what fraction of all decisions were correct? When one class dominates or error costs differ, that question may not match the decision you need to make.

First define the positive class

“Positive” does not mean good. It is the event selected for measurement: a spam message, a defective item, or a document relevant to a query. Reversing that choice changes the interpretation of precision and recall.

Google's classification tutorial defines the confusion-matrix outcomes. A true positive is an actual positive flagged positive. A false positive is an actual negative flagged positive. A false negative is a missed positive. A true negative is a negative correctly left unflagged.

For this example, “positive” means a message requiring manual review. Our synthetic test set contains 1,000 messages, of which 20 actually require review. The system flags 48: 16 correctly and 32 unnecessarily.

Synthetic confusion matrix: rows are actual labels, columns are predictions
Actual classPredicted positivePredicted negativeTotal
Positive16 true positives (TP)4 false negatives (FN)20
Negative32 false positives (FP)948 true negatives (TN)980
Total489521,000

Read each metric as a question

Calculated results at this one operating point
MetricQuestion / formulaExample result
AccuracyHow many decisions were correct? (TP + TN) / N964 / 1,000 = 96.4%
PrecisionOf the flagged cases, how many were positive? TP / (TP + FP)16 / 48 = 33.3%
RecallOf all actual positives, how many did we find? TP / (TP + FN)16 / 20 = 80.0%
SpecificityOf all actual negatives, how many stayed negative? TN / (TN + FP)948 / 980 = 96.7%
F1Harmonic mean of precision and recall: 2TP / (2TP + FP + FN)32 / 68 = 47.1%
Balanced accuracyAverage of recall and specificity in the binary case.(0.8 + 948 / 980) / 2 = 88.4%

The same system can have high accuracy, useful recall, and poor precision. It catches most cases requiring review, but two-thirds of the review queue are false alarms. Whether that trade-off is acceptable depends on review capacity and the consequences of a miss.

An always-negative baseline reaches 98% accuracy here, compared with the model's 96.4%, but has zero recall. This does not prove the model should be deployed; it proves that accuracy alone cannot select the right system.

Precision depends on how common the event is

Christopher K. I. Williams's analysis explains the relationship between class prevalence and precision. Even if a classifier's true-positive and false-positive rates remain fixed, its precision changes when the class mixture changes.

Consider a separate synthetic system with 80% recall and a 2% false-positive rate. In two populations of 10,000 cases, assume those rates stay unchanged:

A controlled prevalence example, not a prediction of real-world transfer
Positive prevalenceTrue positivesFalse positivesPrecision
10%: 1,000 positives8002% of 9,000 = 180800 / 980 = 81.6%
1%: 100 positives802% of 9,900 = 19880 / 278 = 28.8%

The false-positive rate is identical, but the rare-event population contains far fewer true positives to compete with false alarms. A precision score from an artificially balanced test set can therefore misrepresent a real review queue. In practice, rates can also change because the population itself changes, so the fixed-rate assumption must not be taken for granted.

A threshold turns scores into actions

A classifier often produces a score or probability estimate before assigning a label. Scikit-learn's threshold guide separates estimating that score from choosing an action. A default threshold is a convention, not a statement about your error costs.

Model
Produce a score
→Threshold
Choose a cutoff
→Decision
Flag or leave
→Outcome
Measure the cost
The same trained model can produce different decisions under different thresholds. Changing the cutoff does not retrain its ranking function.

For a fixed set of scores, lowering the threshold includes more cases, so recall cannot decrease. Precision need not change monotonically: the added cases may be relatively better or worse. Inspect the actual precision–recall curve instead of treating a trade-off slogan as a guarantee.

Choose the threshold using held-out validation or an appropriate cross-validation procedure. Do not tune it on the final test set and then report that same test as untouched evidence. If one missed case is much more costly than a review, evaluate a high-recall operating point; if the review queue has a hard limit, measure precision and recall within that capacity.

F1 is useful, but not a universal objective

Scikit-learn's metric guide documents F1, balanced accuracy, and different averaging schemes. F1 summarizes positive-class precision and recall, but omits true negatives. It also does not encode the real cost of a false alarm or missed event.

For multiclass tasks, macro averaging gives each class equal weight, while weighted averaging uses class support. A strong weighted score can coexist with a weak minority class. Always show the per-class results and counts before choosing an aggregate that matches the question.

Ranking quality is not the same as operating-point quality

ROC curves compare true-positive and false-positive rates across thresholds. Precision–recall curves compare positive prediction quality and coverage. Saito and Rehmsmeier's study demonstrates why PR plots can be more informative for strongly imbalanced binary problems.

That is not a rule that ROC is always useless. Use the curve that reveals the relevant behavior, and examine the region where your system will actually operate. A high area under a curve does not guarantee that the chosen threshold meets a review budget.

Also label the summary precisely. Average precision and trapezoidal PR area are not interchangeable calculations. Do not compare two numbers called “PR AUC” until you know how each was computed.

Check probability calibration separately

A model can rank cases well while assigning misleading probabilities. Calibration documentation describes comparing predicted probabilities with observed positive frequencies. If cases assigned about 0.8 are positive much less often, the score should not be interpreted as a reliable 80% probability.

Calibration data must be appropriately separated from training. Reliability diagrams, probability-sensitive scores, and enough cases per bin help assess this dimension. Calibration is particularly important when decisions use estimated probabilities rather than just a ranking.

Reproduce the first table in Python

from sklearn.metrics import (
    accuracy_score, balanced_accuracy_score, confusion_matrix,
    f1_score, precision_score, recall_score,
)

# Synthetic labels: 16 TP, 4 FN, 32 FP, and 948 TN.
y_true = [1] * 20 + [0] * 980
y_pred = [1] * 16 + [0] * 4 + [1] * 32 + [0] * 948

print(confusion_matrix(y_true, y_pred, labels=[0, 1]))
metrics = {
    "accuracy": accuracy_score(y_true, y_pred),
    "precision": precision_score(y_true, y_pred, zero_division=0),
    "recall": recall_score(y_true, y_pred, zero_division=0),
    "f1": f1_score(y_true, y_pred, zero_division=0),
    "balanced_accuracy": balanced_accuracy_score(y_true, y_pred),
}
for name, value in metrics.items():
    print(f"{name}: {value:.4f}")

Expected output is [[948, 32], [4, 16]], then 0.9640, 0.3333, 0.8000, 0.4706, 0.8837. Note the label order: scikit-learn puts actual classes on rows and predictions on columns. zero_division=0 is an explicit reporting convention when a denominator is zero; explain such conventions in real reports.

What a responsible evaluation report includes

  • The positive-class definition, test population, date, and class counts.
  • The split strategy and which dataset selected the threshold.
  • A confusion matrix, baseline, per-class metrics, and operating threshold.
  • Review volume, error consequences, and relevant subgroup performance.
  • Uncertainty or limitations, especially with few positive cases.

With only 20 positives, one additional caught case changes recall by five percentage points. Do not present small changes as strong evidence without enough data. A good metric report makes the denominator, assumptions, and decision visible—not merely the best-looking percentage.