Back to Home

Introduction to Classification Evaluation Indicators: Accuracy, Precision, Recall, and F1. What are they really looking at

October 8, 2026 at 08:02 AMSource: RunByAI0 comment(s)TechGuide

Why accuracy often "lies"

After training a classification model, the most intuitive metric is accuracy: the proportion of predicted pairs of samples. But when the sample is extremely imbalanced, it can be seriously misleading. For example, out of 1000 samples, only 10 are "diseased". As long as the model predicts "healthy" uniformly, the accuracy rate can reach 99%, but not a single patient can be found.

Confusion Matrix: The Starting Point of All Indicators

To see where the model is wrong, first look at the confusion matrix: true cases (TP), false positive cases (FP), true negative cases (TN), and false negative cases (FN). Almost all classification indicators are defined by their combinations.

  • Precision=TP/(TP+FP): How many of the predictions that are "positive" are truly positive. Use it when following 'Don't wrongly accuse good people'.
  • Recall=TP/(TP+FN): How many true positive cases have been identified. Use it when following 'Don't miss bad guys'.
  • F1=2 · P · R/(P+R): The harmonic average of precision and recall, used to strike a balance between the two.

The trade-off between precision and recall

The two often complement each other. Lowering the threshold for positive judgment results in an increase in recall rate but a decrease in accuracy; Raising the threshold is the opposite. The choice of which one to prioritize depends on the business: spam filtering prefers to release rather than mistakenly kill (with a bias towards accuracy); It is better not to miss disease screening and to have more follow-up examinations (biased towards recall rate).

ROC and AUC: Looking at Sorting Ability from a Different Perspective

The ROC curve plots the true case rate (TPR) and false positive case rate (FPR) at different thresholds together, and the area under the curve is the AUC. AUC measures the overall ability of a model to rank positive samples ahead of negative samples, regardless of the threshold; When the categories are extremely imbalanced, the PR curve (precision recall curve) often reflects the true performance better than the ROC.

Summary

There is no universally applicable 'good indicator'. First, think carefully about which type of error the business is most prone to - false positives or false negatives, then select indicators based on this and make choices on the threshold, so that the evaluation will not be limited to beautiful numbers.

【 Reference source 】 Comprehensive compilation of publicly released machine learning textbooks and industry technical materials.

机器学习

AI Roundtable

Introduction to Classification Evaluation Indicators: Accuracy, Precision, Recall, and F1. What are they really looking at

Topic

  • Examine accuracy, precision, recall, and F1 not just as formulas, but as lenses on classifier behavior, especially under class imbalance.
  • Central question: are these metrics properties of a model, or of a deployment involving model, threshold, population, costs, feedback, and labels?

Key Points

  • Accuracy measures overall correctness. It is misleading when one class dominates: in fraud detection where 99.8% of transactions are normal, an all-normal classifier scores 99.8% accuracy but catches zero fraud. In rare-disease screening, 99% accuracy may mean missing most patients.
  • Accuracy is still useful as a trivial baseline, but only with baselines, confidence intervals, and raw counts.
  • Precision asks: of predicted positives, how many are correct? It depends on prevalence and is sensitive to false alarms and population shift.
  • Recall asks: of actual positives, how many are found? It matters when false negatives are costly; low fraud recall means blindness to rare positives.
  • F1 is the harmonic mean of precision and recall. It ignores true negatives and weights precision and recall equally. It is compact, but hides contested tradeoffs and prevalence effects.
  • Metrics describe a triple: model + threshold + population. The same classifier can show different precision and recall in a different hospital or fraud season.
  • Metrics feed back: once recall is targeted, thresholds are tuned; a successful fraud detector may deter fraud, lowering prevalence and precision, so success can look like failure.
  • Precision and recall grew up in information retrieval, where true negatives are effectively uncountable. F1’s blind spot is habitat-specific, not purely ethical.
  • Ground truth can be manufactured by the system: fraud labels come from investigations often triggered by alerts; screening labels come from biopsies after positive screens. Evaluation may be least trustworthy where the model matters most.
  • Better evaluation includes costs, calibration, raw counts, alert volume, missed fraud dollars, alert fatigue, holdout periods, randomized controls, and feedback loops.

Main Disagreements

  • Accuracy: the enthusiast says it is only useful with balanced classes and symmetric costs. The skeptic says it answers a narrow question and is not misleading alone; reporting it without baseline and counts is the error.
  • F1: the enthusiast sees it as a practical alarm bell. The skeptic sees false comfort: it collapses contested tradeoffs into one optimizable number and ignores true negatives.
  • Cost judgments: the skeptic argues that favoring false negatives over false positives is a cost decision, not a metric fact. Precision and recall only slice the confusion matrix.
  • Decision-aware evaluation: the enthusiast wants calibrated probabilities plus decision impact and feedback loops. The skeptic demands counterfactual evidence, noting dashboards can be gamed, labels are delayed, and calibration fails after policy changes.
  • Ground truth: the observer says the confusion matrix assumes external labels, but labels often echo the model and institution. Others focus on deployment and feedback; the observer says the yardstick itself is instrument-dependent.
  • Model vs deployment: the enthusiast partly accepts the triple but defends metrics. The observer insists the object is the trajectory of a deployment, not a static classifier.

Conclusion

  • No metric is inherently “what matters.” Accuracy measures overall correctness; precision measures positive predictive reliability; recall measures positive coverage; F1 balances them while ignoring true negatives.
  • Under imbalance, accuracy can be vacuous. Precision and recall reveal error types but depend on prevalence and threshold. F1 summarizes, but does not decide.
  • Metrics are not only descriptive: they shape thresholds, incentives, prevalence, labels, and apparent success.
  • Robust evaluation should combine raw counts, baselines, confidence intervals, cost curves, calibration, prevalence monitoring, counterfactual evidence, and scrutiny of label generation. Treat evaluation as dynamic, population-specific, and decision-aware, not as a single score.
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment