Research Report
Topic
- Examine accuracy, precision, recall, and F1 not just as formulas, but as lenses on classifier behavior, especially under class imbalance.
- Central question: are these metrics properties of a model, or of a deployment involving model, threshold, population, costs, feedback, and labels?
Key Points
- Accuracy measures overall correctness. It is misleading when one class dominates: in fraud detection where 99.8% of transactions are normal, an all-normal classifier scores 99.8% accuracy but catches zero fraud. In rare-disease screening, 99% accuracy may mean missing most patients.
- Accuracy is still useful as a trivial baseline, but only with baselines, confidence intervals, and raw counts.
- Precision asks: of predicted positives, how many are correct? It depends on prevalence and is sensitive to false alarms and population shift.
- Recall asks: of actual positives, how many are found? It matters when false negatives are costly; low fraud recall means blindness to rare positives.
- F1 is the harmonic mean of precision and recall. It ignores true negatives and weights precision and recall equally. It is compact, but hides contested tradeoffs and prevalence effects.
- Metrics describe a triple: model + threshold + population. The same classifier can show different precision and recall in a different hospital or fraud season.
- Metrics feed back: once recall is targeted, thresholds are tuned; a successful fraud detector may deter fraud, lowering prevalence and precision, so success can look like failure.
- Precision and recall grew up in information retrieval, where true negatives are effectively uncountable. F1’s blind spot is habitat-specific, not purely ethical.
- Ground truth can be manufactured by the system: fraud labels come from investigations often triggered by alerts; screening labels come from biopsies after positive screens. Evaluation may be least trustworthy where the model matters most.
- Better evaluation includes costs, calibration, raw counts, alert volume, missed fraud dollars, alert fatigue, holdout periods, randomized controls, and feedback loops.
Main Disagreements
- Accuracy: the enthusiast says it is only useful with balanced classes and symmetric costs. The skeptic says it answers a narrow question and is not misleading alone; reporting it without baseline and counts is the error.
- F1: the enthusiast sees it as a practical alarm bell. The skeptic sees false comfort: it collapses contested tradeoffs into one optimizable number and ignores true negatives.
- Cost judgments: the skeptic argues that favoring false negatives over false positives is a cost decision, not a metric fact. Precision and recall only slice the confusion matrix.
- Decision-aware evaluation: the enthusiast wants calibrated probabilities plus decision impact and feedback loops. The skeptic demands counterfactual evidence, noting dashboards can be gamed, labels are delayed, and calibration fails after policy changes.
- Ground truth: the observer says the confusion matrix assumes external labels, but labels often echo the model and institution. Others focus on deployment and feedback; the observer says the yardstick itself is instrument-dependent.
- Model vs deployment: the enthusiast partly accepts the triple but defends metrics. The observer insists the object is the trajectory of a deployment, not a static classifier.
Conclusion
- No metric is inherently “what matters.” Accuracy measures overall correctness; precision measures positive predictive reliability; recall measures positive coverage; F1 balances them while ignoring true negatives.
- Under imbalance, accuracy can be vacuous. Precision and recall reveal error types but depend on prevalence and threshold. F1 summarizes, but does not decide.
- Metrics are not only descriptive: they shape thresholds, incentives, prevalence, labels, and apparent success.
- Robust evaluation should combine raw counts, baselines, confidence intervals, cost curves, calibration, prevalence monitoring, counterfactual evidence, and scrutiny of label generation. Treat evaluation as dynamic, population-specific, and decision-aware, not as a single score.
Full Transcript
Welcome. Today we ask what accuracy, precision, recall, and F1 are really looking at. Not just formulas, but what each number reveals about a classifier's behavior, and what it hides. My role is to keep us neutral, so I will not favor any metric. Please ground claims in examples, especially where class imbalance changes the story. No jargon without explanation. We will move from accuracy to precision, recall, then F1, then open the floor. Let's begin with a simple question: when is high accuracy actually misleading?
High accuracy is misleading when a model merely learns the majority class. In fraud detection where 99.8 percent of transactions are normal, a classifier that always says normal gets 99.8 percent accuracy while catching zero fraud, so it hides exactly the behavior we care about. In medical screening for rare disease, 99 percent accuracy may mean missing nearly every patient. The lesson is that accuracy is only useful when classes are balanced and the cost of false positives and false negatives is symmetric. We should immediately ask which error matters more, which leads naturally to precision and recall.
I agree the fraud example shows accuracy can be vacuous, but I object to saying accuracy is only useful when classes are balanced and costs symmetric. Accuracy answers one narrow question: overall correctness. In a 99.8 percent normal dataset, 99.8 percent accuracy is not misleading by itself; it reveals the trivial baseline. The error is reporting it without baseline, confidence intervals, or raw counts. Also, deciding false negatives matter more is a cost judgment, not a metric fact. Precision and recall only slice the confusion matrix; they hide prevalence and the burden of false alarms. We should ask what decision costs justify favoring one slice.
Both of you treat a metric as a property of the model. I would say it is a property of a triple: model, threshold, and population. The same classifier moved to a different hospital or a different fraud season shows different precision and recall without changing at all, because prevalence shifted underneath it. These numbers describe a deployment, not a device. Second, nobody has mentioned that metrics act back on the system. Once recall becomes the stated target, thresholds get tuned to it, and the number stops describing behavior and starts steering it, like a thermometer that also heats the room. So perhaps the sharper question is not what each metric looks at, but when it stops observing and starts pushing.
Observer's triple is exactly the future of evaluation: we should report model, threshold, population, and cost curve together. But I reject that this makes accuracy useless or metrics manipulative. If a model has 99.8 percent accuracy only because it predicts normal, that is a powerful negative result: it proves the classifier adds no value over a constant. Precision and recall then show where the signal is hiding: if recall is 0.02 for fraud, the model is blind to rare positives. F1 punishes that blindness, so I still trust it as a practical alarm bell. The answer is not to distrust metrics; it is to make them decision-aware.
I partly accept the triple point, but I challenge the idea that F1 is a practical alarm bell. It is a compromise metric with hidden ethics: the harmonic mean weights precision and recall as equally important and ignores true negatives. In fraud, a low F1 may reflect many cheap alerts, while a moderate F1 can coexist with missing catastrophic cases. F1 also drifts with prevalence because precision depends on base rate. If the target becomes F1, teams can tune thresholds to flatter it without reducing harm. Decision-aware evaluation means costs, calibration, and raw counts, not another single score.
One thing neither side has touched: these metrics have biographies. Precision and recall grew up in information retrieval, where true negatives are effectively infinite, since you cannot count every document you did not retrieve. F1 ignores true negatives not as an ethical choice but because in its native habitat there was nothing to count. Transplant it into a clinic or a bank, where true negatives are real and often the overwhelming majority, and you inherit a blind spot that now looks like a deliberate design decision. So part of this disagreement is not about math but about habitat: we keep asking tools shaped by one ecology to testify about another. I would also correct my own earlier triple: it is not static, because the model feeds back into the population term. A fraud detector that actually works deters fraud, so prevalence falls, so precision falls; the metric punishes the system for succeeding, like a predator whose census numbers crash because it hunted too well. That suggests the real object of evaluation is not a snapshot of a classifier but the trajectory of a whole deployment. The question worth asking is whether any of these numbers can survive contact with their own success.
Observer's biography is not a funeral for metrics but a blueprint for better ones. Precision and recall can survive their own success if we stop treating prevalence as background noise and start modeling it as a response to deployment. In fraud, report alerts per thousand, missed fraud dollars, alert fatigue rate, and how they move when thresholds change. If fraud falls, precision may fall, but the cost curve should reward fewer prevented losses and less customer harm. F1 remains a compact alarm, especially early in development, but the future is calibrated probabilities plus decision impact. The metric should not just watch the classifier; it should watch the whole feedback loop.
I will not call that a blueprint yet. Alerts per thousand and missed fraud dollars sound concrete, but they inherit the same problems: alert volume can be gamed by threshold shifts, missed dollars depend on delayed and incomplete labels, and prevented losses are counterfactual. Without randomized controls or holdout periods, you cannot tell whether precision fell because fraud declined or because the detector became indiscriminate. Calibration is also not a refuge; probabilities calibrated on yesterday's prevalence can be wrong after policy changes. F1 as an early alarm is still false comfort, because it collapses contested tradeoffs into one optimizable number. The hard question is what counterfactual evidence validates the feedback loop, not which dashboard we prefer.
Skeptic asks what counterfactual evidence could validate the loop, but I want to question something beneath both positions: the ground truth itself. Every confusion matrix assumes labels arrive from outside the system. In fraud, labels come from investigations, and investigations are opened mostly by alerts, often from the very model being judged. In screening, the gold standard biopsy tends to go to those who screened positive. So the true row of the matrix is partly manufactured by the attention of the institution doing the measuring. Fisheries science lives with this: catch per hour is read as fish abundance, but it confounds the fish with the fishermen. Maps of known ruins are largely maps of where archaeologists dug. We have covered metrics steering behavior and success eroding prevalence; neither of you has noted that the yardstick itself can be an echo of the instrument. The unsettling consequence is that evaluation may be least trustworthy exactly where the model matters most, because there labels, prevalence, and thresholds all move together.