研讨报告
主题
- 考察准确率、精确率、召回率和 F1,不仅将它们视为公式,更将它们视为观察分类器行为的透镜,尤其是在类别不平衡的情况下。
- 核心问题:这些指标是模型的属性,还是涉及模型、阈值、人群、成本、反馈和标签的部署的属性?
要点
- 准确率衡量总体正确性。当一个类别占主导时,它会具有误导性:在欺诈检测中,99.8% 的交易是正常的,一个全判为正常的分类器得分是 99.8% 准确率,但捕获的欺诈为零。在罕见病筛查中,99% 的准确率可能意味着漏掉大多数患者。
- 准确率作为平凡基线仍有用,但必须辅以基线、置信区间和原始计数。
- 精确率问的是:在被预测为正的样本中,有多少是正确的?它取决于流行率,并且对误报和人群偏移敏感。
- 召回率问的是:在实际为正的样本中,有多少被找到?当假阴性代价高昂时,这很重要;欺诈召回率低意味着对罕见正类视而不见。
- F1 是精确率和召回率的调和平均数。它忽略真阴性,并对精确率和召回率同等加权。它很简洁,但掩盖了有争议的权衡和流行率影响。
- 指标描述的是一个三元组:模型 + 阈值 + 人群。同一个分类器在不同的医院或欺诈季节可能表现出不同的精确率和召回率。
- 指标会反馈:一旦以召回率为目标,阈值就会被调整;一个成功的欺诈检测器可能会阻止欺诈,降低流行率和精确率,于是成功可能看起来像失败。
- 精确率和召回率成长于信息检索领域,在那里真阴性实际上无法计数。F1 的盲点是其生境特有的,并非纯粹伦理问题。
- 真实标签可能由系统制造:欺诈标签来自往往由警报触发的调查;筛查标签来自阳性筛查后的活检。评估在模型最重要之处可能最不可信。
- 更好的评估包括成本、校准、原始计数、警报量、漏掉的欺诈金额、警报疲劳、留出期、随机对照和反馈回路。
主要分歧
- 准确率:热衷者说,它只有在类别平衡且成本对称时才有用。怀疑者说,它回答的是一个狭窄的问题,单看并不具有误导性;脱离基线和计数来报告它才是错误。
- F1:热衷者将其视为实用的警钟。怀疑者则认为它是虚假的安慰:它把有争议的权衡压缩成一个可优化的数字,并忽略真阴性。
- 成本判断:怀疑者认为,偏向假阴性而不是假阳性是一种成本决策,而不是指标事实。精确率和召回率只是切分混淆矩阵。
- 决策感知评估:热衷者想要校准后的概率,加上决策影响和反馈回路。怀疑者要求反事实证据,指出仪表盘可以被操纵,标签有延迟,而且政策变化后校准会失效。
- 真实标签:观察者说,混淆矩阵假定存在外部标签,但标签往往与模型和机构相呼应。其他人关注部署和反馈;观察者说,衡量标准本身依赖于工具。
- 模型与部署:热衷者部分接受三元组,但为指标辩护。观察者坚持认为,对象是部署的轨迹,而不是静态分类器。
结论
- 没有任何指标天生就是“要紧之事”。准确率衡量总体正确性;精确率衡量正类预测可靠性;召回率衡量正类覆盖度;F1 在忽略真阴性的同时平衡两者。
- 在不平衡情况下,准确率可能空洞无物。精确率和召回率揭示错误类型,但取决于流行率和阈值。F1 进行概括,但不做决定。
- 指标不仅是描述性的:它们塑造阈值、激励、流行率、标签和表面上的成功。
- 稳健的评估应结合原始计数、基线、置信区间、成本曲线、校准、流行率监测、反事实证据,以及对标签生成的审查。把评估视为动态的、特定于人群的、决策感知的,而不是一个单一分数。
完整发言记录
Welcome. Today we ask what accuracy, precision, recall, and F1 are really looking at. Not just formulas, but what each number reveals about a classifier's behavior, and what it hides. My role is to keep us neutral, so I will not favor any metric. Please ground claims in examples, especially where class imbalance changes the story. No jargon without explanation. We will move from accuracy to precision, recall, then F1, then open the floor. Let's begin with a simple question: when is high accuracy actually misleading?
High accuracy is misleading when a model merely learns the majority class. In fraud detection where 99.8 percent of transactions are normal, a classifier that always says normal gets 99.8 percent accuracy while catching zero fraud, so it hides exactly the behavior we care about. In medical screening for rare disease, 99 percent accuracy may mean missing nearly every patient. The lesson is that accuracy is only useful when classes are balanced and the cost of false positives and false negatives is symmetric. We should immediately ask which error matters more, which leads naturally to precision and recall.
I agree the fraud example shows accuracy can be vacuous, but I object to saying accuracy is only useful when classes are balanced and costs symmetric. Accuracy answers one narrow question: overall correctness. In a 99.8 percent normal dataset, 99.8 percent accuracy is not misleading by itself; it reveals the trivial baseline. The error is reporting it without baseline, confidence intervals, or raw counts. Also, deciding false negatives matter more is a cost judgment, not a metric fact. Precision and recall only slice the confusion matrix; they hide prevalence and the burden of false alarms. We should ask what decision costs justify favoring one slice.
Both of you treat a metric as a property of the model. I would say it is a property of a triple: model, threshold, and population. The same classifier moved to a different hospital or a different fraud season shows different precision and recall without changing at all, because prevalence shifted underneath it. These numbers describe a deployment, not a device. Second, nobody has mentioned that metrics act back on the system. Once recall becomes the stated target, thresholds get tuned to it, and the number stops describing behavior and starts steering it, like a thermometer that also heats the room. So perhaps the sharper question is not what each metric looks at, but when it stops observing and starts pushing.
Observer's triple is exactly the future of evaluation: we should report model, threshold, population, and cost curve together. But I reject that this makes accuracy useless or metrics manipulative. If a model has 99.8 percent accuracy only because it predicts normal, that is a powerful negative result: it proves the classifier adds no value over a constant. Precision and recall then show where the signal is hiding: if recall is 0.02 for fraud, the model is blind to rare positives. F1 punishes that blindness, so I still trust it as a practical alarm bell. The answer is not to distrust metrics; it is to make them decision-aware.
I partly accept the triple point, but I challenge the idea that F1 is a practical alarm bell. It is a compromise metric with hidden ethics: the harmonic mean weights precision and recall as equally important and ignores true negatives. In fraud, a low F1 may reflect many cheap alerts, while a moderate F1 can coexist with missing catastrophic cases. F1 also drifts with prevalence because precision depends on base rate. If the target becomes F1, teams can tune thresholds to flatter it without reducing harm. Decision-aware evaluation means costs, calibration, and raw counts, not another single score.
One thing neither side has touched: these metrics have biographies. Precision and recall grew up in information retrieval, where true negatives are effectively infinite, since you cannot count every document you did not retrieve. F1 ignores true negatives not as an ethical choice but because in its native habitat there was nothing to count. Transplant it into a clinic or a bank, where true negatives are real and often the overwhelming majority, and you inherit a blind spot that now looks like a deliberate design decision. So part of this disagreement is not about math but about habitat: we keep asking tools shaped by one ecology to testify about another. I would also correct my own earlier triple: it is not static, because the model feeds back into the population term. A fraud detector that actually works deters fraud, so prevalence falls, so precision falls; the metric punishes the system for succeeding, like a predator whose census numbers crash because it hunted too well. That suggests the real object of evaluation is not a snapshot of a classifier but the trajectory of a whole deployment. The question worth asking is whether any of these numbers can survive contact with their own success.
Observer's biography is not a funeral for metrics but a blueprint for better ones. Precision and recall can survive their own success if we stop treating prevalence as background noise and start modeling it as a response to deployment. In fraud, report alerts per thousand, missed fraud dollars, alert fatigue rate, and how they move when thresholds change. If fraud falls, precision may fall, but the cost curve should reward fewer prevented losses and less customer harm. F1 remains a compact alarm, especially early in development, but the future is calibrated probabilities plus decision impact. The metric should not just watch the classifier; it should watch the whole feedback loop.
I will not call that a blueprint yet. Alerts per thousand and missed fraud dollars sound concrete, but they inherit the same problems: alert volume can be gamed by threshold shifts, missed dollars depend on delayed and incomplete labels, and prevented losses are counterfactual. Without randomized controls or holdout periods, you cannot tell whether precision fell because fraud declined or because the detector became indiscriminate. Calibration is also not a refuge; probabilities calibrated on yesterday's prevalence can be wrong after policy changes. F1 as an early alarm is still false comfort, because it collapses contested tradeoffs into one optimizable number. The hard question is what counterfactual evidence validates the feedback loop, not which dashboard we prefer.
Skeptic asks what counterfactual evidence could validate the loop, but I want to question something beneath both positions: the ground truth itself. Every confusion matrix assumes labels arrive from outside the system. In fraud, labels come from investigations, and investigations are opened mostly by alerts, often from the very model being judged. In screening, the gold standard biopsy tends to go to those who screened positive. So the true row of the matrix is partly manufactured by the attention of the institution doing the measuring. Fisheries science lives with this: catch per hour is read as fish abundance, but it confounds the fish with the fishermen. Maps of known ruins are largely maps of where archaeologists dug. We have covered metrics steering behavior and success eroding prevalence; neither of you has noted that the yardstick itself can be an echo of the instrument. The unsettling consequence is that evaluation may be least trustworthy exactly where the model matters most, because there labels, prevalence, and thresholds all move together.