Why accuracy often "lies"
After training a classification model, the most intuitive metric is accuracy: the proportion of predicted pairs of samples. But when the sample is extremely imbalanced, it can be seriously misleading. For example, out of 1000 samples, only 10 are "diseased". As long as the model predicts "healthy" uniformly, the accuracy rate can reach 99%, but not a single patient can be found.
Confusion Matrix: The Starting Point of All Indicators
To see where the model is wrong, first look at the confusion matrix: true cases (TP), false positive cases (FP), true negative cases (TN), and false negative cases (FN). Almost all classification indicators are defined by their combinations.
- Precision=TP/(TP+FP): How many of the predictions that are "positive" are truly positive. Use it when following 'Don't wrongly accuse good people'.
- Recall=TP/(TP+FN): How many true positive cases have been identified. Use it when following 'Don't miss bad guys'.
- F1=2 · P · R/(P+R): The harmonic average of precision and recall, used to strike a balance between the two.
The trade-off between precision and recall
The two often complement each other. Lowering the threshold for positive judgment results in an increase in recall rate but a decrease in accuracy; Raising the threshold is the opposite. The choice of which one to prioritize depends on the business: spam filtering prefers to release rather than mistakenly kill (with a bias towards accuracy); It is better not to miss disease screening and to have more follow-up examinations (biased towards recall rate).
ROC and AUC: Looking at Sorting Ability from a Different Perspective
The ROC curve plots the true case rate (TPR) and false positive case rate (FPR) at different thresholds together, and the area under the curve is the AUC. AUC measures the overall ability of a model to rank positive samples ahead of negative samples, regardless of the threshold; When the categories are extremely imbalanced, the PR curve (precision recall curve) often reflects the true performance better than the ROC.
Summary
There is no universally applicable 'good indicator'. First, think carefully about which type of error the business is most prone to - false positives or false negatives, then select indicators based on this and make choices on the threshold, so that the evaluation will not be limited to beautiful numbers.
【 Reference source 】 Comprehensive compilation of publicly released machine learning textbooks and industry technical materials.