When model manufacturers release new versions, they always mention a series of benchmark scores. What are these abbreviations testing for? Why is there a huge difference in the ranking of the same model on different rankings? This introductory article will help you clarify several of the most common evaluation benchmarks.
1、 What does a benchmark do
A benchmark is a standardized set of questions and scoring methods that allow different models to compare under the same input, thereby answering to some extent the question of 'who is stronger'. But it measures performance on specific tasks and does not necessarily equate to comprehensive abilities in real-world scenarios.
2、 Several classic benchmarks that cannot be bypassed
·MMLU (Massive Multitask Language Understanding): a multiple-choice question covering dozens of disciplines such as mathematics, history, law, medicine, etc., mainly examining the general knowledge and reasoning of the model. It was proposed by Hendrycks et al. in 2020, and the name "Massive Multitask" precisely illustrates its design concept.
·GSM8K: A set of math application problems for primary school difficulty, requiring models to provide problem-solving steps and numerical answers, and testing multi-step reasoning ability. It was proposed by Cobbe et al. in 2021.
·HumanEval: an evaluation set for code generation, providing function signatures and explanations, requiring models to complete code and use unit testing to determine correctness. It was proposed in 2021 along with Codex related work.
·TruthfulQA: Specifically designed to examine whether a model repeats common human errors, focusing on "authenticity" rather than knowledge, proposed by Lin et al. in 2021.
·In addition, there are common sense and reasoning benchmarks such as HellaSwag and ARC, which have long appeared in various technical reports.
3、 Why can't we trust all scores
1. Data pollution. If the test questions happen to appear in the training data, the model is equivalent to "having done the original questions", and the score will be overestimated, which is also the reason why the credibility of the ranking has been questioned for a long time.
2. Ranking brushing and overfitting. In order to make the scores look good, the training process may intentionally or unintentionally conform to the distribution of evaluations.
3. Ceiling effect. The phrase 'too simple a question' can lead to a clustering of high scores, making it difficult to see the real gap. Therefore, the community continues to introduce more difficult alternative benchmarks.
4. Disconnected from real experiences. The benchmarks are mostly multiple-choice questions and short answers, while actual use also requires abilities such as long essay writing and multiple rounds of tool calling, and the two are not completely consistent.
4、 How to view benchmarks correctly
Treat benchmarks as biased signals rather than absolute rankings. When comparing models, prioritize specialized evaluations that are similar to your own tasks. If conditions permit, use your own data for small-scale testing, and make a comprehensive judgment based on latency, cost, privacy, and deployment methods. Score is a starting point, and ultimately it has to fall into specific scenarios.
[Reference source]
Hendrycks et al, Measuring Massive Multitask Language Understanding,2020; Cobbe et al, Training Verifiers to Solve Math Word Problems,2021; Chen et al, Evaluating Large Language Models Trained on Code,2021; Lin et al, TruthfulQA: Measuring How Models Mimic Human Falsehoods,2021。