If accuracy is the transcript of a classification model, then perplexity (PPL) is the most classic ruler of language models. It attempts to answer a question: How unexpected does the model feel when faced with a real text?
1、 Starting from 'predicting the next word'
The core task of language models is to predict the probability distribution of the next token. Given the previous context, the model will assign a probability to each candidate word in the vocabulary. If the model is smart enough, it will give a high probability to the word that actually appears; If the model is poor, real words can only be assigned a very low probability. Confusion is an indicator that summarizes these "unexpected levels".
2、 Definition and Calculation
For a real sequence of length N, the model gives a probability pui for each real token. We first take its negative logarithm, and then average all positions to obtain the cross entropy: cross entropy=- (1/N) × ∑ log (p.ui). The perplexity is its exponent, which is PPL=exp (cross entropy). That is to say, perplexity is an exponential form of "average cross entropy loss per token". The smaller it is, the less unexpected the model is for real text and the stronger its predictive ability.
3、 Intuitive understanding
A common explanation is that perplexity can be roughly understood as the average hesitation between words at each step of the model. If PPL is equal to 10, it is roughly equivalent to the model making choices among 10 equally probable candidate words at each position; The lower the PPL, the more convergent the candidate set, and the more confident the model is. A constant uniformly distributed language model, whose perplexity is approximately equal to the size of the vocabulary, is also a natural reference upper bound.
4、 Pits that must be noted when using
Confusion is strongly correlated with word segmentation methods. The same text can be divided into different word lists and tokenizers with varying numbers of tokens, resulting in different levels of confusion. Therefore, comparing PPL across models and tokenizers is often meaningless and only holds true when compared under the same tokenizer and test set.
Confusion measures the accuracy of predicting the next token, which is not always consistent with downstream task performance. A model with a lower PPL may not necessarily perform better in question answering, reasoning, and instruction following - these abilities largely depend on training data, fine-tuning, and alignment methods, rather than simply prediction probabilities.
Confusion is almost insensitive to sampling quality, creativity, and conversational experience.
5、 It is still important
Although evaluating large models today relies more on various benchmark tests and manual evaluations, perplexity remains the most practical and inexpensive monitoring metric in the training process: it can reflect in real time whether the model converges, whether the data is abnormal, and also help researchers evaluate the effectiveness of different data ratios and hyperparameters. As a "low-level thermometer" for language modeling, it has not been replaced to this day.
[Reference source]
Jurafsky & Martin, Chapter on Language Models and Confusion in Speech and Language Processing (Third Draft); Brown et al, 《An Estimate of an Upper Bound for the Entropy of English》,Computational Linguistics,1992。