The ideal outcome of training a neural network is for the model to stop learning the patterns in the data. However, in reality, it is difficult for us to know in advance when that moment will arrive - training too long can lead to overfitting, while training too little can lead to underfitting. Early Stopping is a simple yet highly effective method: instead of guessing a fixed number of training rounds, it's better to focus on the validation set while training, and call a stop in a timely manner when the model starts to 'deteriorate'.
1、 Why do we need to stop early
On the training set, the loss almost always decreases continuously as the training progresses. But what truly determines the quality of a model is its performance on unseen data. The common process is that the performance on the validation set first improves with training, and then starts to rebound after reaching a certain lowest point - at this point, the model has already started to remember the noise of the training data, which is overfitting. The idea of early stopping is to seize the moment when the validation set performs the best, rather than blindly running the training.
2、 Core approach
The implementation of early stopping is very lightweight, requiring only three things: first, draw a portion from the training data as the validation set (or use cross validation); Secondly, evaluate on the validation set every few training steps (one epoch or fixed number of steps); Thirdly, continuously record the validation set metrics and retain the model parameters with the best performance. When the performance of the validation set does not improve for several consecutive times, the training is stopped and rolled back to the previously saved best weights.
3、 Key parameter: patience
The most important parameter for early stopping is patience, which means' how many consecutive times without improvement are allowed before stopping '. Setting the patience too small (such as 1) may cause premature interruption due to short-term fluctuations in the validation curve, missing out on better results later on; Setting it too large will waste computing power and weaken the significance of preventing overfitting. In practice, it is often taken between 5 and 20, and a min_delta (minimum improvement margin) is used to filter out almost negligible small fluctuations.
4、 Smarter judgment
Just looking at verification losses is not omnipotent. When the categories are imbalanced or the losses are inconsistent with the final indicators, directly monitoring the indicators that are truly of concern (such as F1, AUC) is often more reliable. Another approach is to use the "generalization gap" - the difference between training performance and validation performance - as a stop signal, and when this gap continues to widen, it indicates that overfitting is occurring.
5、 Its relationship with other means
Early stopping is essentially a regularization: it does not change the model structure, but actually limits the model's fit to the training data. It does not conflict with methods such as Dropout, weight decay, and data augmentation, but is often used together. Compared to other regularization methods that require additional parameter tuning, early stopping has almost zero cost, making it one of the default configurations in deep learning training.
6、 One sentence summary
Early stop reminds us that the goal of training is not to maximize the training set, but to make the model perform best in the real world. Knowing when to stop is often just as important as knowing how to move forward.
[Reference source]
Prechelt,《Early Stopping — But When?》, Recorded in Neural Networks: Tricks of the Trade, 1998; Goodfellow、Bengio、Courville,《Deep Learning》,MIT Press, 2016 (Chapter 7 on Regularization and Early Stopping).