When training neural networks, there is occasionally a heart wrenching situation: the loss value suddenly drops from a steady state to NaN, the model parameters instantly "collapse", and the previous hours of training fail. One of the most common causes of such accidents is gradient explosions. Gradient clipping is a simple and effective way to deal with it.
1、 Why do gradients explode
The training of neural networks relies on backpropagation: the gradient of the loss on the parameters is passed back layer by layer along the network, and a chain multiplication is performed every layer. When the network is deep or processes long sequences (RNN/Transformer), if these multiplication factors are generally greater than 1, the gradient will exponentially increase with the number of layers. Once the gradient of a batch is particularly large, a parameter update may push the weight to a very bad position, causing the loss to diverge and even become NaN.
Gradient explosion is particularly common in recurrent neural networks (RNNs) and transformers, and often occurs when the learning rate is high, there are extreme samples in the data, or the mixed precision training values are unstable.
2、 The basic idea of gradient clipping
The core idea is very straightforward: set a "ceiling" for gradients. If the gradient calculated this time is too large, proportionally reduce it to an acceptable range and then hand it over to the optimizer to update the parameters. It does not change the direction of the gradient, only limits its length, so it does not distort the original optimization intention of the model, but only prevents "taking too big a step and dragging it across".
3、 Two common practices
1. Clip by Norm: First, calculate the L2 norm (length) of the entire parameter gradient vector. If the threshold x_norm is exceeded, all gradients are proportionally reduced to make the overall norm equal to x_norm; Otherwise, it remains unchanged. It is currently the most mainstream method, corresponding to torch.nn.tils.comlip_grad_norm_ in PyTorch.
2. Clip by Value: Determine element by element and directly truncate gradients with absolute values exceeding the threshold to [- c, c]. The implementation is simple, but it will change the direction of the gradient. In practice, it is not as common to crop according to the norm.
4、 Key parameter: How to choose threshold
The threshold is too small, which is equivalent to repeatedly flattening useful learning signals, and training will slow down or even stagnate; The threshold is too high and does not provide protection. The common experience value in large model training is around 1.0, which still needs to be adjusted based on the actual performance of the model size, learning rate, and loss curve. A practical approach is to observe the gradient norm distribution before pruning: if it exceeds the threshold for a long time, it indicates that the learning rate may be too high; If the threshold is almost never touched, it means that this "protective umbrella" is basically useless.
5、 Why is it still important today
In the era of Transformer, gradient clipping is almost standard. It is often used in conjunction with college learning rate warm-up: the gradient is unstable in the early stages of training, so it is first trimmed and then preheated to ensure a smooth start to training. It has extremely low cost and almost no additional memory, but can significantly improve the stability of long sequence and deep network training, making it a typical representative of "small changes, big benefits".
6、 One sentence summary
Gradient clipping is like installing a "speed limiter" on the training process: it does not change the direction of progress, only ensuring that each step does not overturn due to excessive force.
[Reference source]
Pascanu、Mikolov、Bengio,《On the difficulty of training recurrent neural networks》,ICML 2013( Propose to crop according to norm gradient); PyTorch official documentation torch.nn.tils.comlip_grad_norm_.