Normalization is a subtle but crucial technique in deep learning. It does not change what the network can represent, but directly affects whether the training can be stable and converge quickly. In large models, the position and form of normalization are even repeatedly adjusted - from LayerNorm to RMSNorm, becoming a part of the current mainstream structure. This article explains what problems normalization aims to solve and the differences in thinking between LayerNorm and RMSNorm.
1、 Why is normalization necessary
During deep network training, the input distribution of each layer will continuously drift as the parameters of the previous layer are updated, which is commonly known as internal covariate shift. The trouble it brings is that each layer has to constantly adapt to new input distributions, training is prone to oscillations, and the learning rate cannot be increased. The normalization method is to readjust the activation values to a relatively stable distribution after each layer calculation (usually by first reducing the mean, then dividing by the standard deviation, and finally scaling and shifting with learnable parameters), so that subsequent layers can receive inputs with controllable scales. In this way, training is more stable and convergence is faster.
2、 The difference between BatchNorm and LayerNorm
The early BatchNorm was very successful on CNN: it statistically analyzed the mean and variance on the batch dimension and normalized each channel. However, in NLP and Transformer scenarios, sequence lengths vary and batch sizes are limited, making batch statistics very unstable. Reasoning also relies on statistical data from the training period, which can be awkward. LayerNorm has taken a different approach: normalizing individual samples in the feature dimension without relying on batch size, resulting in consistent training and inference behavior. This is precisely why Transformer prefers LayerNorm.
3、 The position of LayerNorm in Transformer
There are two common ways to place Transformers: one is to place normalization after sub layers (attention, feedforward), i.e. Post LN; Another approach is to place it before the sub layer, namely Pre LN. Early primitive Transformers used Post LN, but deep training was prone to instability and required fine warmups. Later large models generally switched to Pre LN (or improved Post LN), which has smoother gradients, is more tolerant of learning rates, and is easier to stack deep.
4、 RMSNorm: Remove the mean and leave only scaling
RMSNorm is a simplified version of LayerNorm: it only scales the input based on root mean square (RMS), removing the step of reducing the mean. Intuition is that the model truly cares about the scale of the vector, not its center; After eliminating the mean subtraction, the calculation is more efficient, but the actual test results are basically not discounted. Major models such as LLaMA use RMSNorm, combined with Pre LN placement and SwiGLU feedforward, to form an efficient and stable structural combination.
5、 Key points in practice
Normalization is usually used in conjunction with residual connections, which together make deep networks easier to train: residuals provide a direct path and normalize stable scales. In addition, the learnable scaling parameters of the normalization layer are generally initialized to 1, so that they do not initially disrupt the original distribution. It should be noted that the normalization effect is coupled with parameter initialization and learning rate scheduling, and adjusting it alone may not solve all training problems.
6、 Summary
The value of normalization lies in providing scale controllable inputs to each layer, allowing deep networks to train more stably and quickly. The trend from BatchNorm to LayerNorm and then to RMSNorm is: more suitable for sequence modeling, less dependent on batch statistics, and more computationally efficient. It, together with residual connections and appropriate activation functions, forms the foundation on which modern large-scale models can be built deep and trained.
[Reference source]
- Ba, Kiros & Hinton, Layer Normalization, arXiv:1607.06450
- Zhang & Sennrich, Root Mean Square Layer Normalization, arXiv:1910.07467
- Xiong et al., On Layer Normalization in the Transformer Architecture, arXiv:2002.04745
-Comprehensive compilation of industry information that has been publicly released