Back to Home

Introduction to Learning Rate Scheduling: Warmup, Cosine Attenuation, and Training Stability

September 24, 2026 at 01:33 PMSource: RunByAI0 comment(s)TechGuide

Learning rate is one of the most important hyperparameters when training neural networks, which determines the magnitude of parameter updates at each step. However, a fixed learning rate is often not the optimal solution: careful probing is required in the early stages of training, rapid decline is needed in the middle stages, and fine convergence is needed in the later stages. Learning Rate Scheduling is a set of methods that allow the learning rate to dynamically change with the training process. This article explains several common scheduling strategies and why they are effective.

1、 Why is a fixed learning rate not sufficient

The learning rate is too high, and the loss is prone to oscillation or even divergence; Too small, slow convergence, and prone to stopping at poor solutions. The demand for "step size" varies at different stages of training, leading to the idea of "adjusting the learning rate with training": the overall curve shows an initial increase followed by a decrease, gradually converging.

2、 Warmup: Heat up first and accelerate behind

Warmup refers to gradually increasing the learning rate from a very small value to a set peak in the first few steps of training. It is mainly used to stabilize the initial stage of training: when parameters are randomly initialized, the gradient direction is very unreliable, and using a high learning rate at this time can easily bias the parameters; When working with an adaptive optimizer, the early estimation of second-order statistics may not be accurate, and the high learning rate will amplify this instability. Therefore, Warmup is commonly used for training large models.

3、 Common attenuation strategies

Step Decay: Multiply the learning rate by a coefficient every fixed number of rounds, which is simple and intuitive.

2. Cosine Decay: The learning rate smoothly decreases from the peak along the cosine curve to nearly 0, which is common in current large-scale model training.

3. Linear Decay: Linear decay at the end of training, easy to implement.

4. Exponential Decay: Continuous decay at a fixed rate.

4、 Why 'big first, small later' is intuitive

In the early stages of training, parameters are still far from the target, and larger step sizes can quickly explore; In the later stage of training, the parameters are close to the optimal region and need to be finely adjusted with small steps to avoid jumping back and forth near the optimal point. Cosine, linear, and other decays perfectly fit this "coarse to fine" process.

5、 Several experiences in engineering

-Warmup steps usually only account for a small proportion of the total steps, but are indispensable in large batch or model training.

-The peak learning rate, attenuation curve shape, and Warmup length need to be adjusted together, and looking at one alone often does not yield a conclusion.

-Learning rate scheduling is often coordinated with optimizers (such as AdamW): optimizers are responsible for updating direction, while scheduling is responsible for step size.

Conclusion

Learning rate scheduling does not change what the model can represent, it adjusts the 'pace of learning'. Warmup allows the model to start smoothly, while decay makes it converge more steadily and finely in the later stages. Understanding this rhythm is more valuable than memorizing a specific numerical value.

[Reference source]

-Attention is All You Need (Vaswani et al., 2017, proposed learning rate scheduling with Warmup)

- SGDR: Stochastic Gradient Descent with Warm Restarts(Loshchilov & Hutter, 2016, Cosine annealing)

deep learninglarge model
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment