Back to Home

Introduction to Optimizer: From SGD to AdamW, Why Do Large Model Training Prefer Adaptive Methods

September 23, 2026 at 01:32 PMSource: RunByAI0 comment(s)TechGuide

When training a large model, the parameters can easily reach billions. How to "adjust" these parameters based on the loss function directly determines whether the training can converge and how fast it converges. The optimizer is responsible for this matter. This article outlines the evolution from the most basic gradient descent to AdamW and explains why large model training prefers adaptive optimizers.

1、 What is the optimizer doing

The essence of training a neural network is to find a set of parameters that minimize the loss function as much as possible. Given the gradient of the loss for each parameter (i.e. "which direction to adjust and how much to adjust to reduce the loss"), the optimizer determines how to update the parameters at each step: direction, amplitude, and whether to utilize historical information. The simplest rule is gradient descent - walking along the negative gradient direction at a fixed step size.

2、 From batch gradient descent to stochastic gradient descent (SGD)

Batch gradient descent calculates gradients using all data at once, which is accurate but slow. Stochastic Gradient Descent (SGD) uses only one (or a small batch) sample at a time, with frequent updates and fast speed, but comes at the cost of high gradient noise and path jitter. In practice, small batch SGD is commonly used to strike a balance between efficiency and stability.

SGD also derived the Momentum approach: exponentially weighted averaging of historical gradients, which is equivalent to adding inertia to parameter updates, can suppress jitter and also pass through smaller local depressions. This step is crucial because most of the adaptive methods that follow also drive the quantity.

3、 Adaptive learning rates: AdaGrad, RMSProp, and Adam

There is an awkward thing about fixed learning rates: the appropriate step size for different parameters may vary greatly. The adaptive method allows each parameter to have its own learning rate that dynamically adjusts with training.

AdaGrad: Scale the learning rate based on the sum of squares of the historical gradients of the parameters, and the step size of parameters with larger gradients will automatically decrease. The disadvantage is that this accumulation only increases without decreasing, and the learning rate will monotonically decay to almost zero, making it almost impossible to learn in the later stage.

RMSProp: Change "summation" to exponential moving average, focusing only on recent gradient amplitudes, alleviating the problem of AdaGrad learning rate decay too quickly.

Adam: Simultaneously maintaining the first-order moment (mean, equivalent to momentum) and second-order moment (variance, equivalent to the scale of RMSProp) of the gradient, and performing bias correction on both, is equivalent to a combination of "momentum+adaptive step size". It converges quickly, is insensitive to hyperparameters, and is the default choice for deep learning in the long run.

4、 Adam's Controversy and AdamW

Adam has a widely discussed problem: L2 regularization coupled with adaptive scaling leads to a weakening of the regularization effect. AdamW separates weight decay from gradient updates and directly applies it to parameters to achieve "decoupled weight decay". This change makes regularization more controllable and has better generalization, making it the mainstream choice for training large language models today.

5、 Why does big model training prefer AdamW

1. Fast convergence and robustness: insensitive to hyperparameters such as learning rate, suitable for expensive training that can last for months, and less prone to parameter tuning pitfalls.

2. Adapt to imbalanced gradients: The gradient scales of different layers and parameters vary greatly, making parameter by parameter adaptation more stable.

3. Compatible with large-scale parallelism: Mature implementation, widely supported by mainstream frameworks and distributed training schemes.

In practice, techniques such as learning rate warm-up (Warmup), cosine decay, and gradient clipping are often used to further stabilize training.

6、 Summary

The evolution of optimizers is essentially a continuous improvement between "accurate update direction" and "appropriate update step size": SGD adds momentum to solve directional jitter, AdaGrad/RMSprop introduces adaptive step size, Adam combines the two, and AdamW decouples weight decay. By understanding this main thread, one can comprehend what the optimizer behind that line in the big model training script is doing.

[Reference source] Comprehensive compilation of optimization algorithm papers published publicly and mainstream deep learning framework documents.

deep learninglarge model
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment