Back to Home

Introduction to Regularization: L1, L2 and Weight Attenuation, How to Put the "Anti overfitting" Chain on Models

September 29, 2026 at 08:10 AMSource: RunByAI0 comment(s)TechGuide

In the previous articles, we talked about "training stabilizers" such as Dropout, early stop, and weight initialization. They share a common goal: to prevent the model from memorizing training data, but to truly learn patterns. The regularization that we are going to talk about today is the most fundamental and core concept in this family of methods.

##What is regularization?

Regularization, in simple terms, means adding an additional "penalty" to the training objective to limit the model from becoming too complex. The more complex the model (with more parameters and greater weights), the more capable it is of capturing all the noise and coincidences in the training samples. As a result, it performs almost perfectly on the training set, revealing its true form when encountering new data - this is known as * * overfitting * *.

The idea of regularization is: instead of pursuing only "minimum training error", it is better to also require "parameters to be as simple as possible". So the loss function changed from: Loss=data error to: Loss=data error+λ x complexity penalty. Among them, lambda is a hyperparameter used to balance the two. The larger the λ, the heavier the punishment, and the more conservative the model is.

##L2 regularization and weight decay

The most common is * * L2 regularization * *: the penalty term is equal to half of the sum of squares of all parameters, that is, lambda × ∑ w ². It tends to reduce the overall weight, making the model less sensitive to small changes in input and resulting in smoother output.

For neural networks, L2 regularization in gradient descent is manifested as an additional action: with each update, the parameters are first scaled down a little bit proportionally before being updated normally. So it is often referred to as * * Weight Decay * *. In many deep learning frameworks, these two names are often used interchangeably, but strictly speaking, under standard SGD, the two are equivalent. When replaced with adaptive optimizers such as Adam, their behavior will differ (which also gave rise to the practice of "decoupling" weight decay in AdamW).

##L1 regularization and sparsity

**L1 regularization penalizes the sum of absolute values of the parameters (λ × ∑ | w |). It has an interesting side effect: it will compress a portion of the weight * * to exactly 0 * *, resulting in * * sparse * * solutions. This is very useful in scenarios that require "feature selection" - the model automatically picks out a few important features and discards the unimportant ones.

One sentence comparison: L1 → Sparse, tends to "cut off" unimportant parameters; L2 → Smooth, tending to 'overall reduce' all parameters.

##How to use and how much to use?

-λ is a hyperparameter that needs to be adjusted, usually starting from a very small value (such as the order of 1e-4, 1e-5);

-λ is too small, regularization is almost ineffective, and the model still overfits; If λ is too large, the model is "pressed too hard", and both training and validation errors cannot be reduced, resulting in * * underfitting * *;

-In practical engineering, regularization is often used in conjunction with Dropout, early stopping, and data augmentation, in a multi pronged approach.

##Summary

Regularization is not about making the model "smarter", but about putting an appropriate shackle on it: restricting it from rote memorization while giving it enough freedom to learn real laws. L1 seeks sparsity, L2 seeks smoothness, and weight decay is the specific form of L2 in the optimizer. By combining this set of ideas with Dropout and early parking mentioned earlier, you will have a complete puzzle on how to make the model generalize.

[Reference source]

-Ian Goodfellow, Yoshua Bengio, Aaron Courville, "Deep Learning" (MIT Press, 2016), Chapter 7: Regularization for Deep Learning

-Comprehensive compilation of industry information that has been publicly released

AI机器学习正则化
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment