Back to Home

Introduction to Dropout: Why Neural Networks Intentionally "Randomly Drop Packages"

September 27, 2026 at 01:32 PMSource: RunByAI0 comment(s)TechGuide

Dropout is one of the simplest yet most effective regularization techniques in deep learning: it randomly "shuts off" a portion of neurons during training, using this deliberate "incompleteness" to force the network to learn more steadily, thereby alleviating overfitting.

1、 What is overfitting

Overfitting refers to a model that performs better and better on the training set but deteriorates on unseen data - it memorizes the details and even noise of the training samples without learning the true patterns. For neural networks with billions of parameters, overfitting is almost inevitable, so various regularization methods are needed to constrain it.

2、 The basic method of Dropout

During each forward propagation, Dropout temporarily sets the output of some neurons to zero with a probability p (dropout rate), and these neurons do not participate in forward computation or backpropagation at this step; Another batch of neurons will be discarded in the next iteration. Therefore, the network that "survives" each training step is different.

At the inference (testing) stage, there is no longer random dropout, and all neurons participate in the computation. In order to keep the output scale consistent with the training, one of the two methods is usually adopted: one is to divide the retained activation values by (1-p) during training, that is, inverted dropout, which is also the mainstream practice in modern implementation; The second is to multiply the weights by (1-p) during testing, which is an early practice.

3、 Why is this actually effective

A straightforward analogy is team meetings: if a random number of people are absent each time, the remaining members must have the ability to independently solve problems, and the team will not overly rely on a few 'key gentlemen'. Neural networks are also similar - randomly shutting down neurons weakens co adaptation between neurons to specific peers, forcing each neuron to learn more useful features.

From a mathematical perspective, Dropout is equivalent to sampling a large number of sub networks that share weights with each other during the training process, and approximating an ensemble average of these sub networks during inference. This is a classic idea for improving generalization ability.

4、 Key points for use

The selection of dropout rate p: The hidden layer is usually 0.2~0.5, and the input layer is usually a smaller value (such as 0.1~0.2). A large p-value can lead to underfitting and difficulty in training.

In the training of modern large-scale Transformers, the use of Dropout has significantly decreased due to the popularity of techniques such as LayerNorm and weight decay, but it is still common in fine-tuning, small dataset training, and convolutional networks.

When using Dropout and BatchNorm simultaneously, it is important to note that the variance changes caused by Dropout may interfere with BatchNorm's statistics, and in practice, it is often necessary to adjust the order of the two or make trade-offs.

5、 Summary

Dropout trades "random destruction" for stronger generalization: actively losing packets during training and involving all members during inference. It reminds us that a good model is not about remembering training data, but about being able to provide stable and correct answers even in the presence of noise and disturbances.

[Reference source]

Srivastava, Hinton, Krizhevsky, Sutskever, Salakhutdinov,《Dropout: A Simple Way to Prevent Neural Networks from Overfitting》,Journal of Machine Learning Research,2014。

deep learningLarge Language Model (LLM)
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment