Back to Home

Introduction to Residual Connections: Why Neural Networks Can Stack Up to Hundreds of Layers

September 24, 2026 at 01:32 PMSource: RunByAI0 comment(s)TechGuide

Residual Connection is one of the most inconspicuous yet crucial designs in modern deep networks. It does not increase parameters or change the expressive power of a single layer, but it turns "deepening the network" from a difficult problem to a routine operation. From ResNet to Transformer, almost all models that can stack up to dozens or even hundreds of layers rely on it to maintain trainability internally. This article explains what problem residual connections need to solve, their form, and why they are so important for training stability.

1、 Why is the deep network 'unable to stack'

Intuitively, the deeper the network, the more complex the functions it can express, and the better the effect should be. But experiments have found that when the number of layers increases to a certain extent, the training error does not decrease but instead increases, which is called the degradation problem. It is neither overfitting (training errors are also increasing), nor can it be simply attributed to gradient vanishing, but rather the naive expectation that "deeper networks should not be worse than shallow ones" is difficult to meet.

2、 Form of residual connection

The core formula of residual connection is concise: y=F (x)+x, where x is the input of this layer and F (x) is the transformation learned by this layer. The network no longer directly fits the target mapping, but instead fits the "relative input increment" (residual). When F approaches zero, the output is approximately equal to the input, leaving the network with an identity path of 'doing nothing'.

3、 Why is it effective

1. High speed gradient channel: During backpropagation, the additive structure allows gradients to propagate back to the shallow layer along an identity path almost without loss, alleviating the vanishing gradient in deep networks.

2. Identity mapping is a "safe default": if certain layers cannot learn useful features, the network can make them approximate identity, at least without dragging down the effect.

3. Optimization is simpler: the learning objective becomes a "correction based on input", which is easier than fitting the entire mapping from scratch.

4、 Role in Transformer

Each sublayer of Transformer is written in the form of x+Sublayer (x), and attention and feedforward networks are hung on the residual backbone. It is this identity backbone that enables stable training of transformers with dozens or even hundreds of layers, and also makes normalized positions (Pre LN/Post LN) a design point worth discussing.

5、 A detail that is often overlooked

Residual connection requires that the input and output dimensions be consistent, otherwise projection is needed to align the dimensions. In addition, residuals are "additive" and sensitive to numerical ranges, which is why they often appear together with normalization.

Conclusion

The idea of residual connections is very simple: instead of letting the network learn a complex mapping from scratch, it is better to let it learn how much to add to existing results. It is this design of "leaving a way out for optimization" that forms the engineering foundation for deep learning to continuously deepen.

[Reference source]

- Deep Residual Learning for Image Recognition(He et al., 2015)

- Attention Is All You Need(Vaswani et al., 2017)

deep learninglarge model
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment