Back to Home

Introduction to Activation Functions: How the "Nonlinear Switch" of Neural Networks Works from ReLU to GELU, SwiGLU

September 24, 2026 at 08:02 AMSource: RunByAI0 comment(s)TechGuide

The activation function is the most inconspicuous yet almost ubiquitous component in neural networks. Its task is simple: perform a nonlinear transformation on the weighted sum of each neuron. But it is precisely this layer of nonlinearity that gives the stacked network the ability to fit complex functions - without it, no matter how many layers are stacked, the entire network is equivalent to a linear transformation.

1、 Why is nonlinearity necessary

The basic calculation of a neuron is to first perform weighted summation on the input (z=Wx+b), and then feed it into a function f to obtain the output a=f (z). If f is an identity mapping, then the two layers of the network stacked together still form y=W ₂ (W ₁ x+b ₁)+b ₂, essentially a linear model. Only by introducing nonlinearity can the network truly possess the ability to "bend" decision boundaries and approximate any complex function. This is also the origin of the term 'activation': it determines to what extent a neuron is activated and how much signal it transmits downstream.

2、 ReLU: Simple and Effective Classic

The expression for ReLU (Rectified Linear Unit) is f (x)=max (0, x): if the input is positive, it passes through as is, and if it is negative, it outputs zero. Its advantages are very prominent: extremely fast computation, constant gradient of 1 in the positive interval (alleviating the gradient vanishing problem of Sigmoid/Tanh in the early years), and in practice, it can also bring sparse activation. After AlexNet in 2012, ReLU almost became the default choice for deep networks. But it also has its shortcomings: the negative half axis gradient is zero, and once the input of a neuron is negative for a long time, it may never be updated again (commonly known as "death ReLU"). Therefore, variants such as LeakyReLU and ELU have emerged.

3、 GELU: Smooth version of ReLU, a frequent visitor to Transformers

GELU (Gaussian Error Linear Unit) does not have a one size fits all approach like ReLU, but instead performs smooth weighting based on input size. It can be roughly understood as determining how much proportion to retain based on the size of the input. Intuitively, GELU has a smooth transition near zero, allowing a small number of negative values to pass through, resulting in a more continuous gradient and smoother optimization. A large number of models such as GPT series and BERT have replaced the activation function in feedforward networks with GELU. The cost is that the calculation is slightly more expensive than ReLU, and in engineering, tanh's approximate implementation is usually used.

4、 SwiGLU: Gate Control Approach, the Mainstream of Current Large Models

SwiGLU belongs to the GLU (Gated Linear Unit) family. Its core idea is to use one signal to "gate" the other signal - one part is linearly transformed, and the other part is activated by Swish to act as a gate, and the two are multiplied element by element. Compared to a single activation function that only independently transforms each dimension, the gating structure allows for mutual adjustment between dimensions and has stronger expressive power. The cost is that the feedforward layer requires an additional weight matrix, typically reducing the hidden dimensions to about two-thirds to maintain a comparable number of parameters. The feedforward networks of mainstream large models such as LLaMA and PaLM commonly use SwiGLU, combined with RMSNorm and RoPE, forming one of the "standard formulas" for current large model structures.

5、 How to actually choose

Traditional CNN or small networks: ReLU remains the default option with the highest cost-effectiveness; Transformers and Large Language Models: GELU or SwiGLU are more common and excel in smoothness and expressiveness; If a large number of neurons become inactive during training, variants such as LeakyReLU and ELU can be tried. It should be emphasized that the activation function rarely determines success or failure alone. It is more often combined with normalization, parameter initialization, and learning rate scheduling to jointly affect the stability and final effectiveness of training.

6、 Summary

The activation function is the key switch for neural networks to move from "linear fitting" to "nonlinear expression". The evolution from ReLU to GELU and then to SwiGLU essentially revolves around the three goals of "smoother gradients, stronger expressiveness, and more stable training". Understanding the motivations behind them is more important than memorizing formulas.

[Reference source]

- Hendrycks & Gimpel, Gaussian Error Linear Units (GELUs), arXiv:1606.08415

- Noam Shazeer, GLU Variants Improve Transformer, arXiv:2002.05202

-Comprehensive compilation of industry information that has been publicly released

deep learningTransformer
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment