If there is only matrix multiplication, even the deepest neural network will only be a large linear transformation, and it will not have complex patterns at all. What truly brings neural networks to life is the seemingly insignificant component in each layer - the activation function. It injects nonlinearity into the network, which is the key to deep learning's ability to fit complex functions.
1、 Why do we need activation functions
Linear operations have a fatal limitation: when multiple layers of linear superposition are added, the result is still equivalent to a single layer of linear transformation. That is to say, without an activation function, stacking multiple layers is useless. The activation function inserts nonlinearity in the middle to enable the network to fit any complex mapping.
2、 Common activation functions
1. Sigmoid: Pressing the output to 0~1, commonly used in the early days. The disadvantage is that the gradient at both ends is close to 0, which can easily lead to "gradient vanishing" and make it difficult to train deep networks.
2. Tanh: When compressed to -1~1, it is more symmetric than Sigmoid, but there is also the problem of vanishing gradient in the saturation region.
3. ReLU (Rectified Linear Unit): f (x)=max (0, x), positive half axis linear, negative half axis zeroing. Extremely fast computation and alleviation of gradient vanishing are important turning points for deep networks. The disadvantage is the negative region 'death ReLU' - once the input is negative for a long time and the gradient is 0, the neuron may never update again.
4. Leaky ReLU/PReLU: Leave a small slope on the negative half axis to alleviate the issue of dead ReLU.
5. GELU and SiLU (Swish): With smoother curves and smooth transitions near zero, GELU is the mainstream choice in Transformers and large models, and is widely used in BERT and GPT series.
3、 Softmax is another type
Strictly speaking, Softmax is commonly used in the output layer to convert a set of scores into a probability distribution with a total sum of 1. In classification tasks, it is used to determine which category it belongs to.
4、 How to choose
Traditional CNN uses ReLU and its variants (Leaky ReLU) as the most convenient; Transformers and large models often use GELU and SiLU, and are used in conjunction with normalization; Softmax is used for output layer classification, but it is generally not used for regression tasks.
5、 One sentence summary
The activation function is a switch like presence in neural networks: it transforms linear stacks into truly nonlinear models, giving deep networks the ability to express complex worlds. The evolution history from Sigmoid to ReLU and then to GELU is also a history of deep learning that becomes deeper and more stable as it is trained.
[Reference source] Comprehensive compilation of industry information that has been publicly released (such as relevant public papers and official documents of mainstream deep learning frameworks).