When training neural networks, we often hear the phrase 'forward propagation calculates results, backward propagation changes parameters'. Backpropagation is the cornerstone of deep learning, on which almost all model training is built. This article uses as straightforward language as possible to explain what it is doing and why it can make the model more accurate with practice.
1、 First, clarify the goal: minimize the 'error'
There are a large number of parameters (weights and biases) in neural networks, and the output during random initialization is basically meaningless. The goal of training is to find a set of parameters that make the model's predictions on the training data as close as possible to the true answer. The indicator for measuring "how far apart" is called the loss function (Loss), commonly used in regression with mean square error and in classification with cross entropy. Training is constantly adjusting parameters to minimize losses.
2、 Forward propagation: Calculate predicted values and errors
The input data starts from the input layer, passes through the weighted sum and activation function of each layer, and is transmitted all the way to the output layer to obtain the prediction result. By substituting the predicted results and the actual labels into the loss function, we can determine how wrong the current set of parameters is. This step is called forward propagation, which is essentially a series of matrix operations.
3、 Backpropagation: Spread the error back to each parameter
The key question is: There are thousands of parameters in the network, which one to modify and how many to modify is the most effective? The answer given by backpropagation is to use the chain rule of calculus to calculate the partial derivative of the loss for each parameter layer by layer from the output layer back, that is, the gradient.
Intuitive understanding: The output error is jointly caused by each layer. The chain rule is like "responsibility allocation", where the error is passed back layer by layer along the calculation graph according to its contribution size, and finally the gradient of each parameter is obtained. The gradient tells us in which direction and how much to adjust this parameter, which can quickly reduce the loss.
4、 Gradient descent: descend the mountain according to the gradient
After obtaining the gradient, use gradient descent to update the parameters: new parameters=old parameters - learning rate x gradient. It can be imagined as going downhill blindfolded, with the gradient being the current steepest downhill direction, and the learning rate being the size of each step taken. Repeating the process of "forward → reverse → update" usually results in a gradual decrease in losses.
5、 Several key points in practice
1. Learning rate: If it is too high, it will oscillate or even diverge, and if it is too low, it will converge slowly. It is often dynamically adjusted in conjunction with learning rate scheduling.
2. Gradient vanishing and exploding: When the number of layers is very deep, the gradient may become extremely small or large in the feedback. Residual connection, normalization, and reasonable initialization are commonly used to alleviate this.
3. Batch: Estimate gradients using a small batch of samples (small batch gradient descent), balancing speed and stability.
4. Automatic differentiation: Modern frameworks such as PyTorch and TensorFlow use computation graphs and automatic differentiation to automatically complete backpropagation, and developers only need to define forward computation and loss.
6、 Summary
Backpropagation is not mysticism, it is just an efficient engineering implementation of "chain rule+gradient descent": the error is obtained in the forward direction, the gradient of each parameter is calculated in the reverse direction, and then the parameters are fine tuned along the gradient direction. After understanding this main line, let's take a look at the minor adjustments LoRA、 Concepts such as distributed training will be much smoother.
[Reference source] Comprehensive compilation of publicly released deep learning textbooks and technical documents.