What is a loss function
When training a model, we have to have a "scoring" method: how far off is the prediction of the current set of parameters? The loss function, also known as the cost function, is the yardstick for answering this question. It converts the difference between the predicted values of the model and the true labels into a comparable number. The smaller the number, the less errors the model makes on this batch of data.
Imagine the training process as descending a mountain, where the value of the loss function is the altitude - gradient descent is responsible for finding the way, and the loss function is responsible for telling it where it is currently standing.
Why can't we just look at 'the right few'
The most intuitive indicator is accuracy: how many predictions were made correctly. But it has two fatal issues. Firstly, it is not differentiable and cannot calculate gradients, so gradient descent cannot be used for optimization; Secondly, it loses the information of "how much was wrong" - judging 0.51 as 1 and 0.99 as 1 are exactly the same in terms of accuracy, but the latter should actually be corrected more. So a smooth and differentiable loss function is used during training, and accuracy is only used for final evaluation.
Common loss functions
- Mean Square Error (MSE): For frequent regression tasks, the square of the error for each sample is averaged. The punishment for large errors is heavier because the square will amplify it.
- Mean Absolute Error (MAE): Taking the absolute value of the error and then averaging it makes it more robust to outliers, but at the cost of non smoothness near the zero point.
- Cross Entropy: The mainstream of classification tasks. It measures the difference between the probability distribution provided by the model and the true label distribution, and is often used in conjunction with Softmax output. The more confident the prediction is but the more wrong it is, the greater the punishment.
- Hinge loss: a classic choice for support vector machines (SVM), which only cares about whether the samples are paired and leave sufficient spacing.
How to choose
One sentence rule of thumb: Look at the error distribution in regression problems - use MAE for outliers and MSE or Huber for smooth differentiability; For classification problems, we basically choose cross entropy with closed eyes, Softmax for multi classification, and Sigmoid for binary classification. What really matters is not which one is' more advanced ', but whether it aligns with the goals you care about.
Summary
The loss function is the compass for model training: it translates "good" and "bad" into a differentiable and optimizable number. Only by choosing the right ruler can the model know which direction to progress in.
【 Reference source 】 Comprehensive compilation of industry information publicly released (Ian Goodfellow et al., Deep Learning, MIT Press; Official documentation of scikit learn; Stanford CS231n course lecture notes.