In the past few years, the most counterintuitive and important discovery in the field of AI can be summarized in one sentence: making models bigger, data more, and computing power more abundant often leads to better results. The law of "scale for performance" is called the Scaling Laws. It is not a specific algorithm, but an empirical relationship about "input and output".
The core observation of the law of extension is that the loss of a model (roughly understood as "how accurate the prediction is") will exhibit a fairly smooth and predictable downward trend as the number of parameters, data, and computation increases. In other words, as long as the scale of these three factors is proportionally increased, the loss will follow an extrapolation curve downwards. This means that researchers can fit curves in small-scale experiments and then speculate on "what level I would probably reach if I were to magnify it tenfold," thus making judgments before actually investing huge amounts of computing power.
This pattern has a profound impact on industry strategy. It explains why everyone tends to go in the direction of "bigger" without agreement, and also brings up the discussion of "data is rarer than parameters": if parameters are added blindly and there is not enough data, the model may not be "fed enough", resulting in diminishing marginal returns. So "how to expand more efficiently under limited data" and "whether data quality can replace data quantity" have become new focuses.
But the law of expansion does not mean 'blindly stacking materials wins'. Firstly, it describes the overall trend, and for a specific task, the ability may not necessarily increase smoothly; Some abilities may "suddenly appear" after reaching a certain threshold in scale, while others may not show any improvement for a long time. Secondly, the cost of expanding the scale is extremely high, but the returns may decrease, and cost-effectiveness needs to be calculated. Once again, as the scale increases, it becomes even more difficult to align, control, and reduce inference costs.
Therefore, in recent years, there has been a greater emphasis on "smart scaling": using higher quality data, more reasonable architecture, and training formulas to maximize the effectiveness of each piece of computing power; Spend a large amount of computing power on training, while using distillation, quantization, sparsification, and other methods on the inference end to reduce costs. The law of extension provides direction and expectation, while engineering and algorithms determine how far one can go.
In summary, scale is not omnipotent, but in many cases, it is indeed the "main engine" driving capacity growth - provided that data and algorithms keep up.
[Reference source] Comprehensive compilation of industry information and publicly available materials from research institutions.