Around 2020, a counterintuitive discovery changed the direction of the entire AI industry: as long as models continue to grow, data is fed in, and computing power is piled up, the performance of models can almost steadily improve along a predictable curve. This law is called the Scaling Law, and it is also the true confidence behind the name 'big model'.
1、 What is the scaling law
Simply put, it describes the power-law relationship between the model's loss (roughly understood as the level of surprise when predicting the next word) and several key variables: the number of model parameters (N), the amount of training data (D, usually measured in tokens), and the training power (C, roughly proportional to the product of N and D).
Draw the logarithm of these three quantities, and the loss will decrease along a smooth straight line. In other words, performance improvement does not rely on a sudden inspiration, but on engineering problems that can be estimated in advance and invested according to plan.
2、 From 'bigger is better' to 'proportionally increase together'
In the early days, people tended to blindly pile up parameters, but soon found that if the parameters rose and the data did not keep up, the returns would quickly decline, and the model would even be difficult to learn. So a more refined version emerged - under a given computing budget, there exists a solution that optimizes the ratio of parameters to data.
This conclusion directly gave rise to the "compute optimal" training approach: instead of training a super large but insufficient data model, it is better to train a moderately sized and more data abundant model. Later on, the performance of a batch of open-source models also continuously confirmed the importance of reasonable allocation.
3、 Why is it so important
Predictability: Before investing millions of dollars in training, teams can use small-scale experiments to fit curves and predict the final performance of large models, reducing the risk of "taking a gamble".
Guiding resource allocation: It turns "how much data, how large models, and how much computing power are needed" into a computable engineering problem.
Explanation of "emergence": When the scale crosses certain thresholds, the model suddenly possesses abilities that are almost invisible on small models, and this phenomenon is called "emergence".
4、 Where is the boundary of the scaling law
The scaling law is not omnipotent, and several limitations are being repeatedly discussed: data walls - the total amount of high-quality public text is limited, and simply "feeding more" is difficult to sustain; Declining returns - losses decrease at a slower rate, while training costs increase almost exponentially; Not just looking at losses - what users truly care about are specific abilities such as reasoning, programming, and factual accuracy, and a decrease in losses does not necessarily mean a synchronous improvement in every ability.
Therefore, in the past two years, the focus of the industry has shifted from "infinite amplification pre training" to "post training" and "inference computation" - using finer alignment, retrieval enhancement, tool invocation, and inference models that allow the model to "think more for a while" in exchange for additional capability improvements.
Summary
The scaling law has pushed large models from "alchemy" to "programmable engineering". It explains why manufacturers are willing to invest massive computing power and reminds us that scale is not the only answer. Proportions, data quality, and post training are also key factors determining success or failure.
【 Reference source 】 Comprehensive compilation of research papers related to the scaling law that have been publicly published and public interpretations from mainstream technology media.