In the past two years, the focus of AI computing power has almost always been on "training" - the expansion of training clusters with thousands of cards and tens of thousands of cards, and large-scale pre training. Each round of model parameter increase is accompanied by the expansion of training clusters. But as large models enter the stage of large-scale application, a more realistic bottleneck is emerging: inference computing power.
Inference is the computational power consumed for each question and answer, as well as each generation, after the model is launched. Compared with training, reasoning has completely different characteristics: training is a centralized and pre planned large-scale project; Reasoning is a distributed and ongoing 'daily expense'. When an AI application has a large number of users, inference costs quickly accumulate and become an expenditure item that operators cannot ignore.
This is precisely why inference chips have received attention. There are several directions in the industry trend: firstly, general-purpose GPUs are still the mainstay, but the optimization for inference scenarios continues to deepen - low precision computing, sparsity, and video memory bandwidth optimization all directly affect the cost of single inference; Secondly, dedicated inference chips and neural network acceleration units are developing simultaneously in the cloud and on the edge, allowing AI capabilities to extend to mobile phones PC、 Sinking of automobiles and other equipment; Thirdly, the division of labor between inference and training is becoming increasingly clear, and cloud vendors are starting to redesign server and cluster architectures based on inference workloads.
For developers, the decrease in inference costs directly determines whether the product can be scaled up. For the industry, the second half of the computing power competition is no longer simply a competition of "who has trained a bigger model", but a competition of "who can deliver model capabilities to every user at lower cost and with lower latency".
It can be foreseen that with the continuous improvement of inference efficiency, the application of large models will enter a more practical stage - from technical demonstrations to true commercial closed loops.