Since 2024, a new approach has emerged in the field of big models, parallel to "making models bigger": allowing models to "think for a while" before answering. OpenAI's O-series, DeepSeek's R-series, and other inference models treat the computational complexity of the inference phase itself as a dimension that can be scaled up. This article outlines the basic logic, common implementation methods, and changes it brings to engineering practice of this route.
1、 From 'predicting the next word' to 'writing a draft before answering'
The traditional dialogue model is autoregressive: given context, it directly predicts the next token. For problems that require multi-step deduction (mathematics, code, constrained planning), the amount of computation that can be completed in one forward propagation is limited, and the model is prone to blurting out a seemingly reasonable but incorrect answer.
The core idea of the inference model is to make the process before generating the answer explicit: the model first forms a long intermediate inference (usually referred to as a chain of thought), and then gives the final answer. The middle token itself is an additional calculation - each step reads the previous text and makes new inferences, equivalent to exchanging sequence length for "depth of thought".
2、 Calculation during testing: a new scaling dimension
The industry refers to this additional investment as' test time computation ', which corresponds to the amount of computing power invested during training. The general rule of thumb is that for the same model, allowing a longer chain of thought often leads to higher accuracy in mathematical and coding tasks; However, the benefits are not non-linear, and overthinking simple problems may introduce errors, increase delays, and incur costs. Therefore, most systems will implement "budget control" and allocate thinking length based on the difficulty of the problem.
This is also why the same generation of inference models typically offer different levels of inference strength, allowing callers to choose between accuracy, latency, and cost.
3、 Common implementation paths
1. Long thought chains trained in reinforcement learning: reinforce the model's spontaneous generation of longer reasoning processes through verifiable rewards (whether the answer is correct or not, whether the code can pass the test), rather than relying on manual annotation of each step.
2. Sampling and Rearrangement: Parallel sampling of multiple candidate inference paths, and then selecting the best using scoring models or validators, essentially trading the computational power of inference for accuracy.
3. Search based reasoning: Expand the reasoning process as a tree or graph, and perform pruning and backtracking within it.
4、 Three things to reconsider in engineering
The cost model has changed: output tokens may be more than input, billing and latency budgets need to be re estimated according to the "thinking length", caching and streaming returns are more important.
Observability has changed: the intermediate reasoning process can be presented to users, but it may also contain incorrect directions, and should be treated as debugging information rather than authoritative explanations.
The evaluation has changed: only looking at the final answer accuracy will mask the issue of inference efficiency, and it is necessary to also consider the "computational cost per unit accuracy".
5、 Boundary and Reminder
The improvement of reasoning ability does not mean 'not making mistakes'. The model may still deviate during the intermediate process, especially when there is a lack of external validation methods (tools, retrieval, unit testing). For production systems, a more secure combination is still: inference models+external tools+verifiable inspection processes.
[Reference source] Comprehensive compilation of industry information publicly released.