Most people who have used local large models have a common experience: typing is too slow. Especially in dialogue based generation, the model must squeeze out each token one by one, making it difficult to improve speed. Speculative Decoding is one of the most popular inference acceleration methods in recent years - its core idea is counterintuitive: use a smaller model to "guess first" and then let the larger model "correct".
1、 Why is the generation of large models so slow
The big language model uses autoregressive generation: every time a token is generated, all previous content must be passed through the network again. This process is "serial", and although GPUs have strong computing power, their utilization is severely reduced due to the need to wait one by one. That is to say, 'limited video memory bandwidth and insufficient computing power'.
2、 Speculation on how decoding works
Its approach is divided into three steps:
1. Draft: Use a small, fast draft model with fewer parameters to predict candidate sequences for the next few tokens in one go.
2. Verification: Submit this draft to the target model at once and have it score each position in parallel to determine if these guesses fit its probability distribution.
3. Accept or rollback: Compare from scratch one by one, adopt those that can be matched, and once there is a mismatch, backtrack to the divergence point, and the large model will regenerate that token by itself, and then continue to the next round.
The key is that the computational cost of validating a batch of tokens and generating a single token for a large model is almost the same (both are forward once). So as long as the small model guesses accurately enough, it can "earn" several tokens at once, significantly improving throughput.
3、 Why is it so attractive
The biggest advantage of decoding is that it is "lossless" - theoretically, it can ensure that the final output is consistent with the distribution of the results generated by the large model one by one, without sacrificing quality. This distinguishes it from compression methods such as quantization and pruning, which can cause accuracy loss, and therefore it is popular in large model services.
4、 Common implementation variants
1. Dual model scheme: Directly use small models of the same series as drafts, such as drafting the small-sized version for the large-sized version.
2. Medusa: No additional small models are provided, but multiple "prediction heads" are added to the large model to guess multiple subsequent positions in parallel.
3. EAGLE and others: Utilize feature level information to create more accurate drafts, further improving acceptance rates.
In engineering, inference frameworks such as vLLM already have built-in support for inference decoding.
5、 When will the profits be high
The effectiveness of guessing decoding depends on the "acceptance rate": the more similar the draft model and the large model are, and the more routine the task (such as summarization, translation, code completion) is, the easier it is to guess and the more obvious the acceleration is; On the contrary, if it is a highly creative open generation with low acceptance rate, the revenue will be discounted.
6、 One sentence summary
It is speculated that decoding uses the method of "small model answering and large model judging" to turn the bottleneck of serial generation of large models into parallel verification, which results in doubled throughput without sacrificing quality. Understanding it means understanding a main thread of current large-scale model inference optimization.
[Reference source] Comprehensive compilation of industry information released publicly (such as Speculative Decoration, Speculative Sampling, Medusa, EAGLE, and other public papers, as well as vLLM official documents).