Back to Home

Introduction to Speculative Decoding: Drafting Small Models and Making Decisions on Large Models

September 20, 2026 at 08:01 AMSource: RunByAI0 comment(s)TechGuide

The inference of large models is slow, often in a counterintuitive way: every time a token is generated, the entire model needs to be run once, and the GPU's large amount of computing power is actually idle waiting for video memory read and write. Speculative Decoding is a type of acceleration method proposed to address this bottleneck. Its core idea can be explained in one sentence: let a cheap small model guess first, and then let the big model verify the ticket at once.

1、 Why is word for word generation so slow

Autoregressive generation is serial: the Nth token must wait until the N-1 tokens are computed before starting. The computational cost of each step is not large, but the model weights need to be transferred from the memory to the computing unit - this step is a typical memory bound. The result is that the floating-point computing power of the GPU is largely idle, and the real time is spent on data handling. Since we are stuck in handling rather than calculating, a natural idea is: can we handle it all at once and verify multiple candidate tokens at the same time?

2、 How does speculative decoding work

The process can be divided into three steps:

1. Draft: Use a much smaller parameter draft model to continuously guess K tokens in a normal autoregressive manner.

2. Parallel verification: Send the prompt word along with these K draft tokens to the target large model, and the output distribution for each position can be calculated in one forward iteration.

3. Accept/Reject: Compare from left to right one by one, and reserve draft tokens that are consistent with the distribution of the target model according to the rules; When encountering the first inconsistent position, replace it with the token provided by the target model itself and discard all subsequent drafts.

The key is that verifying these K tokens only takes one forward time for the large model, but it may push several tokens at once.

3、 Why doesn't it change the output result

This is the most beautiful part of speculative decoding. In the standard implementation, the accepted draft token and the token obtained by the target model's own sampling are equivalent in probability distribution (guaranteed by one correction sampling), so the final output distribution is the same as directly generating it word for word with the large model. That is to say, it belongs to lossless acceleration rather than approximation - this is also the biggest difference between it and methods such as quantization, distillation, and pruning that can change model behavior.

4、 When it's useful and when it's useless

-The prerequisite for effectiveness is that the output distribution of the draft model is sufficiently close to that of the target model, resulting in a high acceptance rate. Small models from the same series are most suitable as drafts, and shallow versions of the target model can also be used for drafting.

-The acceleration ratio depends on the acceptance rate and draft length: in public experiments, when the acceptance rate is high and the K value is appropriate, it is often possible to achieve acceleration within several times; If drafts are always rejected, verification costs will become an additional burden.

-Not proficient in scenarios such as high temperature sampling and creative writing, where the output distribution is very scattered, the acceptance rate will significantly decrease.

-Not saving memory: It saves time, but instead requires loading an additional draft model, making it more cost-effective in online services with abundant memory and latency sensitivity.

5、 Several common variants

-Bullish speculation: draft multiple candidate branches at once, choose the most likely one to be accepted, and improve the single round hit rate.

-Self Speculative: Without loading additional small models, use the shallow or skip layer structure of the target model as a draft to save memory.

-Draft based on lookup tables/prompt words: directly copying a paragraph from the context as a candidate, suitable for tasks such as summarization, rewriting, and code completion that involve copying the original text in large quantities.

-Combined with continuous batch processing: In mainstream inference frameworks, speculative decoding is often used in conjunction with PagedAttention, continuous batch processing, etc., further increasing throughput.

6、 Relationship with adjacent technologies

-Quantization: Quantization reduces the cost of each forward iteration, while speculative decoding reduces the number of forward iterations. The two can be combined.

-Compared to distillation: Distillation is the process of compressing the capabilities of a large model into a small model, with the ultimate goal of eliminating the need for a large model; Speculative decoding always retains the large model, while the small model only serves as an auxiliary.

-Compared to batch processing, batch processing improves throughput (while serving more people), while speculative decoding mainly improves single request latency (everyone sees words faster).

Summary in one sentence: Speculative decoding does not change the model or result, but rather replaces taking only one step at a time with guessing a few steps first and verifying again, using the extra computing power to return the wasted memory access time. If your service is stuck in the experience of first character delay and word for word pronunciation, this is worth adding to the optimization list on the deployment side.

[Reference source] Comprehensive compilation of publicly published academic papers and mainstream inference framework public documents.

Inference AccelerationLarge Language Model (LLM)Reasoning optimization
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment