When deploying a large model as a service, "whether it's fast or not, how many people can be served at once" is often more critical than "whether it can run". Besides the model itself, it is often the inference engine that truly determines throughput and cost. This article takes vLLM as an example to explain the core problems that inference engines need to solve, as well as the ideas behind PagedAttention.
1、 The two stages of reasoning: Prefill and Decoding
Large model text generation usually involves two steps:
Prefill: Read the entire prompt at once, calculate and cache the key values (KV) of all tokens. This step involves highly parallel matrix operations, which are computationally intensive.
Decoding: Generate tokens one by one, using the KV of all previous tokens for each step. This step is carried out sequentially for each token, mainly due to the limitation of video memory bandwidth, and GPU computing power is often idle.
Because the output length is not fixed and only one token is generated at each step, the decoding stage is the focus of throughput optimization.
2、 KV Cache and Video Memory Fragmentation
To avoid duplicate calculations, the historical Key/Value will be cached during inference, known as KV Cache. For long context and large models, KV Cache may occupy more memory than model weights, and it continues to grow with generation.
Traditional implementation pre allocates a whole block of contiguous video memory for each request based on the 'maximum possible length'. The problem is also very obvious: firstly, excessive reservation causes waste; secondly, different request lengths vary, resulting in a large number of fragments in the video memory, which severely limits the number of requests that can be served simultaneously.
3、 The core idea of PagedAttention
The PagedAttention proposed by vLLM draws inspiration from the paging concept of virtual memory in operating systems: it cuts KV Cache into fixed sized "blocks", no longer requiring a contiguous segment of video memory, but using a "block table" to map logically contiguous KV to physically dispersed blocks. The advantage is:
1. On demand allocation: How many blocks are needed to occupy how much video memory, reducing reserved waste.
2. Eliminate fragmentation: The block size is fixed, and physical memory can be efficiently reused.
3. Shared prefix: If multiple requests prompt the same (such as system prompt words, few shot examples), they can share the same batch of blocks and only need to save one copy, which is particularly beneficial for batch scenarios of "multiple requests with the same template".
4、 Continuous Batching
Traditional batch processing requires waiting for all requests in a batch to be generated before switching to the next batch, as slow requests can slow down the entire batch. Continuous batch processing (also known as iterative scheduling) re batches after each decoding step: whoever generates it exits, and whose KV block is freed up immediately adds a new request. As a result, the GPU hardly runs idle, significantly improving throughput. The block based management of PagedAttention is the foundation for efficient operation of continuous batch processing.
5、 Benefits and trade-offs in practice
Throughput improvement: Official and multi-party practices have shown that compared to simple implementation, throughput can be improved several times under similar latency, and the benefits are more significant in long context and multi request scenarios.
Memory utilization: Paging and sharing prefixes make more efficient use of video memory, allowing the same graphics card to support more concurrency.
Attention: Parameters such as block size and scheduling strategy can affect performance; Small batch and extremely short output scenarios have limited returns; The ultra long context is still constrained by the upper limit of video memory and requires the use of quantization, prefix caching, and other means.
6、 Summary
The value of inference engines lies in using computing power and video memory bandwidth on the cutting edge: PagedAttention solves the fragmentation and waste of KV Cache with paging, and continuous batch processing prevents GPU from idling. Understanding these two points highlights the key to improving throughput for engines like vLLM.
[Reference source] Comprehensive compilation of vLLM technical documents and related paper materials that have been publicly released.