Back to Home

Introduction to KV Cache: Why does big model inference become slower as we chat, and how does caching save computing power

September 21, 2026 at 08:01 AMSource: RunByAI0 comment(s)TechGuide

Unlike questions that are answered with just one glance, during multiple rounds of conversation, you will find that the response speed of the large model is not stable: the first sentence is always half a beat slow, and the later the conversation, the faster it gets, but the longer the conversation, the higher the video memory usage. The decisive factor behind this is a mechanism that almost all inference frameworks use - KV Cache.

1、 Autoregressive generation: Say only one word at a time

The generation of text by large models is "autoregressive": given all the previous words, predict the next word, then concatenate the new word to continue predicting the next one. That is to say, to generate a response of length N, the model needs to forward calculate N times, adding only one token each time.

The problem lies in the attention mechanism: every time a new word is generated, it must be compared with all the previous words to calculate attention. If there is no cache, the Nth step is to recalculate the Key and Value of the first N-1 words, and the computational complexity increases with the square of the length, so the more we talk, the slower it becomes.

2、 What has KV Cache done

The idea of KV Cache is very straightforward: since the K (Key) and V (Value) of the preceding words will not change in the subsequent steps, then count them once and save them for reuse. In this way, each step only needs to calculate Q, K, V for the new token, and then pay attention to the historical K and V in the cache, reducing the computational complexity from square level to linear level.

The cost is video memory. Cache needs to store a pair of K and V for each layer, attention head, and historical token. The longer the context, the more significant the occupancy; On a 7B scale model, the cache size of a token is typically several hundred KB.

3、 Why does it determine how long we can chat

The memory usage of KV Cache is proportional to the length of the context, while GPU memory is fixed. When the context becomes longer and there are more concurrent requests, the cache will fill up the video memory, becoming a key bottleneck that limits concurrency and maximum context length. This is also the direct reason for the emergence of a series of optimization techniques such as PagedAttention (vLLM), KV Cache quantization, and sliding window attention.

4、 Several common optimization directions

-PagedAttention: Manage KV Cache like an operating system manages memory paging, reducing fragmentation and improving concurrency.

-KV Cache Quantization: Pressurize the cache from FP16 to INT8/INT4, and switch it back to video memory with a little precision.

-Group Query Attention (GQA) and Multi Query Attention (MQA): allowing multiple attention heads to share K and V, directly reducing cache size.

-Sliding window and sparse attention: only retaining the most recent or relevant history, actively discarding distant cache.

Understanding KV Cache actually means understanding where the main cost of large model inference lies: often it's not the weights, but the memory and bandwidth invested in "remembering context".

[Reference source] Comprehensive compilation of industry information publicly released (including vLLM team's public paper on PagedAttention).

Reasoning optimizationLarge Language Model (LLM)
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment