Many large model applications share a common feature: each request must be accompanied by the same long and fixed content - system prompt words, product descriptions, knowledge base fragments, code specifications. This part of the content requires recalculating attention every time, which is both costly and time-consuming. Prompt caching is an optimization method designed for repeating prefixes.
1、 The root cause of the problem: the prefix is repeatedly recalculated
The big model is generated by autoregression, and every word generated needs to look at all the previous words. In order to reuse the same prefix in multiple requests, the engine needs to save the intermediate state (key value in attention, i.e. KV) corresponding to this prefix. The problem is that by default, each new request is independent, and no matter how the prefix is the same, it has to be recalculated from scratch.
2、 Save the intermediate state of the prefix
The idea of prompt word caching is very straightforward: if the beginning of two requests is exactly the same, then the calculation result of this prefix can be reused. Cache the intermediate state of the prefix to video memory or high-speed storage during the first request; When subsequent requests hit the same prefix, skip the pre fill calculation of this prefix and continue generating from the cached position.
3、 Several prerequisites for hitting cache
Cache is not automatically effective, usually there are several conditions: the prefix must be completely consistent (not even one character apart, so fixed content should be placed at the beginning and variable content should be placed at the end); Cache has a lifecycle, and if not accessed for a period of time, it will be cleared; Different service providers have their own regulations on the minimum length that can be cached and the billing method for hits.
4、 Why is it worth doing
For long context applications, the benefits are often direct. The longer the system prompt words and the more frequent the calls, the more obvious the savings brought by caching - it not only reduces the delay caused by repeated calculations, but also may significantly reduce the API cost charged by input tokens. Multiple rounds of dialogue, Agent loops, and batch processing of the same document are all natural applicable scenarios.
5、 Common pitfalls
One is to put mutable content in the prefix, which leads to failed hits every time; The second reason is that the cache is assumed to be permanently valid, but the performance drops after expiration and the reason cannot be found; Thirdly, it ignored the differences in caching rules among different service providers and made ineffective optimizations. Additionally, it should be noted that caching only saves duplicate calculations and does not allow the model to "remember" additional information.
Summary
The prompt word cache transforms the "repeated prefix" from a cost item that needs to be recalculated every time to a fixed expense that can be paid in full at once. For projects that stably call the large model API, it is a good partner with a clear hierarchical design of prompt words - first divide the content into "fixed/variable" layers, and then let the cache save the fixed layer for you.
[Reference source] Comprehensive compilation of industry information publicly released.