LLM Prefix Caching: What Gets Reused?

黎 浩然/ 10 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

Repeated questions about the same long material can waste prefill work if the entire input is recomputed each time. Prefix caching reuses a computed shared beginning. It matches the actual prefix, not text with roughly the same meaning.

Reuse computation state, not answers

vLLM’s automatic prefix caching reuses KV state for a shared prefix, allowing later requests to skip the cached prefill portion. It is not a full-answer cache. Different questions after the same document still require processing new input and generating new answers.

Consider token sequences [10,11,12,13,14] and [10,11,12,15]. Their common prefix has length three. These integers are teaching IDs, not real tokenizer output.

Request A: 10 11 12 | 13 14
Cached prefix: 10 11 12
Request B adds: 15
Original reuse diagram. Cache organization and matching conditions constrain actual reusable state.

Calculate an ideal bound first

Without reuse, the inputs contain 5+4=9 units. Under ideal full-prefix reuse, the second request adds only one, for a total of six. Saving three input units does not imply a one-third latency reduction. Cost is not simply linear in token count, and queuing, scheduling and generation remain.

def common_prefix(a, b):
    n = 0
    for x, y in zip(a, b):
        if x != y: break
        n += 1
    return n

first = [10, 11, 12, 13, 14]
second = [10, 11, 12, 15]
assert common_prefix(first, second) == 3
assert common_prefix(first, [99] + second) == 0
assert common_prefix([], second) == 0
assert common_prefix(first, first) == 5
raw_units = len(first) + len(second)
ideal_units = len(first) + len(second) - common_prefix(first, second)
assert (raw_units, ideal_units) == (9, 6)
print('shared prefix:', 3)
print('input units without / with ideal reuse:', raw_units, ideal_units)

The executed code checks shared prefixes, a changed first token, empty input and identical input. It measures sequence matching and input units, not vLLM performance. Moving common text to another position does not establish a hit under this shared-beginning model.

Changes worth testing

Stable system instructions and repeated material before the varying question may increase shared prefixes. Do not change task meaning merely for caching. Templates, whitespace, message structure and tokenization can also change the actual sequence.

The vLLM guide states that prefix caching reduces prefill work, not the decoding work of generating new tokens. Long outputs, little shared prefix or evicted cache entries can limit the benefit.

Compare cold and warm cache runs separately, controlling input and output lengths. Observe first-token time, total latency and cache hits. Warm-cache gains are not a promise for every request. KV cache memory estimates also illustrate why cache resources require a budget.

中文版

References

Official documentation

Publication note: actually published as a catch-up on October 10, 2026 (Beijing time), retaining the planned 2026-10-10 09:00 article date. The code is an original teaching check, not a real-model performance test.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*