LLM Prefix Caching: What Gets Reused?
Repeated questions about the same long material can waste prefill work if the entire input is recomputed each time. Prefix caching reuses a computed shared beginning. It matches the actual prefix, not text with roughly the same meaning.
Reuse computation state, not answers
vLLM’s automatic prefix caching reuses KV state for a shared prefix, allowing later requests to skip the cached prefill portion. It is not a full-answer cache. Different questions after the same document still require processing new input and generating new answers.
Consider token sequences [10,11,12,13,14] and [10,11,12,15]. Their common prefix has length three. These integers are teaching IDs, not real tokenizer output.
Calculate an ideal bound first
Without reuse, the inputs contain 5+4=9 units. Under ideal full-prefix reuse, the second request adds only one, for a total of six. Saving three input units does not imply a one-third latency reduction. Cost is not simply linear in token count, and queuing, scheduling and generation remain.
def common_prefix(a, b):
n = 0
for x, y in zip(a, b):
if x != y: break
n += 1
return n
first = [10, 11, 12, 13, 14]
second = [10, 11, 12, 15]
assert common_prefix(first, second) == 3
assert common_prefix(first, [99] + second) == 0
assert common_prefix([], second) == 0
assert common_prefix(first, first) == 5
raw_units = len(first) + len(second)
ideal_units = len(first) + len(second) - common_prefix(first, second)
assert (raw_units, ideal_units) == (9, 6)
print('shared prefix:', 3)
print('input units without / with ideal reuse:', raw_units, ideal_units)
The executed code checks shared prefixes, a changed first token, empty input and identical input. It measures sequence matching and input units, not vLLM performance. Moving common text to another position does not establish a hit under this shared-beginning model.
Changes worth testing
Stable system instructions and repeated material before the varying question may increase shared prefixes. Do not change task meaning merely for caching. Templates, whitespace, message structure and tokenization can also change the actual sequence.
The vLLM guide states that prefix caching reduces prefill work, not the decoding work of generating new tokens. Long outputs, little shared prefix or evicted cache entries can limit the benefit.
Compare cold and warm cache runs separately, controlling input and output lengths. Observe first-token time, total latency and cache hits. Warm-cache gains are not a promise for every request. KV cache memory estimates also illustrate why cache resources require a budget.
References
Publication note: actually published as a catch-up on October 10, 2026 (Beijing time), retaining the planned 2026-10-10 09:00 article date. The code is an original teaching check, not a real-model performance test.
Support
If this article helped you, you can support this site.
Click a QR code to enlarge. More options: support page。


