KV Cache: How It Works and Uses VRAM
Fitting model weights into VRAM does not guarantee that a long conversation will fit. For an autoregressive language model using full causal attention, longer inputs and outputs generally require a larger KV cache. Separating weights from runtime state gives a more useful capacity estimate than looking only at the parameter count.
Contents
What does the KV cache store?
Attention maps token representations to queries Q, keys K, and values V. When generating a new token, its query interacts with keys and values from earlier positions. The KV cache retains those historical K and V tensors so the model can reuse them instead of computing the same historical keys and values again.

This is not a cache of completed answers, and it does not skip the historical context entirely. The new token still attends to the visible history. The cache trades storage for less repeated computation.
Why does a longer context need more VRAM?
For a simplified estimate, assume full attention in every layer, the same storage precision for K and V, and no alignment, allocator, or temporary-tensor overhead. Let L be the layer count, B the number of simultaneous sequences, T the cached length, H the number of KV heads, D the head dimension, and S the bytes per element.
KV bytes ≈ 2 × L × B × T × H × D × S
The factor 2 accounts for K and V. T includes the processed prompt and generated positions that must remain cached. For batches of different lengths, padding and cache allocation affect actual usage.
| Variable | Effect | Caveat |
|---|---|---|
| Context length T | Full-attention cache usually grows linearly with length | Sliding-window and hybrid layers cannot always use one T for every layer |
| Sequence count B | More concurrent sequences need more cache | Prefix sharing and scheduling can change actual usage |
| KV head count H | Fewer KV heads reduce this term | Do not substitute the query-head count |
| Storage size S | Fewer bytes reduce theoretical storage | Quantization metadata and retained high-precision portions still use memory |
A reproducible capacity estimate
For 32 layers, one sequence, 8,192 cached positions, eight KV heads, 128 dimensions per head, and two bytes per element, the estimate is 1 GiB. Keeping everything else fixed and using 32 KV heads makes it 4 GiB. This is a hypothetical calculation, not a measurement of a particular model.
def kv_gib(layers, batch, tokens, kv_heads, head_dim, bytes_per_element=2):
return 2 * layers * batch * tokens * kv_heads * head_dim * bytes_per_element / 2**30
assert kv_gib(32, 1, 8192, 8, 128) == 1.0
assert kv_gib(32, 1, 8192, 32, 128) == 4.0
print('8 KV heads:', kv_gib(32, 1, 8192, 8, 128), 'GiB')
print('32 KV heads:', kv_gib(32, 1, 8192, 32, 128), 'GiB')
This pure Python example was executed and its assertions passed:
8 KV heads: 1.0 GiB
32 KV heads: 4.0 GiB
A GiB is 2 to the power of 30 bytes. These numbers cover only K and V tensors, not the entire inference service. Weights, activations, workspaces, cache allocation, and framework overhead require separate budgets.
Weight quantization does not automatically fix the cache
A model described as “4-bit” usually refers to its weights. Whether its KV cache is quantized, and at what precision, is a separate configuration. Even with smaller weights, a long context can make cache storage a major component.
First determine whether length, concurrency, or precision causes the pressure. Removing irrelevant history, limiting simultaneous generation, and choosing a cache strategy address different problems. A memory formula alone cannot promise faster inference.
What do different cache strategies trade?
A dynamic cache grows during generation. A static cache reserves capacity to keep shapes more stable, potentially leaving unused space for short requests. Offloading moves part of the cache to CPU memory, reducing device storage pressure while adding transfers. Cache quantization reduces storage representation with additional processing and precision tradeoffs.
Availability, compilation compatibility, and APIs depend on the model and software version. Check support, then measure time to first token, per-token latency, and peak memory on the target hardware.
The earlier article on language-model deployment and optimization (Chinese) covers other deployment costs. KV cache is a separate budget item, not a universal explanation for every performance problem.
References
- Transformers: How caching works, first-party documentation of the mechanism.
- Transformers: Cache strategies, including support constraints. The main branch changes over time; no specific-version API behavior is guaranteed here.
The Chinese edition retained its planned date of October 7, 2026, 14:00 Beijing time and was published later that day. This article explains established technology rather than breaking news.
English edition added on October 7, 2026, after the Chinese edition. The article date matches the Chinese edition; it is not the actual time this English edition became public.
Support
If this article helped you, you can support this site.
Click a QR code to enlarge. More options: support page。


