KV Cache: How It Works and Uses VRAM

黎 浩然/ 7 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

Fitting model weights into VRAM does not guarantee that a long conversation will fit. For an autoregressive language model using full causal attention, longer inputs and outputs generally require a larger KV cache. Separating weights from runtime state gives a more useful capacity estimate than looking only at the parameter count.

Contents
  1. What does the KV cache store?
  2. Why does a longer context need more VRAM?
  3. A reproducible capacity estimate
  4. Weight quantization does not automatically fix the cache
  5. What do different cache strategies trade?
  6. References

中文版 / Chinese version

What does the KV cache store?

Attention maps token representations to queries Q, keys K, and values V. When generating a new token, its query interacts with keys and values from earlier positions. The KV cache retains those historical K and V tensors so the model can reuse them instead of computing the same historical keys and values again.

KV Cache: How It Works and Uses VRAM: original technical diagram
A conceptual causal decoding flow, not a specific model architecture or a speed benchmark.

This is not a cache of completed answers, and it does not skip the historical context entirely. The new token still attends to the visible history. The cache trades storage for less repeated computation.

Why does a longer context need more VRAM?

For a simplified estimate, assume full attention in every layer, the same storage precision for K and V, and no alignment, allocator, or temporary-tensor overhead. Let L be the layer count, B the number of simultaneous sequences, T the cached length, H the number of KV heads, D the head dimension, and S the bytes per element.

KV bytes ≈ 2 × L × B × T × H × D × S

The factor 2 accounts for K and V. T includes the processed prompt and generated positions that must remain cached. For batches of different lengths, padding and cache allocation affect actual usage.

Variable Effect Caveat
Context length T Full-attention cache usually grows linearly with length Sliding-window and hybrid layers cannot always use one T for every layer
Sequence count B More concurrent sequences need more cache Prefix sharing and scheduling can change actual usage
KV head count H Fewer KV heads reduce this term Do not substitute the query-head count
Storage size S Fewer bytes reduce theoretical storage Quantization metadata and retained high-precision portions still use memory

A reproducible capacity estimate

For 32 layers, one sequence, 8,192 cached positions, eight KV heads, 128 dimensions per head, and two bytes per element, the estimate is 1 GiB. Keeping everything else fixed and using 32 KV heads makes it 4 GiB. This is a hypothetical calculation, not a measurement of a particular model.

def kv_gib(layers, batch, tokens, kv_heads, head_dim, bytes_per_element=2):
    return 2 * layers * batch * tokens * kv_heads * head_dim * bytes_per_element / 2**30
assert kv_gib(32, 1, 8192, 8, 128) == 1.0
assert kv_gib(32, 1, 8192, 32, 128) == 4.0
print('8 KV heads:', kv_gib(32, 1, 8192, 8, 128), 'GiB')
print('32 KV heads:', kv_gib(32, 1, 8192, 32, 128), 'GiB')

This pure Python example was executed and its assertions passed:

8 KV heads: 1.0 GiB
32 KV heads: 4.0 GiB

A GiB is 2 to the power of 30 bytes. These numbers cover only K and V tensors, not the entire inference service. Weights, activations, workspaces, cache allocation, and framework overhead require separate budgets.

Weight quantization does not automatically fix the cache

A model described as “4-bit” usually refers to its weights. Whether its KV cache is quantized, and at what precision, is a separate configuration. Even with smaller weights, a long context can make cache storage a major component.

First determine whether length, concurrency, or precision causes the pressure. Removing irrelevant history, limiting simultaneous generation, and choosing a cache strategy address different problems. A memory formula alone cannot promise faster inference.

What do different cache strategies trade?

A dynamic cache grows during generation. A static cache reserves capacity to keep shapes more stable, potentially leaving unused space for short requests. Offloading moves part of the cache to CPU memory, reducing device storage pressure while adding transfers. Cache quantization reduces storage representation with additional processing and precision tradeoffs.

Availability, compilation compatibility, and APIs depend on the model and software version. Check support, then measure time to first token, per-token latency, and peak memory on the target hardware.

The earlier article on language-model deployment and optimization (Chinese) covers other deployment costs. KV cache is a separate budget item, not a universal explanation for every performance problem.

References

The Chinese edition retained its planned date of October 7, 2026, 14:00 Beijing time and was published later that day. This article explains established technology rather than breaking news.

English edition added on October 7, 2026, after the Chinese edition. The article date matches the Chinese edition; it is not the actual time this English edition became public.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*