GQA: How It Reduces the KV Cache

黎 浩然/ 11 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

When estimating an LLM’s KV cache, query-head count is not necessarily KV-head count. GQA lets groups of query heads share key and value heads, retaining multiple queries while caching fewer KV tensors. It shares K and V; it does not force every query to produce the same output.

Contents
  1. Eight query heads need not mean eight KV pairs
  2. KV-head count determines the raw cache ratio
  3. Shared KV can still produce different outputs
  4. Paper results are conditional evidence
  5. Check the actual deployment
  6. Sources

Eight query heads need not mean eight KV pairs

Consider eight query heads. MHA gives each a corresponding KV-head pair. MQA shares one pair across all queries. GQA sits between them: four queries might share each pair, giving two pairs in total. Here a “KV head” refers to matching key and value heads, not one tensor combining K and V.

Q0 · Q1 · Q2 · Q3
Shared K0 / V0
Q4 · Q5 · Q6 · Q7
Shared K1 / V1
Original grouping diagram: eight query heads share two KV-head pairs. Group numbering is illustrative; it specifies no model checkpoint.
Structure Query heads KV heads Sharing
MHA 8 8 One pair per query
GQA 8 2 Four queries per pair
MQA 8 1 One pair for all queries

The GQA paper was first submitted to arXiv in May 2023 and appeared in the peer-reviewed EMNLP proceedings that December. This is historical research, not a new announcement. The paper describes grouping as an interpolation between MHA and MQA.

KV-head count determines the raw cache ratio

Assume one request, L layers, T cached tokens, Hkv KV heads per layer, d components per head and b bytes per component. Keys and values have matching shapes and compact storage. Raw cache bytes are:

KV bytes = 2 × L × T × Hkv × d × b

With L=32, T=4096, d=128 and b=2, eight KV heads take 512 MiB, two take 128 MiB and one takes 64 MiB. One MiB is 2²⁰ bytes. With all those conditions fixed, two-head GQA stores one quarter of the raw KV data of eight-head MHA.

This excludes weights, activations, alignment, allocator reserves and multi-device replication. It does not establish a fourfold reduction in total memory or inference time. For requests of different lengths, sum their actual cached lengths rather than representing the whole batch with one short request.

Shared KV can still produce different outputs

Queries q₁ and q₂ in the same group can differ. Even with identical K and V, softmax(q₁Kᵀ/√d)V and softmax(q₂Kᵀ/√d)V can differ. KV sharing restricts representational freedom without forcing identical query attention distributions.

The original teaching code below was executed. It checks all three cache sizes, the eight-to-two group mapping and two different queries reading identical keys and values. The final outputs differ. Scalar values keep the example short; full attention usually uses value vectors.

import math

def kv_bytes(layers, tokens, kv_heads, head_dim, bytes_per_value):
    return 2 * layers * tokens * kv_heads * head_dim * bytes_per_value

for name, heads in [("MHA", 8), ("GQA", 2), ("MQA", 1)]:
    size = kv_bytes(32, 4096, heads, 128, 2)
    assert size == {"MHA": 512, "GQA": 128, "MQA": 64}[name] * 1024**2
    print(name, size // (1024 ** 2), "MiB")

query_heads, kv_heads = 8, 2
assert query_heads % kv_heads == 0
mapping = [h // (query_heads // kv_heads) for h in range(query_heads)]
assert mapping == [0,0,0,0,1,1,1,1]

def attend(query):
    keys = [(1.0,0.0), (0.0,1.0)]
    values = [10.0,20.0]
    scores = [sum(a*b for a,b in zip(query,k))/math.sqrt(2)
              for k in keys]
    weights = [math.exp(s-max(scores)) for s in scores]
    return sum(w*v for w,v in zip(weights,values))/sum(weights)

# Two queries share exactly the same keys and values.
a, b = attend((1.0,0.0)), attend((0.0,1.0))
assert not math.isclose(a,b)
print(round(a,6), round(b,6))

This code loads no model, trains no GQA model and uses no GPU. Its arithmetic and local attention assertions are not model-quality or performance benchmarks.

Paper results are conditional evidence

The main experiments use T5.1.1 and examine checkpoint conversion followed by additional training. The authors report quality and inference tradeoffs on summarization, translation and question-answering tasks, with timing on TPUv4. These conditions do not guarantee the same outcome for arbitrary models, GPUs and workloads.

Changing a KV-head configuration field alone does not turn an existing MHA checkpoint into an equally capable GQA model. Weight shapes, conversion and subsequent training must match the architecture. When comparing existing models, do not attribute differences in model size or training data entirely to KV-head count.

Check the actual deployment

Read query-head count, KV-head count, head dimension and cache dtype from the actual model configuration, then calculate the raw budget. Measure peak runtime memory while recording input length, output length, concurrency and cache policy to explain differences from the estimate.

GQA and PagedAttention target different layers: sharing structure versus KV-cache organization. This morning’s LLM quantization article discusses numerical representation and error. Check structure, storage management and bit width separately rather than using one label for every memory optimization.

中文版

Sources

GQA, EMNLP 2023; Peer-reviewed paper; arXiv submission history.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*