Grouped-Query Attention

The practical middle ground: split the heads into a handful of groups, and let each group share one set of keys and values. A dial between multi-head quality and multi-query thrift — and the scheme most open models actually ship.

Large Language Models: From Transformers to Frontier Models

The last two chapters gave us two extremes. Multi-head attention keeps a separate key/value set per head — best quality, biggest cache. Multi-query attention collapses them to a single shared set — smallest cache, degraded quality. Plotted against each other, these are the endpoints of a line, and the obvious question is whether the sweet spot lies somewhere between them. Grouped-query attention (GQA) is that in-between, and it is what most production open models use today.

The idea: share within groups, not across everything

GQA takes the middle path literally. Rather than giving every head its own keys and values (multi-head) or forcing all heads to share one set (multi-query), it splits the heads into a few groups and lets each group share one key/value set.

Say a model has 32 heads. Multi-head keeps 32 distinct key/value sets; multi-query keeps 1. GQA might form 8 groups of 4 heads each, keeping 8 distinct sets — one per group. Within a group, the four heads share keys and values (like a little multi-query); across groups, the sets differ (like multi-head). Queries, as always, stay per-head.

The cache now scales with the number of groups gg rather than the number of heads nn:

cache size  ∝  g⋅h,1≤g≤n.\text{cache size} \;\propto\; g \cdot h, \qquad 1 \le g \le n.

The two endpoints are just special cases: g=1g = 1 is multi-query attention, and g=ng = n (one group per head) is full multi-head attention. gg is a dial you set.

Why the middle is a good place to be

GQA works because the two things in tension respond differently as you turn the dial:

  • Memory falls steeply at first. Going from 32 sets to 8 already cuts the per-head cache factor by four — most of multi-query's saving, captured with only a modest number of groups.
  • Quality holds up well. With 8 distinct key/value sets, the heads still command a real variety of "views" — far more than multi-query's single shared set. Two heads in different groups attend over genuinely different material, so the model retains most of multi-head's ability to track several relationships at once.

So a handful of groups buys you most of the memory saving for only a small quality cost. That favourable curve — big memory win, small quality loss — is why GQA, not multi-query, became the default. Llama's models, for instance, adopted grouped-query attention across their sizes.

Reading the dial

Think of gg as choosing how much head diversity to keep. Small gg (few groups) → smaller cache, less diversity, toward multi-query. Large gg (many groups) → bigger cache, more diversity, toward multi-head. GQA's success is empirical: in practice you can drop to a modest gg and barely dent quality, so you pocket most of the memory saving almost for free.

Where it leaves us

GQA is a genuine improvement, and for most models it is enough. But notice what it is, fundamentally: a compromise on the same axis MQA introduced. It still shrinks the cache by making heads share key/value sets — it just shares less brutally. The more groups you keep for quality, the less memory you save; the tension between diversity and cache size is eased, not escaped.

That raises the ambitious question that closes this module: could we get multi-head-level quality and a multi-query-sized cache at the same time — not a compromise between them, but genuinely both? Answering "yes" requires abandoning the "how many heads share" framing entirely and compressing the keys and values into a small shared representation that each head still expands into its own keys and values. That is multi-head latent attention, and it is where we go next.

Try it yourself
Transformer Lab: MHA, GQA, MQA and MLA side by side →

Change the number of key/value groups and watch the cache shrink.

EasyGQA

How does grouped-query attention position itself between multi-head and multi-query attention?

MediumGQA

Why does a small number of groups capture most of the memory saving at little quality cost?

MediumGQA

In what sense is GQA still a compromise rather than a true escape from the trade-off?