The last two chapters gave us two extremes. Multi-head attention keeps a separate key/value set per head — best quality, biggest cache. Multi-query attention collapses them to a single shared set — smallest cache, degraded quality. Plotted against each other, these are the endpoints of a line, and the obvious question is whether the sweet spot lies somewhere between them. Grouped-query attention (GQA) is that in-between, and it is what most production open models use today.
The idea: share within groups, not across everything
GQA takes the middle path literally. Rather than giving every head its own keys and values (multi-head) or forcing all heads to share one set (multi-query), it splits the heads into a few groups and lets each group share one key/value set.
Say a model has 32 heads. Multi-head keeps 32 distinct key/value sets; multi-query keeps 1. GQA might form 8 groups of 4 heads each, keeping 8 distinct sets — one per group. Within a group, the four heads share keys and values (like a little multi-query); across groups, the sets differ (like multi-head). Queries, as always, stay per-head.
The cache now scales with the number of groups rather than the number of heads :
The two endpoints are just special cases: is multi-query attention, and (one group per head) is full multi-head attention. is a dial you set.
Why the middle is a good place to be
GQA works because the two things in tension respond differently as you turn the dial:
- Memory falls steeply at first. Going from 32 sets to 8 already cuts the per-head cache factor by four — most of multi-query's saving, captured with only a modest number of groups.
- Quality holds up well. With 8 distinct key/value sets, the heads still command a real variety of "views" — far more than multi-query's single shared set. Two heads in different groups attend over genuinely different material, so the model retains most of multi-head's ability to track several relationships at once.
So a handful of groups buys you most of the memory saving for only a small quality cost. That favourable curve — big memory win, small quality loss — is why GQA, not multi-query, became the default. Llama's models, for instance, adopted grouped-query attention across their sizes.
Think of as choosing how much head diversity to keep. Small (few groups) → smaller cache, less diversity, toward multi-query. Large (many groups) → bigger cache, more diversity, toward multi-head. GQA's success is empirical: in practice you can drop to a modest and barely dent quality, so you pocket most of the memory saving almost for free.
Where it leaves us
GQA is a genuine improvement, and for most models it is enough. But notice what it is, fundamentally: a compromise on the same axis MQA introduced. It still shrinks the cache by making heads share key/value sets — it just shares less brutally. The more groups you keep for quality, the less memory you save; the tension between diversity and cache size is eased, not escaped.
That raises the ambitious question that closes this module: could we get multi-head-level quality and a multi-query-sized cache at the same time — not a compromise between them, but genuinely both? Answering "yes" requires abandoning the "how many heads share" framing entirely and compressing the keys and values into a small shared representation that each head still expands into its own keys and values. That is multi-head latent attention, and it is where we go next.
Change the number of key/value groups and watch the cache shrink.