Multi-Query Attention

The boldest way to shrink the cache: let all heads share a single set of keys and values. The cache drops by a factor equal to the number of heads — but the diversity that made multi-head attention powerful collapses with it.

Large Language Models: From Transformers to Frontier Models

The last chapter left us with a clear target. The cache size scales with n⋅hn \cdot h — a key and value stored for every head. If we could store keys and values for fewer heads, the cache would shrink in direct proportion. The first idea in this direction is the most aggressive one, and it is worth studying precisely because its flaw teaches us what the later methods must protect.

The idea: all heads share one set of keys and values

Recall multi-head attention: each of the nn heads has its own learned projections, so each produces its own queries, keys and values and attends in its own way. The cache therefore holds a separate key and value set for every head.

Multi-query attention (MQA) makes a drastic simplification. Keep a separate query for each head — queries are cheap, they are not cached — but give all heads a single shared set of keys and values. Every head still asks its own question, but they all consult the same keys and blend the same values.

The effect on the cache is immediate and large. Instead of storing keys and values for all nn heads, we store them for one. The per-head factor nn in the cache formula collapses to 1:

cache size  ∝  n⋅hn=h.\text{cache size} \;\propto\; \frac{n \cdot h}{n} = h.

For a model with 128 heads, that is a 128-fold reduction. The hundreds-of-gigabytes cache from the last chapter shrinks to a few gigabytes. As a pure memory play, it is spectacular, and it speeds up generation too, since there is far less data to move and fewer key/value projections to compute.

Three side-by-side diagrams: multi-head attention with distinct per-head keys/values, multi-query with one shared key/value set, and grouped-query with per-group shared sets
Multi-head (every head its own K/V), multi-query (all heads share one K/V), and — next chapter — grouped-query (heads share K/V within groups). Colour = distinct key/value content.

Why it hurts: diversity collapses

So why not stop here? Because MQA quietly undoes the very reason multi-head attention exists.

Recall why we wanted many heads: so that different heads could attend in different ways at once — one tracking the subject, another a distant referent, another local syntax. That diversity came from each head having its own keys and values, letting it define its own notion of what is relevant.

MQA forces every head to share one set of keys and values. The queries still differ, so the heads are not identical — but the raw material they attend over is now the same for all of them. The space of distinct "views" the heads can form shrinks sharply. In effect we have kept the many questions but given them all the same reference book to consult. The model's ability to capture multiple relationships at once drops, and with it, quality.

The trade-off in one line

MQA shrinks the cache by the number of heads and loses a meaningful amount of accuracy. It sits at one extreme: smallest cache, weakest attention. Plain multi-head attention sits at the other: largest cache, strongest attention. The useful designs live in between — which is exactly the next chapter.

What this teaches us

MQA frames the whole problem as a spectrum with two quantities in tension:

  • Cache size, set by how many distinct key/value sets we store.
  • Quality, set by how much the heads can differ in what they attend over.

Push all the way toward sharing and you win on memory but lose on quality. The art is to give up far less quality for most of the memory saving — and there are two ways to do that. One is to share partially, in groups, which is grouped-query attention, the next chapter. The other, more radical, is to stop thinking in terms of "how many heads share" altogether and instead compress the keys and values into a small shared form from which each head can still reconstruct its own — which is latent attention, the climax of this module. Both are reactions to the lesson MQA just taught: shrinking the cache by crushing head diversity costs too much.

EasyMQA

How does multi-query attention shrink the KV cache, and by how much?

MediumMQA

Why does sharing keys and values across all heads reduce model quality, even though the queries still differ?

MediumMQA

MQA and plain multi-head attention sit at opposite extremes of a trade-off. What are the two quantities in tension, and where does each method land?