What the Heads Learn — and Whether We Need Them All

Different heads specialise into recognisable jobs, but many are redundant and the keys and values they store dominate memory at generation time. That tension sets up the efficient-attention module.

Large Language Models: From Transformers to Frontier Models

We have built multi-head attention and followed its shapes. This chapter steps back to two questions that matter for the rest of the course: what do the heads actually end up doing, and — since heads are not free — do we need all of them? The second question is the doorway into Module 6.

Heads specialise into jobs

When researchers inspect a trained model's attention, individual heads turn out to take on recognisable, interpretable roles. Commonly observed types include:

  • Positional heads that mostly attend to the previous token, or the one two back — tracking local order.
  • Syntactic heads that link words in grammatical relationships, such as a verb to its subject or object, or a noun to its determiner.
  • Coreference heads that connect a pronoun to the noun it refers to — our it → cat link.
  • Delimiter or "rest" heads that park their attention on a sentence-start token or punctuation, effectively doing little — a kind of no-op the model can fall back on.

No one programs these roles. They emerge because a division of labour lowers the loss: it is easier for the model to dedicate a head to "find the subject" than to make one head do everything. This specialisation is the concrete payoff of the previous chapters — the reason several narrow heads beat one wide one.

But heads are redundant — and expensive

Two uncomfortable facts sit alongside that nice story.

Many heads are redundant. Studies that prune heads from a trained model find that a large fraction can be removed with little loss in quality; often a single head per layer carries most of the weight for a given function. The model spreads capacity generously during training, but the trained result leans on far fewer heads than it has.

Heads are costly at generation time — through their keys and values. Recall from the overview that generating text token by token is made fast by caching each past token's keys and values so they need not be recomputed — the KV cache. Every head stores its own keys and values for every past token in every layer. So the memory cost of generation scales with:

cache size  ∝  (layers)×(heads)×(sequence length)×dk.\text{cache size} \;\propto\; (\text{layers}) \times (\text{heads}) \times (\text{sequence length}) \times d_k.

For long contexts this cache, not the model's weights, becomes the thing that fills the accelerator's memory — and the more heads (specifically, the more distinct key/value sets), the worse it gets.

The tension that drives Module 6

Multi-head attention wants many query patterns for expressiveness, but each distinct key/value set it stores multiplies the generation-time memory. These two pressures pull in opposite directions. The whole family of efficient-attention methods — multi-query, grouped-query and latent attention — is a search for the best trade: keep many queries, but share or compress the keys and values.

The seam between "classic" and "modern"

This is a natural place to pause and see the course's shape. Modules 2–4 built the transformer essentially as it was introduced: embeddings, a stack of blocks, and multi-head scaled, masked attention. It is complete and it works.

But we have just exposed two loose threads that a production system cannot ignore:

  1. Position is still missing — attention is order-blind, and we have been quietly assuming it somehow knows order. Module 5 fixes that with positional encodings.
  2. Keys and values are expensive at generation time, and heads are partly redundant. Module 6 exploits exactly that redundancy.

Everything from Module 5 onward is, in one way or another, an engineering response to a limitation we can now name precisely. That is the difference between knowing what a transformer is and knowing what it takes to make one fast, long-context and affordable.

EasyMulti-head

Give two examples of specialised roles that attention heads are observed to take on.

MediumInference

Why does the KV cache grow with the number of heads, and why does that matter for long contexts?

MediumInference

Multi-head attention needs many query patterns but pays for every distinct key/value set it stores. How does this tension motivate the next efficiency methods?