We have built multi-head attention and followed its shapes. This chapter steps back to two questions that matter for the rest of the course: what do the heads actually end up doing, and — since heads are not free — do we need all of them? The second question is the doorway into Module 6.
Heads specialise into jobs
When researchers inspect a trained model's attention, individual heads turn out to take on recognisable, interpretable roles. Commonly observed types include:
- Positional heads that mostly attend to the previous token, or the one two back — tracking local order.
- Syntactic heads that link words in grammatical relationships, such as a verb to its subject or object, or a noun to its determiner.
- Coreference heads that connect a pronoun to the noun it refers to — our
it→catlink. - Delimiter or "rest" heads that park their attention on a sentence-start token or punctuation, effectively doing little — a kind of no-op the model can fall back on.
No one programs these roles. They emerge because a division of labour lowers the loss: it is easier for the model to dedicate a head to "find the subject" than to make one head do everything. This specialisation is the concrete payoff of the previous chapters — the reason several narrow heads beat one wide one.
But heads are redundant — and expensive
Two uncomfortable facts sit alongside that nice story.
Many heads are redundant. Studies that prune heads from a trained model find that a large fraction can be removed with little loss in quality; often a single head per layer carries most of the weight for a given function. The model spreads capacity generously during training, but the trained result leans on far fewer heads than it has.
Heads are costly at generation time — through their keys and values. Recall from the overview that generating text token by token is made fast by caching each past token's keys and values so they need not be recomputed — the KV cache. Every head stores its own keys and values for every past token in every layer. So the memory cost of generation scales with:
For long contexts this cache, not the model's weights, becomes the thing that fills the accelerator's memory — and the more heads (specifically, the more distinct key/value sets), the worse it gets.
Multi-head attention wants many query patterns for expressiveness, but each distinct key/value set it stores multiplies the generation-time memory. These two pressures pull in opposite directions. The whole family of efficient-attention methods — multi-query, grouped-query and latent attention — is a search for the best trade: keep many queries, but share or compress the keys and values.
The seam between "classic" and "modern"
This is a natural place to pause and see the course's shape. Modules 2–4 built the transformer essentially as it was introduced: embeddings, a stack of blocks, and multi-head scaled, masked attention. It is complete and it works.
But we have just exposed two loose threads that a production system cannot ignore:
- Position is still missing — attention is order-blind, and we have been quietly assuming it somehow knows order. Module 5 fixes that with positional encodings.
- Keys and values are expensive at generation time, and heads are partly redundant. Module 6 exploits exactly that redundancy.
Everything from Module 5 onward is, in one way or another, an engineering response to a limitation we can now name precisely. That is the difference between knowing what a transformer is and knowing what it takes to make one fast, long-context and affordable.