A single attention head computes one set of weights per token — one way of deciding what is relevant. But a token usually needs several kinds of context simultaneously. Take the verb sat in "the cat sat on the mat": to be understood well it wants to know who sat ( cat, its subject), where it sat ( mat, via on), and the grammatical shape of the phrase. A single softmax has to spread its one budget of attention across all of these and inevitably compromises.
Multi-head attention removes the compromise by running several attention operations — heads — in parallel, each free to focus on a different kind of relationship.
The idea: several attentions at once
Instead of one set of projection matrices , the layer has independent sets, one per head. Head number has its own , and therefore computes its own queries, keys, values, scores and blended output — the entire scaled, masked attention from Module 3, run start to finish, times over on the same input.
Because each head has its own learned projections, each can specialise:
- one head might track subject–verb links,
- another might follow adjacent-word syntax,
- another might reach back to a distant referent (our
it→cat), - another might attend to punctuation or sentence boundaries.
They all look at the same tokens, but through different learned lenses, at the same time.
Splitting the budget, not multiplying the cost
Here is the elegant part. Multi-head attention does not make each head as wide as the whole model. Instead the model's dimension is divided among the heads. With heads, each head works in a smaller dimension
So a model of width with heads gives each head a 64-dimensional query, key and value. Each head attends in its own 64-dimensional subspace. The total amount of computation is about the same as one full-width head — we have re-spent the same budget as several narrow, specialised views rather than one wide, unfocused one.
You might expect smaller heads to be weaker. In practice several narrow heads beat one wide head, because attention's bottleneck is not dimension but focus: one softmax can only emphasise one pattern of relevance at a time. Eight softmaxes can emphasise eight patterns at once. Diversity of attention matters more here than the width of any single head.
Combine what the heads found
After all heads produce their blended outputs (each a -dimensional vector per token), the layer must fold them back into a single -dimensional vector to return to the residual stream. It does this in two steps:
- Concatenate the heads' outputs, stacking the vectors of size back into one vector of size .
- Mix them with a final learned output matrix , so that information discovered by different heads can interact rather than sit in separate slots.
The output projection is essential: without it the heads' findings would remain in disjoint sub-blocks of the vector, never combined. With it, the layer produces a single enriched vector that reflects all the heads at once.
The shape to remember
Multi-head attention is one layer that internally fans out into small attentions and fans back in:
- Fan out: project the input into sets of (query, key, value), each of dimension .
- Attend: run scaled, masked attention independently in each head.
- Fan in: concatenate the outputs and mix them with .
The next chapter looks more closely at the bookkeeping of these shapes, and the one after asks what the heads actually end up doing — and whether we need all of them.
Compare four heads with different jobs, and check that the parameter count does not depend on the number of heads.