The Shapes of Multi-Head Attention

Following the dimensions — batch, sequence, heads, head-dimension — turns multi-head attention from a diagram into something you could implement. The key trick is that heads are just a reshaped view of one big projection.

Large Language Models: From Transformers to Frontier Models

The previous chapter gave the idea; this one follows the shapes. Tracking dimensions is the difference between understanding multi-head attention as a picture and understanding it well enough to build. There is no new concept here — just careful bookkeeping, and one neat implementation trick.

The dimensions in play

Four numbers describe the data inside an attention layer:

  • BB — the batch size, how many sequences we process at once.
  • nn — the sequence length, the number of tokens.
  • dd — the model dimension, the width of each token's vector.
  • hh — the number of heads, with per-head dimension dk=d/hd_k = d/h.

The input to the layer is a block of shape B×n×dB \times n \times d: for each of BB sequences, nn tokens, each a dd-vector.

One projection, then a reshape

Conceptually each head has its own WQ(i),WK(i),WV(i)W_Q^{(i)}, W_K^{(i)}, W_V^{(i)}. In practice we do not run hh separate small multiplications. We apply one big projection WQW_Q of shape d×dd \times d to get all the queries at once — output shape B×n×dB \times n \times d — and then reshape that last dimension dd into h×dkh \times d_k:

B×n×d  ⟶  B×n×h×dk.B \times n \times d \;\longrightarrow\; B \times n \times h \times d_k.

Nothing is computed differently; we have simply relabelled the dd numbers as "hh heads of dkd_k numbers each." A final transpose puts the heads next to the batch, giving B×h×n×dkB \times h \times n \times d_k — which reads as "B⋅hB \cdot h independent little attention problems, each of nn tokens in dimension dkd_k." The same is done for keys and values.

Heads are a view, not a loop

This is the practical heart of multi-head attention: the heads are a reshaped view of a single projection, not hh separately coded operations. That is why multi-head attention costs about the same as one full-width head and runs just as efficiently on a GPU — it is the same matrix multiplies with one extra axis.

The per-head computation, by shape

With queries, keys and values all shaped B×h×n×dkB \times h \times n \times d_k, the attention from Module 3 runs on the last two axes, identically for every (B,h)(B, h) pair:

  1. Scores QK⊤Q K^{\top}: multiplying n×dkn \times d_k by its transpose dk×nd_k \times n gives an n×nn \times n score grid per head → shape B×h×n×nB \times h \times n \times n.
  2. Scale by 1/dk1/\sqrt{d_k} and add the causal mask (the same lower-triangular mask, broadcast across batch and heads).
  3. Softmax over the last axis, so each of the nn rows sums to 1.
  4. Blend values: multiply the n×nn \times n weights by the n×dkn \times d_k values → back to B×h×n×dkB \times h \times n \times d_k.

Every head has done a full scaled, masked attention, all as batched matrix multiplies.

Fanning back in

Now reverse the reshape. Transpose the heads back beside the head-dimension and merge them:

B×h×n×dk  ⟶  B×n×h×dk  ⟶  B×n×d.B \times h \times n \times d_k \;\longrightarrow\; B \times n \times h \times d_k \;\longrightarrow\; B \times n \times d.

This is the concatenation of the heads — again just a relabelling of axes. Finally apply the output projection WOW_O (shape d×dd \times d), mixing the heads' contributions, to produce the layer's output at shape B×n×dB \times n \times d — exactly the shape we started with, ready to be added back to the residual stream.

The round trip in one view

B×n×d⏟in→project + reshapeB×h×n×dk⏟per-head Q,K,V→attentionB×h×n×dk⏟per-head out→merge+WOB×n×d⏟out\underbrace{B \times n \times d}_{\text{in}} \xrightarrow{\text{project + reshape}} \underbrace{B \times h \times n \times d_k}_{\text{per-head } Q,K,V} \xrightarrow{\text{attention}} \underbrace{B \times h \times n \times d_k}_{\text{per-head out}} \xrightarrow{\text{merge} + W_O} \underbrace{B \times n \times d}_{\text{out}}

The shape you start with is the shape you end with — the layer enriches each token's vector without changing its size, which is exactly what lets you stack these blocks as deep as you like. A count worth carrying: four projections (WQ,WK,WV,WOW_Q, W_K, W_V, W_O), each d×dd \times d, make up the parameters of a multi-head attention layer.

MediumMulti-head

If each head has its own Q/K/V projections, why can the whole thing use one d×d matrix per role plus a reshape?

MediumMulti-head

What is the shape of the attention-score tensor inside a multi-head layer, and what does each axis mean?

EasyArchitecture

An attention layer takes B×n×d in and returns B×n×d out. Why does preserving the shape matter?