The previous chapter's formula was almost complete. The real attention formula has one extra ingredient — a division inside the softmax:
where is the dimension of the query and key vectors. This factor looks like an arbitrary fudge, but it fixes a real numerical problem. It is worth understanding, because the same reasoning ("keep the numbers going into a softmax from getting too large") recurs throughout deep learning.
The problem: dot products grow with dimension
A score is a dot product of a query and a key: . It is a sum of terms. If the individual components of and are roughly independent with a typical size around 1, then:
- the sum of such terms has a typical magnitude that grows like .
So with the scores are spread over a range around ; with , around . The bigger the head dimension, the bigger the raw scores — purely as an artifact of adding up more terms, not because the match is any more meaningful.
Why large scores hurt: softmax saturates
Now feed large scores into softmax. Softmax exponentiates, so if one score is much larger than the others, its dominates and the output collapses toward a one-hot distribution — essentially all the weight on a single token, near 0 everywhere else.
That has two bad effects:
- It attends too sharply, too early. Before the model has learned anything useful, an accident of large dot products can make attention nearly deterministic, blocking it from considering a spread of tokens.
- Gradients vanish. Where softmax is saturated (outputs near 0 or 1), its slope is almost flat. As the activation functions lesson in the Deep Learning course shows for the sigmoid, a saturated nonlinearity passes almost no gradient back. The query and key matrices then barely learn.
The fix: divide by √dₖ
Dividing the scores by exactly cancels the growth we identified: if the raw sum scales like , dividing by brings the typical score magnitude back to around 1, independent of the head dimension. Softmax then starts in a sensible, non-saturated range, attention begins soft and diffuse, and gradients flow. As training proceeds the model can learn to make certain scores large where sharp attention is genuinely warranted — but it is no longer forced into it by dimensionality.
We divide by the standard deviation of the score, which grows like , not by its variance, which grows like . Dividing by normalizes the spread to a constant; dividing by would over-shrink the scores and make attention uselessly flat. The square root is the right power precisely because a sum of independent terms has standard deviation proportional to .
The complete mechanism
With scaling in place, scaled dot-product attention is the finished single-head operation:
- Project each token to a query, key and value.
- Score every query against every key: .
- Scale by to keep the scores well-sized.
- Softmax each row into attention weights.
- Blend the values by those weights.
This is the exact operation repeated, in parallel copies, inside every attention layer of every model in this course. The next chapter adds the one rule that makes it suitable for generating text: a token may attend only to the tokens that come before it.
Grow the head dimension and watch unscaled softmax collapse onto one key.