Scaling the Scores

One small divisor — √dₖ — sits inside the softmax. Without it, dot products in high dimensions grow large, softmax saturates, and gradients vanish. Here is why that factor is there.

Large Language Models: From Transformers to Frontier Models

The previous chapter's formula was almost complete. The real attention formula has one extra ingredient — a division inside the softmax:

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V,\text{Attention}(Q, K, V) = \text{softmax}\!\left(\frac{Q K^{\top}}{\sqrt{d_k}}\right) V,

where dkd_k is the dimension of the query and key vectors. This factor dk\sqrt{d_k} looks like an arbitrary fudge, but it fixes a real numerical problem. It is worth understanding, because the same reasoning ("keep the numbers going into a softmax from getting too large") recurs throughout deep learning.

The problem: dot products grow with dimension

A score is a dot product of a query and a key: q⋅k=∑m=1dkqmkmq \cdot k = \sum_{m=1}^{d_k} q_m k_m. It is a sum of dkd_k terms. If the individual components of qq and kk are roughly independent with a typical size around 1, then:

  • the sum of dkd_k such terms has a typical magnitude that grows like dk\sqrt{d_k}.

So with dk=64d_k = 64 the scores are spread over a range around ±8\pm 8; with dk=128d_k = 128, around ±11\pm 11. The bigger the head dimension, the bigger the raw scores — purely as an artifact of adding up more terms, not because the match is any more meaningful.

Why large scores hurt: softmax saturates

Now feed large scores into softmax. Softmax exponentiates, so if one score is much larger than the others, its eze^{z} dominates and the output collapses toward a one-hot distribution — essentially all the weight on a single token, near 0 everywhere else.

That has two bad effects:

  1. It attends too sharply, too early. Before the model has learned anything useful, an accident of large dot products can make attention nearly deterministic, blocking it from considering a spread of tokens.
  2. Gradients vanish. Where softmax is saturated (outputs near 0 or 1), its slope is almost flat. As the activation functions lesson in the Deep Learning course shows for the sigmoid, a saturated nonlinearity passes almost no gradient back. The query and key matrices then barely learn.

The fix: divide by √dₖ

Dividing the scores by dk\sqrt{d_k} exactly cancels the growth we identified: if the raw sum scales like dk\sqrt{d_k}, dividing by dk\sqrt{d_k} brings the typical score magnitude back to around 1, independent of the head dimension. Softmax then starts in a sensible, non-saturated range, attention begins soft and diffuse, and gradients flow. As training proceeds the model can learn to make certain scores large where sharp attention is genuinely warranted — but it is no longer forced into it by dimensionality.

Why √dₖ and not dₖ

We divide by the standard deviation of the score, which grows like dk\sqrt{d_k}, not by its variance, which grows like dkd_k. Dividing by dk\sqrt{d_k} normalizes the spread to a constant; dividing by dkd_k would over-shrink the scores and make attention uselessly flat. The square root is the right power precisely because a sum of dkd_k independent terms has standard deviation proportional to dk\sqrt{d_k}.

The complete mechanism

With scaling in place, scaled dot-product attention is the finished single-head operation:

  1. Project each token to a query, key and value.
  2. Score every query against every key: QK⊤QK^{\top}.
  3. Scale by 1/dk1/\sqrt{d_k} to keep the scores well-sized.
  4. Softmax each row into attention weights.
  5. Blend the values by those weights.

This is the exact operation repeated, in parallel copies, inside every attention layer of every model in this course. The next chapter adds the one rule that makes it suitable for generating text: a token may attend only to the tokens that come before it.

Try it yourself
Transformer Lab: why divide by √d →

Grow the head dimension and watch unscaled softmax collapse onto one key.

MediumAttention

Why do the raw attention scores tend to grow as the head dimension dₖ increases?

MediumAttention

What goes wrong if very large scores reach the softmax?

HardAttention

Why divide by √dₖ specifically, rather than by dₖ?