Self-Attention: Queries, Keys and Values

Give each token three learned roles — a query, a key and a value. Match every query against every key to get weights, then blend the values by those weights. That is self-attention, in one formula.

Large Language Models: From Transformers to Frontier Models

The last chapter gave the intuition: each token asks a question, compares it against what the others offer, and pulls back a blend of their contents. Now we make "question", "offer" and "contents" concrete. They become three vectors per token — the query, the key, and the value — each produced by its own learned matrix.

First, why not just compare the embeddings?

Before introducing three matrices, it is worth seeing why the obvious cheap idea fails — because that failure is the whole reason the matrices exist.

The simplest way to score how related two tokens are is to take the dot product of their embeddings: large when the vectors point the same way, small when they do not. So to decide how much a token should attend to another, just dot their embedding vectors and be done. Why not?

Consider the sentence "the dog chased the ball but could not catch it." The final word, it, refers to the ball, not the dog — so when the model processes it, it should attend mostly to ball. But the embedding of it is a fixed vector, and so are the embeddings of dog and ball. Dotting it with dog and it with ball gives two numbers that reflect only the generic, context-free similarity of those words — and there is no reason that would single out ball. The plain dot product has no way to express "in this sentence, as a thing that gets caught, the ball is the relevant one." It measures similarity, but the relationship we need is subtler than similarity, and it changes with context.

When we cannot write down the rule for a relationship, deep learning offers a standard move: don't hand-design it — add trainable weights and let training discover it. That is precisely what the query, key and value matrices are. Rather than comparing the raw embeddings, we first pass each token through learned projections that are free to pull out just the aspects that matter for attending, and compare those instead. The matrices start random and are shaped by gradient descent until the comparison lands where it should — on ball rather than dog.

Three projections of the same vector

Start from a token's vector hh (its current entry in the residual stream). We compute three new vectors from it, using three learned weight matrices WQW_Q, WKW_K, WVW_V:

q=WQ h,k=WK h,v=WV h.q = W_Q\,h, \qquad k = W_K\,h, \qquad v = W_V\,h.
  • The query qq encodes what this token is looking for.
  • The key kk encodes what this token offers to others — how it advertises itself.
  • The value vv encodes the content this token will hand over if attended to.

"Self"-attention means all three come from the same sequence: every token produces a query, a key, and a value, and the tokens attend to each other. (Later, in cross-attention, queries come from one sequence and keys/values from another — but the machinery is identical.)

Step 1 — Score every pair with a dot product

How well does token ii's query match token jj's key? Use the dot product, which is large when two vectors point the same way:

score(i,j)=qi⋅kj.\text{score}(i, j) = q_i \cdot k_j.

For a sequence of nn tokens this gives an n×nn \times n table of scores: row ii holds how much token ii's query matches every token's key. This is the "compare my question to every offer" step, done for all tokens at once.

Step 2 — Turn scores into weights with softmax

The scores are raw numbers. To blend, we need weights that are positive and sum to 1, so we apply softmax across each row:

αij=e score(i,j)∑j′e score(i,j′).\alpha_{ij} = \frac{e^{\,\text{score}(i,j)}}{\sum_{j'} e^{\,\text{score}(i,j')}}.

Now αij\alpha_{ij} is the attention weight: the fraction of its attention that token ii pays to token jj. Each row sums to 1 — token ii distributes a total of 100% of its attention across all tokens, most of it on the keys its query matched best.

Step 3 — Blend the values

Finally, token ii's output is the weighted sum of values, using those weights:

outi=∑jαij vj.\text{out}_i = \sum_{j} \alpha_{ij}\, v_j.

A token it attends to strongly contributes most of its value; a token it ignores contributes almost nothing. This blended vector — token ii, enriched with exactly the context its query sought — is added back into the residual stream.

The whole thing in one line

Stacking the queries, keys and values for all tokens into matrices QQ, KK, VV (one row per token), the three steps collapse to the compact formula at the heart of the transformer:

Attention(Q,K,V)=softmax ⁣(QK⊤)V.\text{Attention}(Q, K, V) = \text{softmax}\!\left(Q K^{\top}\right) V.

Read it right to left in words: score every query against every key (QK⊤QK^{\top}), normalize each row into weights (softmax), blend the values (multiply by VV). (We will add one small scaling factor inside the softmax in the next chapter — it matters for stability but not for the idea.)

One query compared by dot product against all keys, softmax turning the scores into weights, and those weights blending the value vectors into one output
Self-attention for one token: its query scores against every key, softmax makes the scores into weights, and the weights blend the values into the token's enriched output.

Why three separate roles

A natural question: why not just use the token's vector directly for all three? Because the same token plays different parts depending on the direction of the interaction. What a token looks for (query) is generally different from what it advertises (key), which is different again from the content it contributes (value). Giving each role its own learned matrix lets the model tune them independently — and, crucially, WQW_Q, WKW_K, WVW_V are trained by the ordinary gradient descent and backpropagation you already know; attention adds no new training machinery, only a new wiring of matrix multiplies and a softmax.

Keep the three verbs

Query = what I want. Key = what I offer. Value = what I give. Score queries against keys, softmax into weights, blend the values. Almost every attention variant later in the course changes how the keys and values are stored or shared — but this query–key–value skeleton never changes.

The next chapter explains the one missing detail — why we scale the scores before the softmax — and the chapter after adds the rule that a token may only look at the tokens before it.

Try it yourself
Transformer Lab: self-attention →

Tap a token to see what it attends to, and watch “it” switch between “animal” and “street”.

EasyAttention

What do the query, key and value of a token each represent?

MediumAttention

In softmax(QKᵀ)V, what does each of the two matrix products contribute?

MediumAttention

Why give a token three separately learned projections instead of using its vector directly for querying, keying and valuing?