We have built the whole base transformer — embeddings, attention, multiple heads — and quietly got away with a lie. We kept saying the model "knows" that one word comes before another. It does not. Left to itself, attention is blind to order, and this chapter is where we face that and fix it.
Attention sees a bag, not a sequence
Look again at the attention formula, . Every token's query is compared against every token's key, and the results are summed. Nothing in that computation refers to where a token sits. If you shuffled the input tokens, you would shuffle the rows of , and in the same way — and get exactly the same set of outputs, just reordered. Attention treats its input as an unordered bag of tokens.
For language, that is a disaster, because order is meaning:
the dog chased the cat the cat chased the dog
Same five tokens, opposite meaning. A model that cannot tell these apart cannot understand language at all.
A sharper example
The trouble is not only between sentences; it bites inside a single one. Take:
the dog chased another dog
There are two occurrences of dog. The embedding lookup returns the same vector for both — the table knows the word, not its place. So both dog tokens enter attention identically, undergo identical operations, and come out with identical enriched vectors. The model literally cannot distinguish the dog doing the chasing from the dog being chased. (You can verify this with a real model: strip out positions, run self-attention, and the output vectors for two copies of the same word are bit-for-bit equal.)
That is intolerable. We need the first dog and the second dog to leave attention as different vectors, because they play different roles. The only thing that differs between them is position, so position is the information we must supply.
What a position signal must do
Before reaching for a solution, let us be first-principles about the job. A scheme for encoding position should:
- Give different positions distinguishable signals. Position 2 and position 5 must end up different, or we are back to the identical-dogs problem.
- Stay bounded. Whatever we add must not swamp the token's meaning. Embeddings carry semantics in numbers clustered near zero; a position signal that grows huge with sentence length would drown them out.
- Generalise across lengths. The signal for "position 5" should mean the same thing in a short sentence and a long one, and ideally keep working past the lengths seen in training.
- Make relative position easy to use. What usually matters is not that a word is at absolute position 428, but that it is three words after another. A good scheme lets the model recover "how far apart" cheaply.
Those four requirements are a surprisingly tight brief, and the next three chapters are the story of living up to them. We will try the obvious thing (just use the position number), watch it fail requirement 2, fix that with a binary trick that reveals a deep idea, smooth it into the famous sinusoidal encoding, and finally arrive at rotary encodings (RoPE) — the scheme almost every current model uses, which nails requirement 4 in a beautiful way.
In the classic transformer, position is injected right at the start: after the embedding lookup, a positional vector is added to each token's embedding before the stack. Keep that location in mind — part of the journey to RoPE is realising this is not the best place to do it, and moving the operation inside attention instead.