The Problem of Order

Attention treats its input as a bag of tokens — scramble the words and the maths is unchanged. Since meaning depends on order, we must inject position by hand. This chapter pins down exactly what a position signal has to do.

Large Language Models: From Transformers to Frontier Models

We have built the whole base transformer — embeddings, attention, multiple heads — and quietly got away with a lie. We kept saying the model "knows" that one word comes before another. It does not. Left to itself, attention is blind to order, and this chapter is where we face that and fix it.

Attention sees a bag, not a sequence

Look again at the attention formula, softmax(QK⊤/dk)V\text{softmax}(QK^{\top}/\sqrt{d_k})V. Every token's query is compared against every token's key, and the results are summed. Nothing in that computation refers to where a token sits. If you shuffled the input tokens, you would shuffle the rows of QQ, KK and VV in the same way — and get exactly the same set of outputs, just reordered. Attention treats its input as an unordered bag of tokens.

For language, that is a disaster, because order is meaning:

the dog chased the cat the cat chased the dog

Same five tokens, opposite meaning. A model that cannot tell these apart cannot understand language at all.

A sharper example

The trouble is not only between sentences; it bites inside a single one. Take:

the dog chased another dog

There are two occurrences of dog. The embedding lookup returns the same vector for both — the table knows the word, not its place. So both dog tokens enter attention identically, undergo identical operations, and come out with identical enriched vectors. The model literally cannot distinguish the dog doing the chasing from the dog being chased. (You can verify this with a real model: strip out positions, run self-attention, and the output vectors for two copies of the same word are bit-for-bit equal.)

That is intolerable. We need the first dog and the second dog to leave attention as different vectors, because they play different roles. The only thing that differs between them is position, so position is the information we must supply.

What a position signal must do

Before reaching for a solution, let us be first-principles about the job. A scheme for encoding position should:

  1. Give different positions distinguishable signals. Position 2 and position 5 must end up different, or we are back to the identical-dogs problem.
  2. Stay bounded. Whatever we add must not swamp the token's meaning. Embeddings carry semantics in numbers clustered near zero; a position signal that grows huge with sentence length would drown them out.
  3. Generalise across lengths. The signal for "position 5" should mean the same thing in a short sentence and a long one, and ideally keep working past the lengths seen in training.
  4. Make relative position easy to use. What usually matters is not that a word is at absolute position 428, but that it is three words after another. A good scheme lets the model recover "how far apart" cheaply.

Those four requirements are a surprisingly tight brief, and the next three chapters are the story of living up to them. We will try the obvious thing (just use the position number), watch it fail requirement 2, fix that with a binary trick that reveals a deep idea, smooth it into the famous sinusoidal encoding, and finally arrive at rotary encodings (RoPE) — the scheme almost every current model uses, which nails requirement 4 in a beautiful way.

Where position gets added — for now

In the classic transformer, position is injected right at the start: after the embedding lookup, a positional vector is added to each token's embedding before the stack. Keep that location in mind — part of the journey to RoPE is realising this is not the best place to do it, and moving the operation inside attention instead.

Two identical 'dog' token vectors entering self-attention and producing identical output vectors, illustrating that without position the model cannot tell them apart
Without a position signal, two copies of the same token are processed identically and leave attention identical. Order has to be injected by hand.
MediumPositional

Why does shuffling the input tokens leave the set of attention outputs essentially unchanged?

EasyPositional

In 'the dog chased another dog', why do the two ' dog' tokens come out of attention identical without positional information?

MediumPositional

Why must a positional signal stay bounded rather than simply grow with position?