The last chapter left us with two dissatisfactions and the ideas to cure them: inject position inside attention, in the queries and keys, and do it by rotating rather than adding. Follow both and you get rotary positional encoding, or RoPE — the scheme used by most current large language models. It is a deeply visual idea, so we will build it from the picture rather than the algebra.
The setup: rotate query and key vectors by an angle set by position
Recall where position actually bites: inside attention, when a query meets a key to produce a score, . That is the one place the model decides who attends to whom, so that is where position belongs. RoPE leaves the token embeddings completely untouched and instead modifies each query and key vector just before they are compared.
How does it modify them? By rotation. Take a query vector and split its components into pairs: , , and so on. Each pair is the two coordinates of a point in a plane. RoPE rotates each pair by an angle proportional to the token's position:
A token at position turns each pair by a small angle; a token at turns it fifty times as far. The frequencies are the same graded set we have seen since the binary table: the first pairs spin fast as position grows, later pairs spin slowly. The key and query are rotated the same way, and only then is the dot product taken.
Why rotation is exactly the right move
Two properties fall out, and they are precisely the fixes we wanted.
The magnitude never changes. Rotation moves a vector around a circle — it changes direction, never length. So RoPE writes position into the angle of the query and key while leaving their size exactly as the model computed it. Nothing of the token's learned content is overwritten; position is added in a dimension (orientation) that was free to use. Contrast this with adding a vector, which disturbs both length and direction.
The embeddings stay pure. Because the rotation happens on queries and keys inside attention, the token embeddings flow into the stack carrying only meaning, undiluted by position. We have moved the position injection from the wrong place (the input) to the right place (the comparison).
The payoff: attention sees relative position for free
Here is the elegant consequence. When a query at position is compared with a key at position , each has been rotated by an angle proportional to its own position. Because of how rotations compose, their dot product ends up depending only on the difference — how far apart the two tokens are — not on their absolute positions. The rotations of the individual vectors cancel down to a relative one.
This is exactly requirement 4 from the start of the module, delivered cleanly: the attention score between two tokens automatically reflects how far apart they are. And it brings a graceful side effect often called long-term decay — because far-apart tokens are rotated by very different angles, their query–key alignment tends to fall off with distance, matching the intuition that distant words are usually less related.
The two frequency scales, intuitively
The fast and slow pairs each do a recognisable job — the same division of labour we have met at every stage:
- Fast pairs (low index) capture small shifts. Compare "I just told her the truth" with "I told just her the truth." The word told moves by one position and the meaning changes; the fast-rotating pairs, which turn sharply even for a one-step move, register that small shift.
- Slow pairs (high index) capture long-range links. In "Einstein developed the theory of relativity; this breakthrough reshaped physics," this breakthrough refers back many words. The slow-rotating pairs barely change over short spans, so they keep a stable relationship across long distances — the fast pairs would have spun too far to be useful there.
So RoPE encodes position at every scale at once, just as the binary bits and the sine waves did, but now in a form that preserves magnitude, keeps embeddings clean, and hands attention relative distance directly.
Split each query and key into coordinate pairs; rotate every pair by an angle equal to position times a per-pair frequency; then do attention as usual. Lengths are preserved, embeddings stay pure, and query–key scores end up depending on the distance between tokens. That is the whole idea.
Why this is the modern default
RoPE was introduced several years after the original sinusoidal scheme, and it has largely replaced it — Llama, and most other recent open models, use it. The reasons are everything above: it keeps the semantic channel clean, it is parameter-free (no learned position table to run out of), it extends gracefully to longer sequences, and relative position comes built in. It is also the piece we will need in Module 6: the most advanced attention variant there, latent attention, has to be carefully reconciled with RoPE, and you now understand both halves.
With position handled, the base transformer is genuinely complete — it can read order and predict text. Everything from the next module on is about making it fast, long-context, big, and cheap enough to be a frontier system, starting with the memory crisis that long contexts create at generation time.
Rotate a query and a key by position and shift both: the score depends only on their distance.