We ended the last chapter with a precise wish: keep binary's oscillations at graded frequencies, but make every signal smooth. The answer is to use sine and cosine waves, and the result is the sinusoidal positional encoding introduced in the original transformer paper. Its formula looks forbidding the first time you meet it — but after the last two chapters, every piece of it will make sense.
The formula, read slowly
For a token at position , the positional vector's value at dimension index is:
Do not let it intimidate you. It depends on exactly the two variables from the binary picture: the position , and the dimension index . Everything else is detail:
- It is a sine (or cosine), so the value is smooth and bounded between and . That is our continuity wish granted — no more staircase.
- The index sits in the denominator of the angle. Rewrite the angle as with frequency . Because is in the denominator, low indices give high frequency (fast oscillation) and high indices give low frequency (slow oscillation) — exactly the binary structure, now with smooth waves instead of flipping bits.
- The 10000 is just a scale knob. It spreads the frequencies over a useful range so that, across the hundreds of dimensions, oscillations fade gradually from very fast to nearly flat. The exact value is an empirical choice, not a law.
So the scary formula is simply: each dimension is a wave; the waves slow down as you move to higher dimensions. If you ever lose the thread, picture the binary table and replace each column's bit-flips with a smooth wave of the matching speed.
Why a sine and a cosine? The hidden rotation
One detail still looks arbitrary: why pair a sine on even indices with a cosine on the odd one next to it? This is the most beautiful part, and it is the bridge to everything that follows.
Think of each such pair as the two coordinates of a point in a plane: , where . That point sits on a circle, and its angle is set by the position . Now ask: if I move from position to position , what happens to the point? Its angle becomes . The point simply rotates by a fixed angle that depends only on the shift , not on where you started.
That is a remarkable property. It means the encoding of a nearby position is a rotation of the current one. The model does not have to memorise each position in isolation; to relate position to position , it only needs to apply the rotation for . In the paper's words, the scheme lets the model attend by relative position easily — which was requirement 4 from two chapters ago, and the hardest one to satisfy.
And none of this would work with a lone sine. Rotating a point needs both a cosine (the x-coordinate) and a sine (the y-coordinate). That is the entire reason the formula pairs them. The sine/cosine pairing is not decoration — it is what makes relative position a clean rotation.
The fast, low-index waves resolve small shifts — the difference between a word at position 3 and the same word at position 4. The slow, high-index waves barely change over short distances, so they preserve a sense of long-range relationship between far-apart tokens. Together they describe position at every scale at once — fine and coarse — just as the binary bits did.
The flaw that motivates the next step
Sinusoidal encoding is smooth, bounded, generalises to new lengths, and makes relative position a rotation. It powered a wave of early language models. So why does nearly every modern model use something else?
Because of where it is applied. Like integer and binary encoding before it, the sinusoidal vector is added to the token embedding at the very start, before the stack. Even though its values are modest, adding anything to the embedding mixes position into the vector that is supposed to carry meaning — it dilutes the semantics, if only a little, and that mixture then rides through every layer.
Two dissatisfactions crystallise from this, and they are the seeds of rotary encoding:
- Position is injected in the wrong place. The spot where position actually matters is inside attention — where queries meet keys to decide who attends to whom. Why not inject it there, in the queries and keys, and leave the token embeddings pure?
- Adding is the wrong operation. Adding a vector changes a vector's length as well as its direction. But we just discovered that position is naturally a rotation — and a rotation changes direction while leaving length untouched. Why add at all, when we could rotate?
Follow those two ideas — inject position inside attention, and do it by rotating rather than adding — and you arrive directly at rotary positional encoding, the subject of the next chapter.
See the fast and slow waves across dimensions and how similar nearby positions are.