Sinusoidal Positional Encodings

Replace the binary staircase with smooth sine and cosine waves whose frequencies slow down across the dimensions. The famous formula is just that idea — and the sine/cosine pairing hides a rotation that makes relative position easy.

Large Language Models: From Transformers to Frontier Models

We ended the last chapter with a precise wish: keep binary's oscillations at graded frequencies, but make every signal smooth. The answer is to use sine and cosine waves, and the result is the sinusoidal positional encoding introduced in the original transformer paper. Its formula looks forbidding the first time you meet it — but after the last two chapters, every piece of it will make sense.

The formula, read slowly

For a token at position pp, the positional vector's value at dimension index ii is:

PE(p,2i)=sin⁡ ⁣(p10000 2i/d),PE(p,2i+1)=cos⁡ ⁣(p10000 2i/d).PE(p, 2i) = \sin\!\left(\frac{p}{10000^{\,2i/d}}\right), \qquad PE(p, 2i+1) = \cos\!\left(\frac{p}{10000^{\,2i/d}}\right).

Do not let it intimidate you. It depends on exactly the two variables from the binary picture: the position pp, and the dimension index ii. Everything else is detail:

  • It is a sine (or cosine), so the value is smooth and bounded between −1-1 and 11. That is our continuity wish granted — no more staircase.
  • The index ii sits in the denominator of the angle. Rewrite the angle as ωi p\omega_i \, p with frequency ωi=1/10000 2i/d\omega_i = 1/10000^{\,2i/d}. Because ii is in the denominator, low indices give high frequency (fast oscillation) and high indices give low frequency (slow oscillation) — exactly the binary structure, now with smooth waves instead of flipping bits.
  • The 10000 is just a scale knob. It spreads the frequencies over a useful range so that, across the hundreds of dimensions, oscillations fade gradually from very fast to nearly flat. The exact value is an empirical choice, not a law.

So the scary formula is simply: each dimension is a wave; the waves slow down as you move to higher dimensions. If you ever lose the thread, picture the binary table and replace each column's bit-flips with a smooth wave of the matching speed.

Several sine waves stacked, each of lower frequency than the one above, representing positional encoding dimensions from low index (fast) to high index (slow)
Each dimension is a wave; low indices oscillate fast, high indices slowly. It is the binary staircase made smooth — the same multi-frequency code, now differentiable.

Why a sine and a cosine? The hidden rotation

One detail still looks arbitrary: why pair a sine on even indices with a cosine on the odd one next to it? This is the most beautiful part, and it is the bridge to everything that follows.

Think of each such pair as the two coordinates of a point in a plane: (cos⁡θ,sin⁡θ)(\cos\theta, \sin\theta), where θ=ωi p\theta = \omega_i\, p. That point sits on a circle, and its angle is set by the position pp. Now ask: if I move from position pp to position p+kp+k, what happens to the point? Its angle becomes ωi(p+k)=θ+ωik\omega_i(p+k) = \theta + \omega_i k. The point simply rotates by a fixed angle that depends only on the shift kk, not on where you started.

That is a remarkable property. It means the encoding of a nearby position is a rotation of the current one. The model does not have to memorise each position in isolation; to relate position pp to position p+kp+k, it only needs to apply the rotation for kk. In the paper's words, the scheme lets the model attend by relative position easily — which was requirement 4 from two chapters ago, and the hardest one to satisfy.

And none of this would work with a lone sine. Rotating a point needs both a cosine (the x-coordinate) and a sine (the y-coordinate). That is the entire reason the formula pairs them. The sine/cosine pairing is not decoration — it is what makes relative position a clean rotation.

Two intuitions, one for each end of the spectrum

The fast, low-index waves resolve small shifts — the difference between a word at position 3 and the same word at position 4. The slow, high-index waves barely change over short distances, so they preserve a sense of long-range relationship between far-apart tokens. Together they describe position at every scale at once — fine and coarse — just as the binary bits did.

The flaw that motivates the next step

Sinusoidal encoding is smooth, bounded, generalises to new lengths, and makes relative position a rotation. It powered a wave of early language models. So why does nearly every modern model use something else?

Because of where it is applied. Like integer and binary encoding before it, the sinusoidal vector is added to the token embedding at the very start, before the stack. Even though its values are modest, adding anything to the embedding mixes position into the vector that is supposed to carry meaning — it dilutes the semantics, if only a little, and that mixture then rides through every layer.

Two dissatisfactions crystallise from this, and they are the seeds of rotary encoding:

  1. Position is injected in the wrong place. The spot where position actually matters is inside attention — where queries meet keys to decide who attends to whom. Why not inject it there, in the queries and keys, and leave the token embeddings pure?
  2. Adding is the wrong operation. Adding a vector changes a vector's length as well as its direction. But we just discovered that position is naturally a rotation — and a rotation changes direction while leaving length untouched. Why add at all, when we could rotate?

Follow those two ideas — inject position inside attention, and do it by rotating rather than adding — and you arrive directly at rotary positional encoding, the subject of the next chapter.

Try it yourself
Transformer Lab: sinusoidal encodings →

See the fast and slow waves across dimensions and how similar nearby positions are.

MediumPositional

In the sinusoidal formula, why does the dimension index i appear in the denominator of the angle?

HardPositional

Why does the formula pair a sine with a cosine instead of using sine alone?

MediumPositional

Sinusoidal encoding works well, so why do modern models move away from adding it to the embeddings?