We ended the last chapter with a loose thread. Latent attention's whole payoff comes from the absorption trick: because the up-projections are fixed weights, they fold into the neighbouring matrices, so attention can act straight on the cached latent. But modern models carry position with rotary encodings (RoPE) from Module 5 — and RoPE is a rotation whose angle depends on the token's position. A position-dependent rotation is not a fixed weight, so it cannot be absorbed. This chapter is how the two are reconciled, a design called decoupled RoPE, and it is the piece that makes latent attention usable in a real model.
Why RoPE breaks absorption
Recall the shape of the absorption trick. The attention score is a query times a key. In latent attention the key is reconstructed from the latent by a fixed up-projection, so the query's projection and the key's up-projection — both fixed — can be multiplied together once, ahead of time, leaving a single matrix that acts directly on the cached latent. No key ever has to be rebuilt at inference.
Now insert RoPE. To give the keys position, we would rotate each key by an angle set by its position before the dot product. That rotation sits between the query projection and the key up-projection. And because it changes with every position, it is not a constant we can fold in. The two fixed matrices are no longer adjacent — a position-dependent rotation stands between them — so they cannot be absorbed. The consequence is exactly what latent attention was built to avoid: we would have to rebuild and cache a full key for every token again, surrendering the whole saving. As the originating work put it plainly, rotary position is incompatible with the key compression — apply it naively and the absorption trick dies.
The fix: split into a content half and a position half
The resolution is almost disarmingly simple: do not make the same vectors carry both jobs. Split each query and each key into two parts:
- A content part, which carries no position. It is produced from the latent exactly as before, so the absorption trick still works on it untouched.
- A small position part, which carries RoPE. It is computed separately, with the rotation applied, and is kept deliberately small.
Because a dot product of two concatenated vectors is just the sum of the dot products of their halves, the attention score cleanly separates:
The first term keeps all of latent attention's magic — compressed cache, no key rebuilding. The second term is a small, ordinary RoPE attention that we simply compute the old way. We give up the absorption trick only on the small position part, and accept its little extra cost, while the bulk of the computation — the content part — stays compressed. That is the whole idea; the name "decoupled" is just this separation of position from content.
Two small, telling details
Two engineering choices show how carefully the position part is kept cheap:
- The position key is shared across all heads. Unlike the content keys — which stay distinct per head, preserving the diversity we fought to keep — the small RoPE key is the same for every head. So the cache gains only one shared position vector per token, not one per head.
- Only two things are cached. During generation the model stores the compact latent (for the content part) and the shared position key (for the RoPE part). Nothing else is rebuilt. The cache is still dramatically smaller than storing per-head keys and values, now with position folded in.
The scoreboard, complete
With decoupled RoPE, latent attention delivers on its promise in a real, position-aware model. Measured against the alternatives, it holds multi-head-level quality — because content keys and values stay per-head distinct — while its per-token cache is a small fraction of multi-head attention's, a reduction on the order of tens-fold at frontier scale. Multi-query and grouped-query attention bought cache savings with accuracy; latent attention, plus this decoupling, keeps the accuracy.
That closes the efficiency-at-inference story. We started with the KV cache (speed, but ballooning memory), tried sharing keys and values (MQA, GQA — memory at the cost of quality), and arrived at compressing them instead (MLA), reconciled with position (decoupled RoPE) to keep it practical. The next module turns from the attention half of the transformer block to the feed-forward half, and to a very different pressure: how to make the model vastly bigger without making every token pay for it — mixture-of-experts.