Every module so far changed the model's architecture. This one changes its training objective. The whole course has rested on one task: predict the next token. Multi-token prediction (MTP) asks the model to predict the next several tokens at each position during training. It is a small-sounding change with surprisingly deep effects, and it is the third major efficiency idea in modern models — alongside efficient attention and mixture-of-experts.
The change in one picture
In ordinary training, at each position the model produces one prediction — the immediate next token — and the loss compares it to the one true next token. Over a sentence, position 1 is trained to predict token 2, position 2 to predict token 3, and so on. One prediction, one target, per position.
Multi-token prediction widens the target. At each position the model predicts the next tokens — say the next 3 — and the loss compares all predictions against the actual following tokens. Position 1 is now trained to predict tokens 2, 3 and 4; position 2 predicts tokens 3, 4 and 5; and so on. We pick a depth (how far ahead to look) and supervise every step against that whole short horizon.
Nothing about the base model needs to change for this — it is the supervision that is richer. And richer supervision turns out to buy four distinct things.
Benefit 1 — a denser training signal
With single-token prediction, each position teaches the model exactly one fact: what comes immediately next. With depth- prediction, each position teaches facts at once, about structure reaching several steps ahead. The gradient from each training example is denser and more informative — the model is pushed to encode longer-range structure (grammar, coherence, how a phrase will resolve) directly from every position, instead of only the one-step-ahead relationship. More learning is extracted from the same text.
Benefit 2 — better data efficiency
Because each position carries more signal, the model learns more per token of training data. Measured on standard benchmarks, models trained with multi-token prediction reach higher scores than single-token models given the same amount of data — the effect is especially pronounced on code, and it grows with model size (small models see little or no gain; large ones benefit clearly). Same data, more capability.
Benefit 3 — a nudge toward planning
This one is subtle and elegant. Most next tokens are easy and nearly forced by the previous word; a few are choice points — tokens where the text could genuinely go several ways and which shape everything after. Under multi-token prediction, a consequential token shows up in the target window of several earlier positions (it is "token " for one position, "token " for the one before, and so on). So errors on that token are counted multiple times across the loss. The training therefore places more implicit weight on consequential tokens and less on the inconsequential, filler ones — gently teaching the model to look ahead and get the pivotal decisions right. That is as close as training gets to rewarding planning.
A model that only ever predicts one token can be short-sighted: it optimises the immediate word and lets the future fall where it may. Supervising a few steps ahead forces its internal representation at each position to already "know" something about where the text is heading. It is still next-token prediction at heart — but trained to keep the near future in view, which is what makes outputs more coherent over longer spans.
Benefit 4 — a path to faster generation
A model that can propose several future tokens at once opens the door to speculative decoding: draft a handful of tokens quickly, then verify them in a single pass, accepting the ones that are correct. When the drafts are mostly right, several tokens are produced for roughly the cost of one, giving a real generation speed-up. Multi-token prediction provides exactly the machinery speculative decoding needs.
A key practical point
These four benefits do not all have to be used. In practice, a model is often trained with multi-token prediction to reap the training benefits — denser signal, data efficiency, planning — and then, at inference, the extra prediction machinery is simply switched off, so generation reverts to ordinary one-token-at-a-time decoding. The multi-token modules were scaffolding that improved the model during training; they can be discarded afterwards, or kept and repurposed for speculative decoding if the speed-up is wanted. Either way, the model you ship is as good as the richer objective made it.
The next chapter opens up how a model predicts several tokens at once — the small stack of prediction modules that does it, and the one design choice that makes the modern version work better than the original.