Beyond Next-Token Prediction

Training a model to predict the next few tokens at once — not just the immediate one — gives a denser learning signal, better data efficiency, a nudge toward planning, and a path to faster generation. The objective changes; the model does not have to.

Large Language Models: From Transformers to Frontier Models

Every module so far changed the model's architecture. This one changes its training objective. The whole course has rested on one task: predict the next token. Multi-token prediction (MTP) asks the model to predict the next several tokens at each position during training. It is a small-sounding change with surprisingly deep effects, and it is the third major efficiency idea in modern models — alongside efficient attention and mixture-of-experts.

The change in one picture

In ordinary training, at each position the model produces one prediction — the immediate next token — and the loss compares it to the one true next token. Over a sentence, position 1 is trained to predict token 2, position 2 to predict token 3, and so on. One prediction, one target, per position.

Multi-token prediction widens the target. At each position the model predicts the next kk tokens — say the next 3 — and the loss compares all kk predictions against the kk actual following tokens. Position 1 is now trained to predict tokens 2, 3 and 4; position 2 predicts tokens 3, 4 and 5; and so on. We pick a depth kk (how far ahead to look) and supervise every step against that whole short horizon.

Single-token prediction supervising one next token per position, versus multi-token prediction supervising the next three tokens per position
Single-token training supervises one next token per position; multi-token training supervises the next k. Same data, a richer target at every step.

Nothing about the base model needs to change for this — it is the supervision that is richer. And richer supervision turns out to buy four distinct things.

Benefit 1 — a denser training signal

With single-token prediction, each position teaches the model exactly one fact: what comes immediately next. With depth-kk prediction, each position teaches kk facts at once, about structure reaching several steps ahead. The gradient from each training example is denser and more informative — the model is pushed to encode longer-range structure (grammar, coherence, how a phrase will resolve) directly from every position, instead of only the one-step-ahead relationship. More learning is extracted from the same text.

Benefit 2 — better data efficiency

Because each position carries more signal, the model learns more per token of training data. Measured on standard benchmarks, models trained with multi-token prediction reach higher scores than single-token models given the same amount of data — the effect is especially pronounced on code, and it grows with model size (small models see little or no gain; large ones benefit clearly). Same data, more capability.

Benefit 3 — a nudge toward planning

This one is subtle and elegant. Most next tokens are easy and nearly forced by the previous word; a few are choice points — tokens where the text could genuinely go several ways and which shape everything after. Under multi-token prediction, a consequential token shows up in the target window of several earlier positions (it is "token i+1i+1" for one position, "token i+2i+2" for the one before, and so on). So errors on that token are counted multiple times across the loss. The training therefore places more implicit weight on consequential tokens and less on the inconsequential, filler ones — gently teaching the model to look ahead and get the pivotal decisions right. That is as close as training gets to rewarding planning.

Why 'planning' and not just 'prediction'

A model that only ever predicts one token can be short-sighted: it optimises the immediate word and lets the future fall where it may. Supervising a few steps ahead forces its internal representation at each position to already "know" something about where the text is heading. It is still next-token prediction at heart — but trained to keep the near future in view, which is what makes outputs more coherent over longer spans.

Benefit 4 — a path to faster generation

A model that can propose several future tokens at once opens the door to speculative decoding: draft a handful of tokens quickly, then verify them in a single pass, accepting the ones that are correct. When the drafts are mostly right, several tokens are produced for roughly the cost of one, giving a real generation speed-up. Multi-token prediction provides exactly the machinery speculative decoding needs.

A key practical point

These four benefits do not all have to be used. In practice, a model is often trained with multi-token prediction to reap the training benefits — denser signal, data efficiency, planning — and then, at inference, the extra prediction machinery is simply switched off, so generation reverts to ordinary one-token-at-a-time decoding. The multi-token modules were scaffolding that improved the model during training; they can be discarded afterwards, or kept and repurposed for speculative decoding if the speed-up is wanted. Either way, the model you ship is as good as the richer objective made it.

The next chapter opens up how a model predicts several tokens at once — the small stack of prediction modules that does it, and the one design choice that makes the modern version work better than the original.

EasyMTP

How does the training target differ between single-token and multi-token prediction?

HardMTPplanning

Why does multi-token prediction place more implicit weight on 'choice point' tokens?

MediumMTP

Why can the multi-token prediction machinery be discarded at inference without losing its benefits?