Perplexity

The oldest measure of a language model asks one question: how surprised was it by text it had never seen? Perplexity turns that surprise into a single number — and it is the same number the model was trained to minimise.

Large Language Models: From Transformers to Frontier Models

You can now build a model, make it fast and cheap, and adapt it to a task. What we have not addressed once is how you would know whether any of it is any good. That is an uncomfortable gap, and this module closes it. We will look at three ways of measuring a language model — and the honest story is that each one answers a different question and each one is blind to something the others catch, which is exactly why people argue about benchmarks as much as they do.

Start with the oldest and most fundamental.

What evaluating a model means

Evaluation is a pipeline: take data the model has never seen, run it through, and compute a number from what comes out.

The held-out part is not a formality. A model has memorising capacity in the billions, so its performance on training data says almost nothing about its performance in the world. Scores are only meaningful on a test set kept apart from training, compared against ground truth — the text that actually followed, or a reference answer a person wrote.

Two broad families of measurement exist, and the distinction is worth keeping:

  • Intrinsic metrics score the model at the language-modelling task itself: how well it predicts text. Perplexity is the canonical example.
  • Extrinsic metrics score it on something you want done — answering questions, summarising a document, writing code that runs.

They can disagree sharply. A model can predict text beautifully and be useless at your task, which is why no serious evaluation relies on one number. This module covers the intrinsic measure first, then two ways of judging actual outputs.

Perplexity: the exponent of the training loss

An autoregressive model does exactly one thing: given the tokens so far, it produces a probability distribution over what comes next. So a natural question is — on real text, what probability did it assign to the token that actually appeared?

Average the log of that probability over a test set, negate it, and exponentiate:

PPL=exp⁡ ⁣(−1N∑i=1Nlog⁡p(wi∣w<i))\text{PPL} = \exp\!\left( -\frac{1}{N} \sum_{i=1}^{N} \log p(w_i \mid w_{<i}) \right)

where NN is the number of tokens, wiw_i is the ii-th token, w<iw_{<i} is everything before it, and p(wi∣w<i)p(w_i \mid w_{<i}) is the probability the model gave the correct token.

Read it from the inside out. log⁡p\log p of the right token is high (close to zero) when the model was confident and correct, and very negative when it was caught out. Averaging gives the typical log-probability per token; negating makes a lower number better; exponentiating converts it back from log space.

That inner quantity — average negative log-likelihood — is cross-entropy, which is exactly what training minimises. Perplexity is the exponentiated training loss, measured on held-out data. The metric and the objective are the same quantity in different clothing, which is why validation perplexity is the standard thing to watch while a model trains: if it is falling, the model is genuinely getting better at predicting language; when it stops falling, you have stopped learning.

What the number means

The name comes from information theory, and it has a concrete reading: perplexity is an effective branching factor — the number of options the model is effectively choosing between at each token.

A perplexity of 100 means the model is as uncertain, on average, as someone picking uniformly among 100 possibilities. A perplexity of 10 means it has narrowed the field to about ten. Put that way, the direction is obvious: fewer live options means a model that knows better what comes next.

Two next-token probability distributions over a vocabulary: a flat one labelled with a high branching factor and a peaked one labelled with a low branching factor
Perplexity as an effective branching factor. A flat distribution leaves many options live and scores high; a peaked distribution on the token that actually appeared scores low.
The floor is 1, not 0

A common slip is to assume perfect means zero. Perplexity is exp⁡\exp of a non-negative quantity, and exp⁡(0)=1\exp(0) = 1, so the best possible perplexity is 1 — reached only if the model assigns probability exactly 1 to every correct token, with no uncertainty anywhere. That is not achievable on natural language, where many continuations are genuinely valid; real text has irreducible uncertainty. Strong general models sit in the low tens on standard corpora, and that is excellent rather than mediocre.

Why it stayed the default

Four practical properties keep perplexity in use despite everything it misses.

It matches the objective exactly. No proxy, no reference answers, no judgement calls — the thing being measured is the thing being optimised.

It is cheap. One forward pass over the test text gives every log-probability needed; the rest is arithmetic. There is no generation, no sampling, no beam search, no human in the loop. You can evaluate on millions of tokens for the cost of reading them once, which matters when the model is large.

It is a single comparable number. Under matched conditions, a model with lower test perplexity is the better predictor, full stop. That makes it the natural reporting metric for architecture and training changes.

It connects to generation. Decoding strategies that search for high-probability continuations are, in effect, searching for low-perplexity outputs — so the metric is measuring the same quantity generation is trying to maximise.

Where it fails

It only sees local prediction. Perplexity averages per-token surprise. It has no view of whether a passage is coherent, factually right, well argued or self-contradictory. A model can score well by being reliably good at the easy, high-frequency parts of language — the function words that make up much of any corpus — while producing text that falls apart over a paragraph.

It depends on the tokenizer. Probabilities are per token, so changing how text is split changes how many predictions there are and how hard each one is. Two models with different vocabularies produce perplexities that are not comparable, and this is the most common way the metric is misused. A perplexity comparison is only valid with the same tokenizer on the same data.

It depends on the test set. Perplexity is measured against a particular distribution of text. The same model may score well on edited prose and badly on conversational or technical text — not because it got worse, but because the text did. There is no absolute scale; the number only means something relative to a fixed corpus.

Low perplexity is not understanding. Much of a strong score comes from statistical regularity, which can be captured without anything resembling comprehension. Nothing in the metric distinguishes a well-predicted true statement from a well-predicted false one.

It does not apply to most tasks. Perplexity needs a distribution over next tokens. For classification, retrieval, extraction or anything judged by its output rather than its probabilities, it is simply the wrong instrument — hence the next two lessons.

Masked models need a different recipe

The formula assumes a left-to-right model producing a chain of conditional probabilities. Masked models, which see context on both sides and predict a hidden token, never produce that chain, so true perplexity is undefined for them.

The substitute is pseudo-perplexity. Mask one token at a time, run the model, record the probability it assigns to the token that was really there, and aggregate those log-probabilities the same way. It gives a comparable feel for how surprised the model is — at a steep price. A sequence of NN tokens needs NN forward passes instead of one, so it is far too slow for large-scale evaluation, and its values should not be compared against ordinary perplexity from an autoregressive model.

MediumEvaluationperplexity

Why is perplexity the exponentiated training loss, and why does that make it convenient?

MediumEvaluationperplexity

Why can't you compare perplexity between two models with different tokenizers?

EasyEvaluationperplexity

What does a perplexity of 20 actually mean, and what is the best possible value?

HardEvaluationperplexity

A model has excellent perplexity but produces incoherent paragraphs. How is that possible?