BLEU and ROUGE

Perplexity scores the model's probabilities; these two score what it actually wrote. BLEU asks how much of the output was right, ROUGE asks how much of the reference was covered — and both are blind to meaning.

Large Language Models: From Transformers to Frontier Models

Notice what perplexity never did: it never once looked at anything the model wrote. It measures the probabilities a model assigns to text someone else produced, which tells you a great deal about its grasp of language and nothing whatsoever about the quality of its answers. The moment you care about output a human being will read — a translation, a summary, a reply — you need a different kind of measurement, one that compares what the model actually wrote against what it should have written.

The two classical metrics for that are BLEU and ROUGE. They work the same way — count overlapping word sequences with a reference — and differ in the direction they count, which turns out to matter a great deal.

Counting n-grams against a reference

Both metrics compare a candidate (the model's output) to one or more references (text a human produced) by counting shared n-grams: contiguous runs of nn words. Unigrams catch whether the right words appear; bigrams and longer catch whether they appear in the right arrangement.

The difference is the denominator, and it encodes a different question:

  • BLEU is precision: of the n-grams the model produced, how many were in the reference? How much of what it said was right?
  • ROUGE is recall: of the n-grams in the reference, how many did the model produce? How much of what should have been said was covered?
Two overlapping sets of n-grams, candidate and reference, with precision highlighted as the overlap over the candidate set and recall as the overlap over the reference set
The same overlap, two denominators. BLEU divides by what the model produced; ROUGE divides by what the reference contains.

BLEU, and the two tricks that keep it honest

BLEU was built for machine translation. Its raw form — the fraction of candidate n-grams found in the reference — is naive enough to be gamed in two obvious ways, and the fixes for both are the parts worth understanding.

Clipping. Plain precision rewards repetition. If the reference contains "the" twice, a candidate reading "the the the the the the" scores a perfect unigram precision, since every word it produced does appear in the reference. BLEU therefore uses modified precision: each n-gram's match count is capped at the number of times it occurs in the reference. Six copies of "the" against a reference containing two earns two matches, not six. Without clipping the metric is meaningless; with it, repetition stops paying.

Brevity penalty. Precision also rewards saying less. A one-word output that happens to be correct scores 100%, which is not a good translation of a sentence. BLEU multiplies the score by a brevity penalty that reduces it when the candidate is shorter than the reference, and leaves it alone when the candidate is as long or longer. Length is handled by penalty rather than by recall, which keeps BLEU precision-shaped.

Worked through on a short pair:

Reference: the cat is on the mat Candidate: the cat sat on the mat

The candidate has six unigrams. Five of them — "the" twice, "cat", "on", "mat" — appear in the reference within their clipped limits; only "sat" does not. Modified unigram precision is therefore 5/6≈0.8335/6 \approx 0.833. Both sentences are six words long, so the brevity penalty is 1, and BLEU-1 is 0.8330.833. One substituted word, one-sixth off the score.

Full BLEU does this for several n-gram sizes at once — conventionally 1 through 4 — and combines them:

BLEU=BP×exp⁡ ⁣(∑n=1Nwnlog⁡pn)\text{BLEU} = \text{BP} \times \exp\!\left( \sum_{n=1}^{N} w_n \log p_n \right)

where pnp_n is the modified precision at order nn, wnw_n are weights that usually sum to one, and BP is the brevity penalty. That expression is a geometric mean of the precisions, and the choice is deliberate: a geometric mean is zero if any term is zero. A candidate with good word choice but no matching 4-grams has not assembled those words into the right phrases, and BLEU drives its score to zero rather than averaging the failure away. Every n-gram order must contribute something.

ROUGE, and why summaries need recall

ROUGE was built for summarisation, where the failure that matters is different. A summary that omits the main point is bad even if every word it contains is impeccable — so the question becomes coverage, and the denominator becomes the reference.

Reference: Alice won the contest Candidate: Alice won the contest yesterday

All four reference unigrams appear in the candidate, so ROUGE-1 recall is 1.0. The candidate added a word, so precision is 4/5=0.84/5 = 0.8, and the F1 — the harmonic mean — is about 0.890.89. In practice ROUGE is reported as all three, because recall alone can be gamed by writing a long summary that contains everything, and the F1 keeps that in check.

The family has several members:

  • ROUGE-N — n-gram overlap, most often ROUGE-1 and ROUGE-2.
  • ROUGE-L — based on the longest common subsequence: the longest sequence of words appearing in both texts in the same order, though not necessarily adjacently. This rewards preserved ordering without demanding contiguous matches, so a correct summary that inserts an extra clause is not punished the way a bigram count would punish it.
  • ROUGE-S and ROUGE-W — skip-bigram and weighted-LCS variants, which loosen or re-weight what counts as a match.

When several references exist, the usual convention is to score against each and take the best. Different people summarise differently; matching any one valid summary well is success, not something to be averaged down by the others.

Precision or recall is a question about the task

Neither metric is the better one — they are answers to different questions, and the task picks the question.

In translation, adding content is the serious error. A translation that invents a clause is wrong in a way that a slightly terse one is not, so precision leads and length is handled by a penalty. In summarisation, omission is the serious error. A summary that drops the conclusion has failed at its job, so recall leads, with precision reported alongside to discourage padding.

BLEUROUGE
DirectionPrecision — of what was producedRecall — of what was expected
Built forTranslationSummarisation
PunishesInvented content, repetition, brevityOmitted content
Length handlingBrevity penaltyReported with precision and F1
Combining ordersGeometric mean across n-gram sizesVariants reported separately

What both metrics cannot see

Everything above counts word overlap, and that is the ceiling. Both metrics are surface matchers that need reference text, and the consequences are serious enough that they rule out whole categories of use.

Paraphrase is punished. "I'm fine, thanks" and "I'm doing well, thank you" share almost no n-grams and mean the same thing. Any system that expresses the right content in different words is marked down. Synonyms, reordering, a different register — all cost score while costing no quality.

One reference cannot cover a valid space. For translation there are many acceptable renderings, and for an open question there are thousands. Writing a handful of references cannot enumerate them, so a good answer that happens not to resemble the one on file scores badly.

There is no notion of meaning. The metrics match strings. They cannot tell a true statement from a false one assembled from the same vocabulary, which means a confident hallucination that reuses reference wording can score well. Nothing in either metric checks whether the output is correct.

There is no notion of coherence. Logical flow, consistency across paragraphs, whether an argument holds together — invisible. The counts are local, and the quality of long text is not.

The pressures can be gamed in opposite directions. BLEU rewards being conservative and terse; ROUGE rewards being comprehensive and long. Optimising a system against either one directly tends to produce exactly the distortion it rewards.

For translation and summarisation against good references, these metrics remain useful: they are fast, deterministic, cheap enough to run on every build, and correlate reasonably with human judgement over a large corpus. For dialogue, creative writing, factual question answering, reasoning, or anything where many different answers are right, they do not work — the reference-matching assumption has simply broken down. That gap is what the next lesson addresses, by replacing the reference with a human preference.

MediumEvaluationBLEU

Why does BLEU clip n-gram counts, and what breaks without it?

HardEvaluationBLEU

Why does BLEU combine n-gram precisions with a geometric rather than an arithmetic mean?

MediumEvaluationROUGEBLEU

Why is BLEU standard for translation and ROUGE for summarisation?

MediumEvaluationlimitations

Why do BLEU and ROUGE break down for open-ended dialogue?