"How would you evaluate this?" is asked at the end of almost every system design round, and it is where otherwise strong candidates stop short. Building something is the visible half of the job; knowing whether it works is the half that decides whether it ships.
What gets asked
Perplexity. The exponentiated average negative log-likelihood on held-out text — the same quantity training minimises, measured on data the model has not seen. Read it as an effective branching factor: perplexity 20 means the model is about as uncertain as choosing among 20 options per token. The floor is 1, not 0, since it is the exponential of a non-negative quantity. Two limits get asked: it depends entirely on the tokenizer, so models with different vocabularies cannot be compared; and it only measures local prediction, saying nothing about coherence or truth.
BLEU and ROUGE. Both count overlapping n-grams against a reference; they differ in the denominator. BLEU is precision — of what the model produced, how much was in the reference — with clipped counts to stop repetition gaming it and a brevity penalty to stop terseness gaming it. ROUGE is recall — of what the reference contains, how much the model covered. BLEU for translation, where adding wrong content is the serious error; ROUGE for summarisation, where omission is.
Why they break for open-ended work. Both need reference text and both only see surface overlap. Paraphrase is punished, a correct answer phrased differently scores badly, and a confident falsehood built from reference vocabulary scores well. For dialogue, explanation or anything with many valid answers, the assumption has failed.
Preference-based ranking. When there is no reference, the signal left is which of two outputs a person prefers. Elo turns thousands of pairwise votes into one ranking: models start at a baseline, each comparison produces an expected score from current ratings, and the update is proportional to how surprising the result was. A rating gap converts directly into a win probability, which is what makes it readable.
RAG evaluation. Retrieval and generation must be measured separately. Retrieval metrics ask whether the right passage was in the returned set at all; generation metrics ask whether the answer was grounded in what was retrieved. A combined score hides which half failed.
The follow-ups that catch people
Why can't you compare perplexity across two models? Wanted: tokenization. The metric averages over tokens, and different vocabularies produce different numbers of differently-hard predictions.
Does Elo need every model to play every other? Wanted: no. Evidence propagates — if A beats B and B beats C, the ratings already encode something about A versus C — so sampled matchups between similarly-rated models converge with far fewer comparisons than a round robin.
Humans are biased. Why not use automatic metrics instead? Wanted: for open-ended generation the automatic metrics measure the wrong thing. The response is to engineer around human bias — multiple raters, randomised presentation order, anonymised models, down-weighting unreliable voters — not to substitute a precise measurement of something irrelevant.
A model scores well on your benchmark and users complain. What happened? Wanted: the benchmark is not measuring what users care about; possible contamination; distribution mismatch between evaluation set and real traffic. This is a judgement question and a good one.
How would you evaluate a summarisation feature with no reference summaries? Wanted: pairwise human preference, an LLM judge calibrated against human ratings, and task-specific checks — did it retain the key entities, is it faithful to the source.
How it gets worded
Perplexity
- "How do you measure whether a language model is any good at all?"
- "Define perplexity and show me how it is computed. Why is lower better, and why does it bottom out at 1 rather than 0?"
- "What is the relationship between cross-entropy loss and perplexity?"
- "Can you compare perplexity across two models with different tokenizers?"
- "A model has excellent perplexity. Does it write well?"
- "How would you compute perplexity for a bidirectional model rather than a causal one?"
- "What are the limits of reporting perplexity on its own?"
Overlap metrics
- "Explain how BLEU is computed. Why does it need a brevity penalty, and why are the counts clipped?"
- "BLEU or ROUGE — what separates them, and which one for translation, which for summarisation?"
- "What is ROUGE-L seeing that ROUGE-1 is not?"
- "Implement BLEU."
- "The output is correct and worded completely differently from the reference. How does BLEU score it, and what does that tell you about overlap metrics in general?"
- "Several answers are valid. How do you handle multiple references — and what do you do when there are none at all?"
Everything else
- "How would you evaluate a translation system? A summariser? A chatbot?"
- "What has replaced these metrics for modern models, and how does a preference ranking turn votes into a number?"
- "The benchmark score went up and users complained. What happened?"
Reading path
- Perplexity — what evaluation means, the metric, and its four limits.
- BLEU and ROUGE — the overlap metrics, the clipping and brevity tricks, and where they stop working.
- Elo Ratings for Language Models — preference ranking, and the comparison table of all three.
- Retrieval Metrics — measuring the retrieval half on its own.
- Evaluating the Whole Pipeline — end-to-end measurement, including judging groundedness.
- Tracing and Debugging — the observability underneath all of it.
The answer that works in a design round
Three things, in this order.
Separate the stages. Say what you would measure at each — retrieval quality, then generation quality, then the end-to-end outcome. A single number cannot tell you which part failed.
Name an offline and an online measure. Offline: a held-out evaluation set with known-good answers, run on every change. Online: user-visible signals — thumbs, escalation to a human, task completion, follow-up rate.
Say what you would do about disagreement. When the offline score improves and user satisfaction does not, the evaluation set is not representative. Mentioning this unprompted is the strongest version of this answer, because it is the actual experience of running these systems.
And carry the honest point forward from the metrics themselves: every one of them judges the model from outside. None can distinguish a model that has learned the task from one that has learned what scores well — which is why interpretability is the subject of the next chapter.