A model does not output text. It outputs a probability distribution over its vocabulary, and something else turns that into a token. Interviewers ask about this because it is the part engineers actually touch — temperature and top-p are parameters in every API call — and because many candidates use them daily without being able to say what they do.
What gets asked
From vector to word. The final layer projects the last hidden vector onto the vocabulary, producing one logit per token. Softmax turns logits into probabilities. A sampling rule picks one. That token is appended and the whole thing repeats. Be able to narrate this loop.
Temperature. Divides the logits before the softmax. Below 1 sharpens the distribution toward the most likely tokens; above 1 flattens it toward uniform; at 0 it reduces to always taking the maximum. The useful framing is that temperature does not add knowledge or creativity — it only changes how much probability mass the sampler is willing to look at.
Top-k and top-p. Truncation rules applied before sampling. Top-k keeps the k most likely tokens; top-p (nucleus) keeps the smallest set whose probabilities sum past p. The follow-up is why top-p is usually preferred: it adapts to the shape of the distribution, keeping few candidates where the model is confident and more where it is uncertain, whereas a fixed k is wrong in one direction or the other.
Greedy decoding. Always take the highest-probability token. Deterministic and often worse than it sounds, because a locally optimal token can lead into a poor continuation.
When to use which. Deterministic settings for extraction, classification, structured output and anything a program parses; sampling for open-ended writing. This is a judgement question and it is common.
The follow-ups that catch people
Does temperature 0 guarantee identical outputs? Wanted: it makes the sampling deterministic, but reproducibility in practice can still be affected by batching and floating-point non-determinism in the serving stack. Candidates who say "yes, always" are usually corrected.
Why is greedy decoding not simply the best choice? Wanted: each step is locally optimal, which does not make the sequence optimal — the highest-probability first token can commit you to a worse overall continuation.
Your model emits malformed JSON occasionally. What do you change? Wanted: lower the temperature or go deterministic first, then constrained decoding or schema validation with a retry. Discussing prompt wording before sampling settings misses the cheaper fix.
What does the model's probability tell you about correctness? Wanted: very little. Confidence is not accuracy — the model is confident about fluent continuations, not true ones, which is the hallucination argument.
How it gets worded
- "Compare greedy decoding, beam search and sampling."
- "Walk me through beam search step by step. What does widening the beam change, and why is wider not simply better?"
- "Implement a small beam search over log-probabilities."
- "Why does beam search prefer short outputs, and how do you correct for it?"
- "Why does it fall into repetition loops?"
- "When would you choose a beam search over nucleus sampling — and why do chat models almost never use one?"
- "What is the trade between searching more widely and producing something a person wants to read?"
- "Explain what temperature does to the logits, and what top-p does that top-k cannot."
- "An endpoint returns malformed JSON one time in fifty. Which knob do you reach for first?"
- "Does temperature zero guarantee you the same output twice?"
Reading path
- From Logits to Tokens — the whole pipeline from final vector to sampled word, including temperature and top-k/top-p.
- What a Language Model Is — if the objective is not fresh, this grounds why the distribution looks as it does.
- The KV Cache — what makes this loop fast enough to be practical, and the subject of the next chapter.
- Perplexity — the same probabilities, viewed as a measurement rather than a sampler.
The Transformer Lab has a sampling tab. Moving temperature and top-p while watching the candidate distribution change is the clearest way to internalise the difference between them.
What is not covered yet
Beam search — keeping several candidate sequences alive and expanding the most promising, rather than committing to one token at a time — does not yet have a lesson on this site. It still comes up, particularly for translation and other tasks with one right answer, so it is worth reading about elsewhere. The short version: it searches for a high-probability sequence rather than a high-probability next token, which helps where there is a single correct output and hurts for open-ended generation, where it produces bland, repetitive text.
Speculative decoding — using a small model to draft tokens that a large model verifies in parallel — appears in efficiency discussions. The multi-token prediction lesson covers the idea that makes it work.