This is the chapter that separates candidates who have deployed something from candidates who have studied. A model is trained once and served forever, so inference is the cost that matters, and almost every architectural choice in a modern model exists to reduce it. Interviewers ask here because the answers are quantitative and hard to fake.
What gets asked
Why generation is slow without a cache. Producing each new token means attending over everything before it. Done naively, every step recomputes keys and values for the entire prefix, making total work quadratic in the output length. The KV cache stores them, so each step computes only the new token's vectors — turning quadratic into linear.
The cost of the cache. You traded computation for memory, and the memory is substantial. The size is the product of layers, batch size, heads times head dimension, sequence length, two (keys and values) and the bytes per number. At frontier scale with a long context, that can exceed the model's own weights. The follow-up that rewards preparation: this is why long-context requests are priced higher.
Shrinking the cache. Multi-query attention has all heads share one set of keys and values — smallest cache, some quality loss. Grouped-query attention shares within groups — most of the saving, little of the loss, which is why it is the common choice. Latent attention compresses to a smaller representation and expands on use.
Mixture-of-experts. Replace the feed-forward layer with many experts and route each token to a couple. Total parameters grow, active parameters per token do not. Know the distinction between those two numbers and the engineering problem that comes with it — keeping experts balanced so routing does not collapse onto a few.
Low precision. Fewer bits per number means less memory and faster arithmetic. Know the formats and the range-versus-precision trade, mixed precision, and the outlier problem — one extreme value setting the scale for a whole tensor destroys the small values, which is why scaling is done per block.
Compression. Pruning removes weights; distillation trains a smaller model to imitate a larger one. Know which is cheap, which is lossy, and which needs data.
The follow-ups that catch people
Estimate the KV cache for this model. A genuine question, and the arithmetic is small. Practise it once with real numbers so you are not deriving the formula under pressure.
Why does unstructured pruning often give no speed-up? Wanted: a stored zero is still a number. Unless the hardware can skip it, memory and arithmetic are unchanged. Real gains need patterns the hardware supports, or structured removal that actually shrinks the tensors.
Quantization, pruning or distillation? Wanted: quantization first — cheapest, no data needed. Pruning if the hardware rewards the sparsity pattern. Distillation when you need a large latency reduction and have data, since only a genuinely smaller model delivers that. And they compose.
Why not use multi-query attention everywhere, since the cache is smallest? Wanted: collapsing to one key-value set removes head diversity and costs quality; grouped-query keeps most of the memory saving with much less loss.
How it gets worded
Caching, sparsity and precision
- "Why is generation slow without a cache, and what exactly does caching keys and values buy you?"
- "Estimate the cache for this model at this context length. Now tell me why long-context requests cost more."
- "Multi-query, grouped-query, latent — what does each give up to shrink the cache?"
- "What is the difference between a mixture-of-experts model's total and active parameters, and what breaks if routing is left to itself?"
- "Fewer bits per weight: what do you gain, and what is the outlier problem?"
Flash attention
- "How does flash attention get its speed-up? Is it an approximation, or the same answer computed differently?"
- "Attention is quadratic. Does flash attention change that? If not, what does it change?"
- "Take me through the memory-hierarchy argument — where is the bottleneck really?"
- "How is softmax computed when you never hold the full score matrix in memory?"
- "How does it differ from approximate attention schemes, and when would you choose a sparse pattern instead?"
- "What hardware does it need, and how do you turn it on in a real serving stack?"
- "What changed between the first version and the later ones?"
- "When does flash attention buy you nothing at all?"
- "Write a memory-efficient attention kernel from scratch — talk me through the blocking."
Distillation
- "What is knowledge distillation, and what problem are you using it to solve?"
- "How does a student learn from a teacher? What do soft targets carry that hard labels do not?"
- "What is temperature doing in the distillation loss?"
- "What can be distilled besides the output distribution?"
- "Quantize, prune or distil — how do you choose, and can you do all three?"
- "How does distillation differ from transfer learning?"
- "Can a student beat its teacher? Under what conditions?"
- "How would you collapse an ensemble into one deployable model?"
- "The teacher is mediocre, or biased. What reaches the student?"
- "Design a pipeline where the student has a different architecture from the teacher."
- "Where does distillation show up in recent open-weight releases?"
- "Where does distillation fail?"
Reading path
From Large Language Models, modules 6 to 9:
- The KV Cache — the mechanism, the formula, and the memory problem. If you read one lesson here, this one.
- Multi-Query Attention and Grouped-Query Attention — the two ends and the sensible middle.
- Multi-Head Latent Attention — the compression approach, for depth.
- Why Mixture-of-Experts, Routing and the Gating Network, Keeping the Experts Balanced — sparsity and its engineering.
- Why Low Precision and Making Low Precision Work — formats, mixed precision, outliers.
- Pruning and Distillation — the other two compression levers, and when each applies.
- Anatomy of a Frontier Model — all of it assembled, which is a good revision pass.
Beyond Next-Token covers multi-token prediction, worth reading if speculative decoding comes up.
The framing that works in interviews
Group the techniques by what they change, and the whole area becomes one answer rather than six:
- Recompute less — the KV cache.
- Store less — MQA, GQA, latent attention.
- Activate less — mixture-of-experts.
- Use fewer bits — quantization.
- Have fewer weights — pruning and distillation.
Then add the sentence that makes it a senior answer: these compose, and the right order is cheapest-and-safest first. Quantize, then prune if the hardware rewards it, and only distil when nothing else gets you to the latency target.
The flash-attention family — restructuring the attention computation to avoid writing the full score matrix to memory — belongs in this chapter and does not yet have a lesson here. It is worth knowing that it is exact, not an approximation: the same result, computed with far less memory traffic.