Every interview for a role that puts a model in front of users reaches this. The questions sound like safety questions and are really mechanism questions: the wrong answers are the ones that locate the problem in training data rather than in what the model was optimised to do.
What gets asked
Why models hallucinate. The common answers — not enough training data, no internet access — are wrong, and saying so is the strong move. Models hallucinate on subjects they have seen extensively. The cause is an objective mismatch: training maximises the likelihood of plausible continuations, not truth. Asked something it cannot answer, the model has no "no result" state; it produces the text that best fits the shape of an answer. It has no mechanism separating "I know this" from "this sounds right", which is why hallucinations are confident.
Refusals, and the dial. A pretrained model does not refuse — refusal is installed by alignment. That leaves two drives, helpfulness and caution, and which wins determines the failure. Helpfulness winning where the model does not know is a hallucination; caution winning where it was not needed is an over-refusal. They are two directions on one dial, not separate bugs.
Jailbreaks. Prompts that get a model to produce what it would normally decline — rephrasing, roleplay, hiding the request in a puzzle. The interesting finding is that during a successful jailbreak the safety response usually does fire; it simply arrives late. Having begun a sentence, the pressure to complete it coherently carries generation forward for several tokens before the refusal takes over, which is why jailbroken outputs often start complying and then reverse. It is a failure of priority and timing, not of missing safety training.
Why jailbreaks are hard to eliminate. Alignment installs statistical associations — inputs resembling these get responses resembling those — rather than rules checked before output. More safety data covers more phrasings; the space of adversarial inputs is unbounded.
Mitigation. Grounding in retrieved sources so recall has something to work from; calibrated uncertainty, so declining is rewarded; verification outside the model, where it can block output rather than compete with it; and monitoring. Note what is not on the list: finding and deleting the responsible component. Fluent continuation and confident phrasing are what produce readable answers at all.
Interpretability. Probing internal states, isolating interpretable features, and intervening causally to test what a component does. Two findings get asked about: models compute intermediate steps they never state, and they plan ahead — the rhyming word is represented in the hidden state several tokens before it appears, and suppressing that representation changes the output.
Faithfulness. The sharpest practical point here. A model's stated reasoning need not reflect the computation that produced its answer. On a problem it can compute, internal traces match the stated steps. On one it cannot, it produces reasoning of the same quality and confidence, asserting intermediate values it never computed — an answer reached some other way with a justification built backwards. The two are indistinguishable from outside.
The follow-ups that catch people
Can you remove the hallucination circuit? Wanted: no. The capabilities involved are the ones doing the useful work, and hallucination emerges from their interaction rather than sitting in a dedicated component.
How would you reduce hallucination in a product? Wanted: retrieval grounding first, then instructing the model to decline when the context does not support an answer, then verification or citation so a user can check, then measurement. An answer that is only "better prompting" is thin.
Are attention weights an explanation? Wanted: no. Attention correlates poorly with influence on the output. Useful for hypotheses; causal intervention is what establishes effect.
If chain-of-thought can be unfaithful, is it useless? Wanted: no — both things are true. The tokens genuinely help the model compute, and they are not a reliable account of how it computed. Helping and describing are different functions.
How it gets worded
- "Why do large language models make things up?"
- "What has hallucination got to do with the training objective?"
- "How does a jailbreak work mechanically? What is going on inside the model while it complies?"
- "Describe the conflict that happens during a successful jailbreak."
- "How do competing objectives inside a model produce an output nobody asked for?"
- "What is a refusal, and how is it related to hallucination?"
- "Why can't you locate the circuit responsible and remove it?"
- "What is the difference between safety learned in training and a rule applied outside the model?"
- "What has interpretability research told us about why these failures happen?"
- "How does understanding a model's internals help you build a safer product?"
- "What does interpretability contribute to alignment that behavioural testing cannot?"
- "How would you reduce hallucination in something you are shipping next month?"
Reading path
- Hallucinations and Jailbreaks — the objective mismatch, the helpfulness-caution dial, competing internal pressures, and the jailbreak timing finding.
- Model Interpretability — the six technique families, what has been found, and faithfulness in full.
- Why Retrieval — the main practical mitigation, from the other direction.
- Model Alignment — what installs refusal in the first place.
- Evaluating the Whole Pipeline — measuring groundedness, which is how you find out whether any of it worked.
Closing the loop
This chapter is the end of the map, and it is worth seeing why it sits last. Everything earlier was about making a model work: train it, make it efficient, adapt it, build a system around it, measure it. This chapter is about the fact that all of those judge the model from outside — and that a system which behaves well on everything you measured can still fail in ways your measurements could not see.
Interviewers ask about hallucination and jailbreaks partly for the mechanism and partly to find out whether you treat them as embarrassing edge cases or as the predictable consequence of what these models are. The second framing is the right one, and it is the one that produces sensible engineering: bound the system, ground it, verify its output, and watch it in production.
From here, go back to whichever chapter you could not explain aloud.