Hallucinations and Jailbreaks

A model that invents a citation and a model that is talked past its own safety training are failing in related ways: both are cases of competing internal pressures, where the drive to produce fluent, obliging text wins over the drive to be accurate or to decline.

Large Language Models: From Transformers to Frontier Models

The last module was about measuring models. This one is about watching them fail, and specifically about the two failures that dominate every serious conversation about putting these systems in front of users. A model states something completely false with total confidence. A model that would normally refuse a request is sweet-talked into answering it. At first glance these look like two unrelated problems with two unrelated fixes. They are not. Mechanically they are close cousins, and seeing why is the point of this chapter.

Hallucination is the objective working as designed

A hallucination is output that is fluent, confident, specific and not true: an invented citation, a fabricated biographical detail, a plausible-sounding description of how a system works that no system works like.

The common explanations are wrong in an instructive way. It is not fundamentally a shortage of training data — models hallucinate about subjects they have seen extensively. It is not a lack of internet access — a model with live search can still misread or embellish what it retrieves. The real cause is an objective mismatch.

The model was trained to predict likely continuations. Nothing in that objective mentions truth. Given a question it cannot answer from what it absorbed, the model does not have a "no result" state to fall into — it does what it always does, and produces the continuation that best matches the shape of an answer to that kind of question. An invented citation looks exactly like a real citation, because the model learned what citations look like and not which ones exist.

Put sharply: the model has no mechanism that distinguishes "I know this" from "this sounds right." Both feel the same from inside a next-token predictor. That is why hallucinations are confident; confidence was never a measure of knowledge in the first place.

Refusal is a trained override

A refusal is the model declining — because a request is disallowed, or because it is not sure enough to answer.

A purely pretrained model does not refuse. It was trained to continue text, so it continues. Refusal is added by post-training: the alignment stage from Model Alignment teaches the model that certain requests are met with a decline, and that admitting uncertainty is sometimes preferred to guessing.

That leaves an aligned model carrying two drives at once:

  1. Be helpful — answer the question, follow the instruction, produce something useful.
  2. Be safe and honest — don't fabricate, don't comply with disallowed requests.

Most of the time these agree. When they conflict, one wins, and which one wins determines the failure you see. If helpfulness wins where the model does not know, you get a hallucination. If safety wins where it is not warranted, you get an unnecessary refusal — the over-cautious model that declines a reasonable question. Hallucination and over-refusal are not separate bugs to be fixed independently; they are the two directions of error on one dial.

Competing pressures inside the network

It helps to picture the model's internal state as several influences pushing at once on what the next token should be. This is a simplification — the network has no labelled modules — but interpretability research does find identifiable internal features that behave this way.

Some of the pressures that matter here:

  • Knowledge recall — surfacing what the model actually absorbed about the subject.
  • Fluent continuation — producing coherent, well-formed text that follows naturally from what came before.
  • Confidence framing — how certain the output sounds.
  • Safety and refusal — the trained response to content that should be declined.
Four internal pressures — knowledge recall, fluent continuation, confidence, and safety — pushing on the next-token decision, with outcomes labelled hallucination, good answer, refusal and over-refusal
The next token is produced under several competing pressures. Which failure appears depends on which pressure dominates at that moment.

A hallucination is what you get when fluent continuation and confidence framing outrun knowledge recall: the model produces well-formed, assured text with nothing underneath it. A refusal is what you get when the safety pressure dominates.

None of this is programmed. It emerges from training, where the model absorbed both "questions of this shape are followed by answers" and "requests of this shape are followed by declines," as statistical tendencies rather than rules.

There is no hallucination circuit to delete

The natural response to this picture is to find the component responsible and remove it. It does not work, because the capabilities involved are the same ones that make the model useful. Fluent continuation, pattern completion and confident phrasing are not a defect — they are what produces readable answers at all. You cannot excise "makes things up" without also excising "writes coherent text."

The remedies are therefore systemic rather than surgical: training objectives that reward calibrated uncertainty, so the model is rewarded for declining when it should; grounding the model in retrieved sources so that recall has something to recall from, which is the argument of the RAG course; and verification outside the model. Interpretability explains the dynamics; it does not provide a switch.

Jailbreaks: safety outcompeted, not absent

A jailbreak is a prompt that gets a model to produce something its training would normally refuse — by rephrasing the request so it is not recognised, hiding it inside a roleplay or a puzzle, or exploiting the model's own willingness to be helpful.

The term came from phone hacking, where it meant bypassing the manufacturer's restrictions. Applied to language models from around 2022, early versions were crude — "pretend you have no restrictions" — and grew more sophisticated as models improved, into an ongoing exchange between people finding bypasses and labs closing them.

The interesting finding is what happens internally during a successful one. The safety response is usually not asleep. The model frequently registers the problem — the internal signals associated with refusal do activate. What happens is that another pressure beats them to the output.

The clearest demonstrated case works like this. A user hides a forbidden word as an acrostic — an innocuous-looking list of words whose initials spell it — then asks the model to assemble the initials and explain how to make the thing, with an instruction to answer immediately and not reason step by step. The model decodes the word correctly. It then begins to answer, produces the opening of a genuine instruction, and mid-passage reverses: it breaks off and refuses, saying it cannot continue.

Reading that trace tells you a lot. The model understood the request — it was not fooled about what was being asked. The safety response did fire. But having begun a sentence, the pressure to complete it coherently carried the output forward for a few tokens before the refusal took control. The model said "yes" slightly before it managed to say "no."

Two things follow.

The failure is one of priority and timing, not of missing safety training. Safety training was present and strong — it is why the refusal arrived at all. Coherence training was also present and strong, and for a brief window it won. Models are heavily shaped to finish sentences; text that stops mid-clause is rare in everything they learned from. A jailbreak exploits that by manoeuvring the model into starting the wrong sentence, after which momentum does the rest.

Safety behaves like a vote, not a veto. Alignment installs statistical tendencies: inputs resembling these get responses resembling those. It does not install a rule that is checked before output. So a request phrased unlike anything in the safety training may simply not activate the association strongly enough, soon enough.

This also explains why "train harder" is an incomplete answer. More safety data covers more phrasings, which helps, but the space of adversarial inputs is unbounded and the mechanism remains associative. Proposals that address the mechanism rather than the coverage look different: letting the model abandon a sentence mid-way without the fluency pressure fighting it, raising the priority of the safety signal in the generation loop, or placing verification outside the generating model, where it can block output rather than compete with it.

Why interpretability is the thread here

Everything above is a claim about what happens inside the model, and the reason those claims can be made at all is that researchers can now inspect internal activations rather than just reading the output.

That matters practically. "The model was jailbroken" tells you a failure happened. "The model detected the violation, but the fluency pressure dominated for several tokens before the refusal asserted itself" tells you what to change. The first is a bug report; the second is a design brief. If an internal signal reliably precedes a policy breach, it can in principle be monitored and acted on at generation time — catching the failure as it forms rather than after it is printed.

How that inspection actually works — probing internal states, isolating interpretable features, and tracing which components drive a particular output — is the subject of the next lesson.

MediumHallucination

Why is 'not enough training data' the wrong explanation for hallucination?

MediumHallucinationRefusal

How are hallucination and over-refusal related?

HardJailbreakSafety

During a successful jailbreak, is the model's safety mechanism inactive?

MediumInterpretabilityHallucination

Why can't you fix hallucination by locating and removing the responsible component?