In the last chapter we made some fairly specific claims about what goes on inside a model during a hallucination or a jailbreak — that the safety signals do fire but arrive a few tokens too late, that the pressure to finish a sentence wins in the meantime. You would be right to ask how anyone could possibly know that. Those are not guesses read off the output; they come from looking inside the model. This chapter is about how that looking is actually done, and about the genuinely strange things it has turned up.
What interpretability is trying to do
A language model is a function with an enormous number of parameters mapping input text to output text. Given a prompt it produces a continuation, and nothing about that process explains itself. For simpler models, explanation means feature importances or a decision path. For a transformer it means something harder: accounting for the behaviour in terms of neuron activations, attention patterns and the internal representations the model builds as it computes.
The goal is explanations a person can hold: not "these 70 billion numbers produced that token", but "the model recognised the subject as a city, retrieved the associated country, and used it". Behaviour explained in concepts rather than coefficients.
The motivation is practical, in three directions. Debugging — a model fabricates a citation, and you want to know whether it lacked the fact or had it and was overridden. Safety — you want to know whether a model holds representations that would produce harmful output under the right prompt, before a user finds that prompt. Alignment — you want to know what the model is actually optimising for, which its output alone cannot tell you, because a model that has learned to produce approved-looking answers looks identical from outside to one that has learned what you wanted.
It is worth being accurate about the state of the field. Researchers can identify specific circuits, isolate interpretable features, and explain particular behaviours in particular models. Set against hundreds of billions of parameters, this is a small mapped fraction. There is no complete account of how any frontier model works, and claims that one model has been "understood" should be read narrowly — some mechanism, in some layers, for some behaviour. What exists is enough to debug specific failures and check specific properties. It is not a map of the model.
The toolkit
Six families of method, roughly in order of how directly they establish cause.
Activation analysis asks which units respond to what. Run many inputs through the model, record which neurons fire strongly, and look for a pattern in the inputs that excite each one. This surfaces apparent concept detectors — a unit that responds to text in a particular language, to food, to code. It is the cheapest thing to do and the easiest to over-read, since a unit that correlates with a concept is not necessarily computing with it.
Attention visualisation maps which earlier tokens a given position attends to while producing an output. It is useful and routinely misinterpreted.
The intuitive reading — the model attended to this token, therefore this token mattered — does not hold. Attention weights have been shown repeatedly to correlate poorly with actual influence on the output: a token can receive strong attention and barely affect the result, or receive little and be decisive. Attention is one operation among many in a block, and what flows through it is shaped by everything downstream. As a source of hypotheses and a debugging aid it is valuable; as evidence of what drove an output it is not. Gradient-based attribution and causal intervention answer that question properly.
Probing classifiers test whether information is present. Freeze the model, take the hidden states at some layer, and train a small classifier on them to predict a property you care about — the grammatical number of the subject, the sentiment, whether the statement is true. If the probe succeeds, that information is linearly available in the representation. The caveat is the same as ever: available is not used. A probe shows the model could read that property off its own state, not that it does.
Causal interventions are the step that settles it. Change the internal state mid-computation and see whether the output changes. Ablation zeroes a component to test whether it was necessary. Activation patching copies activations from a run on one input into a run on another, isolating which components carry which effect. This is the closest thing in the field to an experiment rather than an observation, and it is why causal methods are the standard of evidence: they measure effect, not correlation.
Sparse autoencoders attack a structural obstacle. Models represent far more concepts than they have neurons, so concepts are stored in superposition — spread across overlapping sets of units, with any individual neuron participating in many unrelated things. That is why single neurons so often look almost-but-not-quite interpretable. A sparse autoencoder is trained to re-express dense activation vectors as a sparse combination of many more features, few active at once. The features that emerge are frequently far more interpretable than the raw neurons, and this has become one of the main ways of getting usable units of analysis out of a model.
Circuit analysis composes the rest: find the minimal subnetwork — specific heads, specific layers, specific features — responsible for a behaviour, and work out mechanically how it implements it.
What looking inside has shown
Models compute intermediate steps they never say. Asked for the capital of the state containing a given city, a model might be expected to retrieve a memorised association. Tracing the computation shows something else: a representation of the city's state appears internally, and the capital is then retrieved from that. The model composes two facts, through an intermediate it never writes down. Nobody built that decomposition; it emerged from training. It is evidence that these models are doing more than lookup — and a reminder that most of their computation is invisible in the output.
Models plan ahead, despite predicting one token at a time. Asked to complete a rhyming couplet, a model writes the second line word by word, which suggests it cannot be aiming at a rhyme it has not reached. Inspecting the hidden state partway through the line shows the representation of the eventual rhyming word already active, several tokens early. The model selected its ending first and steered toward it.
Intervention confirms this is causal rather than coincidental. Suppress that representation, and the model completes the line with a different, still-rhyming word. Inject a different concept, and the line bends fluently toward the injected word instead. The plan is a real object in the model's state, and editing it edits the output.
This deserves precision rather than enthusiasm. The hidden state encodes information about where the generation is going, and that information causally shapes later tokens — planning in a functional sense, since the behaviour satisfies constraints that require foresight. It does not follow that the model deliberates or intends. Training on text full of long-range structure produced circuits that set up compatible representations in advance; that is the mechanism, and it is sufficient to explain the behaviour. The useful statement is the narrow one: next-token prediction does not imply next-token thinking, because the state being carried forward can encode a target the model has not yet reached.
Faithfulness: the explanation and the computation are different objects
This is the finding with the sharpest practical edge.
A faithful explanation reflects the computation that actually produced the answer. An unfaithful one is a plausible account that played no causal role. Models are extremely good at producing the second, because fluent explanation is a pattern they learned thoroughly, and nothing in training required it to match their own internals.
The cleanest demonstration contrasts two problems given to the same model, both asked for step-by-step reasoning.
Given one it can actually do — a square root, a multiplication, a floor — the model states its steps and arrives at the answer, and tracing the internals finds representations corresponding to those intermediate quantities. The reasoning is faithful: the stated steps are the computed steps.
Given one it cannot do — a trigonometric function of a large, arbitrary argument — the model produces reasoning of exactly the same quality and confidence. It states the function's range, applies it, asserts a specific intermediate value, and completes the arithmetic to a final answer. Inside, there is no trace of that computation: no representation of the trigonometric value, nothing that could have produced the number it asserted. It settled on an answer by other means — often latching onto something in the prompt — and constructed a justification backwards. The explanation is a façade. Every step reads correctly and none of them happened.
Note the asymmetry that makes this dangerous: the two outputs are indistinguishable from outside. The unfaithful one is not hesitant or vague. It is just as clean, just as confident, and wrong about its own process.
The pattern is well documented in people. Experiments on patients whose brain hemispheres had been surgically separated found them producing confident, coherent explanations for actions initiated by the hemisphere that could not speak — explanations invented after the fact and sincerely believed. Models may be doing something structurally similar: trained on explanations, they produce explanations, with no mechanism tying the account to the process. Generating a correct report of one's own computation is a different and harder task than generating a plausible one.
How you might check. Internal inspection is the rigorous answer: look for activity corresponding to the claimed steps. That is a research procedure, not something to run per request. Cheaper proxies help in practice. Perturb the inputs — change a number or a condition and see whether the reasoning and conclusion move together as the stated logic requires; an explanation that is decoration tends to survive changes that should have broken it. Check consistency across rephrasings. Where the task allows, verify the answer independently rather than trusting the derivation.
The connection back to the previous lesson is direct. A chain of thought is output, and output is what the model was trained to make plausible. That is why reasoning traces cannot be treated as audit logs, and why chain-of-thought prompting genuinely improving accuracy and chain-of-thought being an unreliable account of the model's process are both true at once. The tokens help the model compute; they do not thereby describe the computation.
Why this matters for the rest of the stack
Interpretability changes what can be said about a failure. "The model hallucinated" is an observation. "The model held the correct fact but the fluency pressure dominated" is a diagnosis, and it implies different fixes than "the model never had the fact". Likewise for safety: an internal signal that reliably precedes a policy breach is something a system can watch for during generation, rather than discovering afterwards in the output.
It also sets a limit on what evaluation can establish. The evaluation module measured models entirely from outside — probabilities, overlap with references, human preference between outputs. All three judge behaviour. None can distinguish a model that has learned the task from one that has learned what scores well, and that distinction is precisely what matters as models get more capable. Looking inside is the only approach that addresses it, which is why a field that began as curiosity about neurons is now treated as safety infrastructure.