Model Alignment

Supervised fine-tuning teaches a model to imitate good answers, but imitation is not the same as being helpful. Alignment adds a second signal — human preferences — through RLHF or its simpler descendant, DPO.

Large Language Models: From Transformers to Frontier Models

Everything up to this point built a next-token predictor: an architecture, trained on a vast amount of text, that will happily continue whatever you hand it. Useful, and not yet an assistant. This module is about what happens after that — the stage called post-training, where the raw predictor is turned into something you would actually want to talk to. The first and most important idea in it is alignment, which is nothing more complicated than teaching the model to give not merely a plausible answer, but the answer a person would have chosen.

Two kinds of post-training signal

Post-training usually begins with supervised fine-tuning (SFT). You collect prompts paired with well-written responses and continue training on them with the same objective as pretraining — predict the next token. Nothing about the architecture or the loss changes; only the data does, from raw web text to curated demonstrations. The effect is large and immediate: the model stops drifting off into unrelated continuations and starts answering the question in front of it, in a consistent voice.

Then it stops improving in a particular way. An SFT model is a very good imitator, and imitation has a ceiling. The second signal — preference data, in which people compare two candidate answers and say which is better — is what lifts the model past it. Understanding exactly why that second signal is needed is one of the most common conceptual questions asked about modern LLMs, so it is worth being precise.

Why imitation is not enough

Three distinct gaps separate "answers like the demonstrations" from "answers the way people want".

The objective is at the wrong level. SFT minimises cross-entropy token by token. Every token in a demonstration counts the same, and the model is rewarded for assigning high likelihood to text that resembles the training answers. But the thing a reader judges is the whole response — is it correct, is it the right length, does it address what was asked. A model can reach excellent perplexity and still produce a confident, fluent, wrong answer. Low token-level loss and high sequence-level quality are related but not the same quantity, and SFT only optimises the first.

Demonstrations encode style, not judgement. A good demonstration shows what a strong answer looks like. It rarely shows why it beat the alternatives — why the shorter version was better, why the hedge was necessary, why a request should be declined rather than answered carefully. Those are comparative judgements, and a dataset of single correct answers has no way to express them.

Coverage runs out. No demonstration set, at any size, anticipates every ambiguous, adversarial or borderline request a real user will send. Faced with something unlike anything it was shown, an SFT model falls back on what it does by default: produce the most plausible-looking continuation. Plausible is not the same as helpful, and on sensitive requests the difference matters.

Would a perfect SFT set remove the need for alignment?

A fair challenge: suppose you had ten million flawless demonstrations — helpful, accurate, safe. Would preference training still add anything? Largely, yes, because the first gap above does not close with more data. The model would still be trained to maximise the likelihood of those tokens rather than to maximise the quality of its response, and the two come apart exactly where it matters: on the answers where sounding right and being right diverge.

The practical reason preference data is so useful is almost mundane: comparing is easier than composing. Asking an annotator to write the ideal answer to a hard question is slow and inconsistent. Asking them which of two answers they prefer is fast, and people agree with each other far more often. A signal that is cheap to collect at scale is a signal you can actually train on.

RLHF: learning from human rankings

Reinforcement learning from human feedback (RLHF) was the method that made this work at scale, and it remains the reference point every newer technique is compared against. It runs in three stages.

Three stages: demonstrations training an SFT model, ranked comparisons training a reward model, and reinforcement learning optimising the policy against the reward model with a KL penalty
The RLHF pipeline. Demonstrations produce the SFT model; human rankings produce a reward model; reinforcement learning then optimises the model against that reward, held near the SFT model by a KL penalty.

Stage one — demonstrations. Collect prompts, have people write strong responses, and fine-tune the pretrained model on those pairs. This is the SFT step described above, and it produces the starting point for everything that follows.

Stage two — a reward model. Sample several responses from the SFT model for the same prompt and ask annotators to rank them from best to worst. Those rankings train a second network — the reward model — usually built from the same base model but with the language-modelling head replaced by a single scalar output. It is trained with a pairwise ranking loss: for every comparison, the response the human preferred should receive the higher score. What you end up with is an automatic, approximate stand-in for human judgement: hand it a prompt and a response, and it returns a number estimating how much a person would like it.

Stage three — reinforcement learning. Now optimise the language model, treated as a policy, to produce responses the reward model scores highly. The model generates, the reward model scores, and the parameters shift toward the higher-scoring behaviour; proximal policy optimization (PPO) is the usual algorithm because it updates conservatively enough to stay stable. One extra term matters: a KL-divergence penalty that measures how far the policy has drifted from the SFT model it started as. Without it, the policy chases the reward model into strange corners of language and the writing degrades. With it, the model is allowed to improve, but not to stop sounding like a language model.

The payoff is real and was demonstrated early: in the InstructGPT work, a comparatively small model that had been through this pipeline was preferred by human raters over a GPT-3 model many times its size that had only been fine-tuned supervised. Preference training buys quality that scale alone does not.

Where the idea came from

RLHF did not start with chatbots. It came out of reinforcement-learning research on tasks nobody could specify by hand — robot behaviours that were easy to recognise as good but very hard to write a reward function for. The workaround was to let people compare pairs of behaviours and learn the reward from those comparisons. Language turned out to have exactly the same shape of problem, which is how a robotics technique ended up at the centre of how assistants are trained.

What RLHF costs

RLHF works, but nothing about it is cheap or simple.

  • It is three trainings, not one. A reward model to train and maintain, and an RL loop that generates samples, scores them and updates — sensitive to hyperparameters, and capable of collapsing into repetitive or bizarre output when they are wrong.
  • The reward model can be gamed. It is an imperfect model learned from finite data, and the policy is explicitly optimising against it. Any quirk becomes a target. The classic case is length bias: if raters leaned slightly toward longer answers, the reward model learns to like length, and the policy learns to pad. This is reward hacking — a higher score that is not a better answer.
  • The reward model goes stale. It was trained on responses from the old policy. As the policy improves, it generates text the reward model never saw, and its scores become less trustworthy exactly where they are being pushed hardest. Keeping RLHF honest means refreshing comparisons and retraining, round after round.
  • There is an alignment tax. Optimising hard for rated helpfulness can cost performance on capabilities nobody is rating — standard benchmarks sometimes slip after the RL stage. Mixing pretraining data back into the RL objective reduces the effect, at the price of yet another moving part.
  • People are the bottleneck. Every round needs fresh human comparisons. For a small team, that cost alone can put the full pipeline out of reach.

DPO: the same preferences, without the loop

Direct preference optimization (DPO), introduced in 2023, asks an obvious question. The reward model is only ever a device for turning preference pairs into a training signal, and the RL loop is only a device for following that signal. If both are intermediaries, can the model be optimised on the preference pairs directly?

It can. The result underlying DPO is that for the KL-constrained objective RLHF is solving, the optimal policy can be written in closed form in terms of the reward — which can be turned around, so that the reward is expressed in terms of the policy itself. Substitute that back into the reward model's ranking loss and the reward model disappears. What is left is an ordinary supervised loss over the model's own probabilities.

Concretely, you need the same data as stage two of RLHF: triples of a prompt xx, a preferred response y+y^+ and a rejected one y−y^-. Starting from the SFT model, DPO minimises

LDPO=−log⁡σ ⁣(βlog⁡πθ(y+∣x)πref(y+∣x)−βlog⁡πθ(y−∣x)πref(y−∣x))\mathcal{L}_{\text{DPO}} = -\log \sigma\!\left( \beta \log \frac{\pi_\theta(y^+ \mid x)}{\pi_{\text{ref}}(y^+ \mid x)} - \beta \log \frac{\pi_\theta(y^- \mid x)}{\pi_{\text{ref}}(y^- \mid x)} \right)

where πθ\pi_\theta is the model being trained, πref\pi_{\text{ref}} is the frozen SFT model it started from, σ\sigma is the sigmoid, and β\beta controls how far the model may move away from the reference.

Read past the notation and it says something simple. Each response gets a score: how much more likely this model makes it than the reference model did. The loss wants the preferred response's score to exceed the rejected one's. If it already does by a comfortable margin, the loss is near zero and little changes; if the model has it backwards, the gradient pushes up the probability of y+y^+ and down the probability of y−y^-. It is a binary classification loss over pairs — can you tell the better answer from the worse one — applied to the model's own likelihoods.

The reference model in the denominators is doing the job the KL penalty did in RLHF. The model is never rewarded for raising a response's probability in absolute terms, only relative to where it began, so it cannot wander far from the SFT model's language. The constraint that needed a separate penalty term in the RL loop is built into the loss here.

RLHF path through a reward model and a sampling-and-scoring reinforcement learning loop, beside the DPO path going straight from preference pairs to a single supervised loss
Both methods consume the same preference pairs. RLHF routes them through a reward model and an RL loop; DPO folds both steps into one supervised loss on the pairs.

What this removes is substantial. There is no second network to train, no generation inside the training loop and therefore none of its variance, and no PPO hyperparameters to balance. Training is a forward and backward pass over a fixed dataset — the kind of job the whole deep-learning stack is already good at running stably. The original results showed DPO matching or beating PPO-based RLHF on summarisation and dialogue quality while being far easier to implement, which is why it spread quickly and now ships in the standard fine-tuning libraries.

What DPO gives up

DPO is not strictly better, and knowing where it is weaker is the difference between repeating a headline and understanding the trade.

It is offline. RLHF can be run as a live loop: the model generates, a judge scores those generations, the model improves, and the next round explores new behaviour. DPO trains on a fixed set of pairs collected beforehand, so it can only learn the distinctions that set happens to contain. The usual answer is iterative DPO — generate from the improved model, collect fresh comparisons, train again — which recovers much of the loop at the cost of some of the simplicity.

Its signal is coarser. A reward model returns a continuous score, and anything that can be scored can be optimised, including automatically-checkable rewards like did this code run or is this answer correct. That flexibility is what reasoning-focused training is built on. DPO consumes binary comparisons, which is usually enough for helpfulness and tone but less expressive when the quality you care about is graded rather than pairwise.

It inherits the data problem. DPO simplifies how preference data is used, not how it is gathered. Both methods stand on the same expensive foundation of human comparisons, and neither works better than the judgement encoded in them.

Stating the difference precisely

The cleanest way to hold these three apart is by what each one maximises. SFT maximises the likelihood of good answers. RLHF and DPO both maximise agreement with human preferences, and differ only in route — RLHF through a learned reward model and a reinforcement-learning loop, DPO through one supervised loss derived in closed form from the same objective. DPO removes the reward model, the instability and the drift; RLHF keeps online exploration and richer reward signals. Both rest on the same human comparison data.

Where this leaves the model

Alignment is the step that changes what a model is for. Pretraining made it able to continue text; SFT made it answer questions; preference training makes it answer them the way a person would have chosen among the options. Every major assistant goes through some version of this stage, which is why the raw pretrained checkpoints released alongside them behave so differently from the products built on top.

The next lesson turns to the other half of post-training: fine-tuning a model for a specific task or domain, and the low-rank methods that make it affordable.

MediumAlignmentRLHF

Why isn't supervised fine-tuning enough to align a language model?

MediumRLHF

What are the three stages of RLHF, and what does the KL penalty do?

HardDPOAlignment

How does DPO reach the same goal without a reward model?

MediumDPORLHF

Name one thing RLHF can do that offline DPO cannot.