Alignment

Why a pretrained model is not yet an assistant, what supervised fine-tuning can and cannot do, and how RLHF and DPO use human preferences to close the gap.

How to Crack the AI Engineer Interview

"Why isn't supervised fine-tuning enough?" is one of the most reliable questions in a GenAI interview. It is a good one, because answering it well requires understanding what training objectives actually optimise — and the common wrong answers reveal exactly where someone's understanding stops.

What gets asked

Why a pretrained model is not an assistant. It was trained to continue text, so it continues text — not necessarily helpfully, safely or in answer to the question. Everything that makes a model usable comes after pretraining.

Supervised fine-tuning. Continue training on prompt-and-good-response pairs. Same objective as pretraining, different data. It produces instruction-following and a consistent voice, and then stops improving in a particular way.

Why SFT has a ceiling. Three reasons, and the first is the one interviewers want. The objective is at the wrong level: SFT maximises token-level likelihood while quality is a property of the whole response, so a model can achieve low loss and still answer confidently and wrongly. Beyond that, demonstrations show what a good answer looks like but never why it beat the alternatives, and no demonstration set covers every ambiguous or adversarial request.

RLHF. Three stages: supervised fine-tuning for a starting point; human rankings of sampled responses to train a reward model; reinforcement learning to optimise against that reward, with a KL penalty holding the policy near where it started. Know what the penalty is for — without it the policy chases reward-model quirks and the language degrades.

DPO. Same preference data, no reward model and no RL loop. Because the optimal policy for RLHF's objective can be written in closed form, the reward can be expressed in terms of the policy itself, collapsing everything into one supervised loss on preferred-versus-rejected pairs. The reference model in that loss does the job the KL penalty did.

The trade. DPO removes the reward model, the instability and the reward drift. RLHF keeps online exploration and richer, continuous reward signals — including automatically-checkable ones, which is what reasoning training is built on. Both need the same expensive human comparisons.

The follow-ups that catch people

If you had ten million perfect demonstrations, would you still need preference training? Wanted: largely yes, because the objective mismatch does not close with more data. The model is still maximising the likelihood of tokens rather than the quality of responses.

What is reward hacking? Wanted: the policy exploits flaws in an imperfect learned reward. The standard example is length — if raters mildly preferred longer answers, the reward model learns to like length and the policy learns to pad.

What is the alignment tax? Wanted: optimising hard for rated helpfulness can degrade capabilities nobody is rating, so benchmark scores sometimes fall after the RL stage.

Why is comparison data easier to collect than demonstrations? Wanted: writing the ideal answer is slow and inconsistent; choosing between two is fast and raters agree far more often. Same reason Elo-style evaluation works.

How it gets worded

  • "Why is supervised fine-tuning not enough on its own?"
  • "Name the specific things an SFT-only model gets wrong."
  • "Explain why a chat assistant needs a stage beyond supervised fine-tuning at all."
  • "How do human preferences end up in the weights?"
  • "Why train a reward model? What job is it doing that the preference pairs cannot do directly?"
  • "What is the KL penalty for, and what happens to the model without it?"
  • "What does it mean to over-optimise a reward model? Give me a concrete case of reward hacking, and tell me how you would catch it."
  • "How does DPO get preference learning done without a reward model, and what does it give up in exchange?"
  • "RLHF or DPO — when would you reach for each?"
  • "What is still missing after alignment, and where does that leave safety?"
  • "How would you collect preference data you actually trust?"
  • "The model has been live for a year and usage has drifted. How do you keep it aligned?"

Reading path

  1. Model Alignment — the whole topic: why SFT falls short, RLHF's three stages and its costs, DPO and what it gives up. The core reading for this chapter.
  2. What a Language Model Is — if the pretraining objective is not fresh, the mismatch argument needs it.
  3. Hallucinations and Jailbreaks — what alignment is trying to prevent, and how it is outcompeted.
  4. Where to Go Next — places alignment in the wider post-training picture, including reasoning models.

How to structure the answer

The cleanest version of the big question runs: what each method optimises, then the route, then the trade.

SFT maximises the likelihood of good answers. RLHF and DPO both maximise agreement with human preferences — RLHF via a learned reward model and an RL loop, DPO via one supervised loss derived from the same objective. DPO is simpler and more stable; RLHF keeps exploration and richer rewards. Both depend entirely on the quality of the human comparisons.

That is four sentences and it covers everything the question is testing. Then let them pick the follow-up.

One connection worth making unprompted: hallucination and over-refusal are the same dial. An aligned model balances helpfulness against caution, and tuning away from fabrication pushes toward unnecessary refusal. Candidates who treat them as two separate bugs are missing what alignment actually does.