Fine-Tuning, LoRA and Transfer

What full fine-tuning costs, why a low-rank update is usually enough, how QLoRA fits a large model onto one accelerator, and when transfer learning is the wrong instinct.

How to Crack the AI Engineer Interview

Adapting an existing model is most of the job in most AI engineering roles, so this gets asked everywhere — in fundamentals rounds as a mechanism question, in design rounds as a cost question. LoRA in particular has become a near-guaranteed topic.

What gets asked

Why full fine-tuning is expensive. Not the weights — the gradients, the optimizer state (Adam keeps two moments per parameter), and the stored activations. The working total is roughly twelve to twenty times the parameter size, which is why large models cannot be fully fine-tuned on modest hardware. You also get a complete model copy per task, and you risk degrading capabilities you are not training.

The PEFT idea. Freeze the pretrained weights, add a small number of trainable parameters, train only those. Gradients and optimizer state are then needed only for the new parameters, which is where the saving comes from.

LoRA. The claim is about the update, not the weights: adapting to a task moves the weights in a few consistent directions, so ΔW\Delta W can be factored into two thin matrices with inner dimension rr. A full update needs d×kd \times k parameters; the factored one needs d×r+r×kd \times r + r \times k. In practice this trains well under 1% of the model. Know that one factor is initialised to zero so the adapter starts as a no-op.

Merging. The adapter is a linear correction, so it can be folded into the weights after training — zero inference overhead. Or kept separate, so one base model serves many tasks by swapping small adapter files, which is how per-customer deployments work.

Why LoRA often matches full fine-tuning. The low-rank constraint acts as regularisation: the model cannot memorise a small dataset as freely, and the frozen base guarantees the original capability survives. On small datasets it sometimes wins outright.

QLoRA. LoRA still needs the frozen base in memory. QLoRA quantizes it to 4-bit — NF4, placing quantization levels to match the bell-shaped distribution of weights — plus double quantization of the scale constants and paged optimizers, with the adapters kept at higher precision. A 65B model becomes fine-tunable on a single large accelerator.

Transfer learning generally. Feature extraction freezes everything and trains a new head; fine-tuning unfreezes some or all of the model. The less data you have and the closer the task, the more you freeze.

The follow-ups that catch people

Why is LoRA free at inference when adapter modules are not? Wanted: a LoRA update is linear and folds into the weight matrix; adapter modules are non-linear networks that cannot be folded, so they stay as extra layers and add latency.

What does rank control? Wanted: capacity against cost. Too small and the update cannot express what the task needs; raising it costs proportionally. A task far from the model's pretraining may hit the low-rank ceiling, where full fine-tuning is genuinely the right answer.

Prompt tuning uses even fewer parameters. Why not use that? Wanted: soft prompts consume context-window space on every request, forever, and are less expressive because they only influence the model through its input. LoRA changes internal computation and disappears into the weights.

When is QLoRA the wrong choice? Wanted: when the model fits without quantization — plain LoRA is simpler. Also flag serving: either dequantize and lose the memory benefit, or run an inference stack with real 4-bit kernels.

Fine-tune or retrieve? Wanted: behaviour versus knowledge. Fine-tune for format, tone, domain conventions and reliability at a skill; retrieve for facts, especially facts that change.

How it gets worded

Transfer

  • "When would you adapt a pretrained model rather than train one?"
  • "Your data is highly specialised and looks nothing like web text. Does transfer still help?"
  • "Why is training from scratch so rare in industry?"
  • "Can transfer learning make things worse? How would you notice?"
  • "Suppose the commercial models are off the table for legal reasons. What is your plan?"
  • "With unlimited compute, would you still start from pretrained weights?"

LoRA and QLoRA

  • "Why does LoRA train such a small share of the parameters, and what does 'low rank' mean in this context?"
  • "How does LoRA keep the base model's knowledge intact?"
  • "How do you pick the rank? What tells you it is set too low?"
  • "Which matrices in a transformer do you attach adapters to?"
  • "Why can a LoRA update be folded into the weights when an adapter module cannot?"
  • "What does mergeability let you do at deployment time that you could not otherwise?"
  • "How does QLoRA fit a very large model onto a single accelerator?"
  • "Where does LoRA lose to full fine-tuning?"

Reading path

  1. Fine-Tuning and LoRA — full fine-tuning's costs, LoRA's mechanism, merging, QLoRA's four parts, and the comparison with prompt tuning, prefix tuning and adapters.
  2. Transfer Learning — the general principle: what transfers, feature extraction versus fine-tuning, when to train from scratch, and negative transfer.
  3. RAG, Fine-Tuning or Long Context — the decision itself, which is the design-round version of this topic.
  4. Why Low Precision — the background QLoRA assumes.
  5. Unsupervised Pre-training and Autoencoders — where the "start from learned weights" idea came from, useful for the historical framing.

Two traps

Assuming transfer learning is always right. It is usually right and sometimes harmful. If the source and target are unrelated, the pretrained weights encode structure that must be unlearned before anything useful happens — negative transfer, which can be worse than random initialisation. The test is whether the learned features carry information about the target, not whether both domains can be described with similar words.

Reaching for fine-tuning too early. The order is: prompt, then retrieve, then fine-tune. A surprising number of problems that sound like fine-tuning problems are prompt or retrieval problems, and saying so is a stronger answer than describing a training pipeline nobody needed.