Fine-Tuning and LoRA

Updating every weight to teach a model a new task is slow, memory-hungry and produces a full model copy each time. Low-rank adaptation freezes the original weights and learns a small correction instead — at a fraction of the cost, and often with no loss of quality.

Large Language Models: From Transformers to Frontier Models

Alignment changes how a model behaves in general. Fine-tuning is the other half of post-training, and it is the half most engineers actually end up doing: taking a capable general model and adapting it to your problem — legal question answering, your company's support tickets, a language the base model handles badly. The interesting question here is not whether to fine-tune. It is how much of the model you have to disturb to do it, and the answer turns out to be far less than you would expect.

What full fine-tuning actually costs

Full fine-tuning continues training with every parameter unfrozen. It is the most direct approach and, for small models, the obvious one. At scale it runs into three walls.

Memory. The weights are the small part. Training also needs gradients for every parameter, plus optimizer state — Adam keeps two running averages per weight — plus activations stored for the backward pass. The working total lands at something like twelve to twenty times the size of the parameters themselves. For a 65-billion-parameter model that is hundreds of gigabytes, which is a cluster, not a workstation.

Storage and serving. Fine-tuning produces a complete new model. Ten tasks means ten full copies, each the size of the original, each needing its own slot in memory to serve.

Stability. Every weight is free to move, so the model can drift away from capabilities nobody is currently training — the general ability it was expensive to acquire in the first place. This is catastrophic forgetting, and it is worst exactly where fine-tuning is most tempting: a narrow task with a small dataset.

The parameter-efficient idea

Parameter-efficient fine-tuning (PEFT) starts from a simple observation. Adapting a model to one task is a small change to what it already knows — so why pay to re-learn all of it? Freeze the pretrained weights, add a small number of new trainable parameters, and train only those. Gradients and optimizer state are then needed for the new parameters alone, which is where nearly all the saving comes from.

Low-rank adaptation (LoRA) is the method that made this the default.

LoRA: learn the correction, not the weights

Take any weight matrix in the model, W0W_0, with shape d×kd \times k. Fine-tuning changes it to W0+ΔWW_0 + \Delta W. LoRA's claim is that the update ΔW\Delta W, unlike the weights themselves, does not need to be a full-rank matrix: adapting to a task pushes the weights in a few consistent directions, not in all d×kd \times k of them. So factor it:

ΔW=WAWB,WA∈Rd×r,  WB∈Rr×k,r≪min⁡(d,k)\Delta W = W_A W_B, \qquad W_A \in \mathbb{R}^{d \times r},\; W_B \in \mathbb{R}^{r \times k}, \qquad r \ll \min(d, k)

The product can have rank at most rr, so this builds the update out of rr directions rather than a dense matrix. W0W_0 stays frozen; only WAW_A and WBW_B are trained. In the forward pass the layer computes

h=W0x+αrWA(WBx)h = W_0 x + \frac{\alpha}{r} W_A (W_B x)

where α\alpha is a scaling constant that controls how strongly the adapter speaks relative to the frozen path.

An input flowing through a large frozen weight matrix and in parallel through two small trainable matrices of rank r, the two outputs added together
LoRA runs a second, narrow path beside the frozen weight matrix. Only the two small matrices are trained; their product is the weight update the model would otherwise have had to learn in full.

The arithmetic is the whole argument. A full update to W0W_0 means d×kd \times k parameters. The factored version needs d×r+r×kd \times r + r \times k. With dd and kk in the thousands and rr a small number like 8 or 16, that is a reduction of several orders of magnitude. Across a whole model, LoRA typically trains well under 1% of the parameters — in the original work on a 175-billion-parameter model, roughly ten thousand times fewer trainable parameters, and about a third of the GPU memory of full fine-tuning, because gradients and optimizer state are only kept for the adapters.

One detail matters for training stability: one of the two factors is initialised to zero and the other randomly. Their product is then zero at step one, so the adapted model starts out numerically identical to the pretrained one and the adapters grow a correction from there, rather than jolting the model on the first batch.

Three choices are yours to make: the rank rr (capacity against cost), the scaling α\alpha, and which matrices get adapters — most often the attention projections, sometimes the feed-forward weights as well.

Merging, and why it means free inference

Because the adapter is a linear correction, once training is done you can compute W0+αrWAWBW_0 + \frac{\alpha}{r}W_A W_B and store the result as an ordinary weight matrix. The merged model has exactly the architecture it started with: no extra layers, no extra latency, nothing to special-case at serving time. LoRA costs nothing at inference once merged.

You can also choose not to merge, and that is often the better call. The adapter for a task is a file of megabytes against a base model of gigabytes, so one loaded base model can be steered to many tasks by swapping adapters in and out. For a service that fine-tunes per customer, this is the difference between a feasible deployment and an impossible one.

Why constraining the update often helps

LoRA is cheaper, which you would expect. What surprises people is that it frequently matches full fine-tuning, and sometimes beats it on small datasets.

The reason is that the low-rank restriction behaves as regularisation. Full fine-tuning can move every weight independently, which gives it ample freedom to memorise a small training set and to overwrite pretrained knowledge on the way. LoRA can only express updates inside a low-rank subspace, so it is forced toward changes that are broadly useful rather than sample-specific — and the frozen base weights guarantee the original capability is still intact underneath. Fewer parameters is not only a cost saving here; it is a constraint that generalises.

The flip side is real. A task that genuinely requires reshaping the model's behaviour — a distant domain, a new language, a large high-quality dataset — may hit the ceiling of what a low-rank update can represent. Raising rr buys capacity back at proportional cost, and at some point full fine-tuning is simply the right tool.

QLoRA: fine-tuning a model that doesn't fit

LoRA removes the cost of training parameters, but the frozen base model still has to be held in memory. For the largest models that alone is the blocker. QLoRA removes it by quantizing the frozen base to 4 bits and training adapters on top of it.

Four pieces make it work:

  • 4-bit NormalFloat (NF4). Ordinary quantization spaces its levels evenly, which assumes the values are spread evenly. Model weights are not — they are roughly bell-shaped around zero. NF4 places its sixteen levels so that each receives a similar share of a normal distribution, putting resolution where the weights actually are.
  • Double quantization. Quantization produces scale constants, usually stored in full precision, and a large model has an enormous number of them. QLoRA quantizes those constants too. It buys no accuracy, only memory — but a meaningful amount of it.
  • Paged optimizers. Memory spikes during training no longer crash the run; the driver pages optimizer state between accelerator and host memory the way an operating system pages to disk.
  • High-precision adapters. The base is 4-bit and frozen; the LoRA matrices stay in a 16-bit format and are the only thing trained. Gradients flow through the quantized weights to reach the adapters, but never update them.

The result is that a 65-billion-parameter model can be fine-tuned on a single 48 GB accelerator, with quality close to 16-bit fine-tuning. The background on NF4 and the outlier problems that make 4-bit delicate is in Why Low Precision and Making Low Precision Work.

Serving a QLoRA model is its own decision

Training in 4 bits does not mean serving in 4 bits comes free. Either you dequantize back to a higher precision for inference — simple, and it surrenders the memory advantage — or you serve quantized, which requires an inference stack with real 4-bit kernels. Without them you get on-the-fly dequantization and the latency that comes with it. And if the model is small enough to fine-tune without quantization, plain LoRA is the better choice; the "Q" earns its complexity only at extreme scale or on a single device.

The rest of the PEFT family

LoRA is one answer among several, and the differences come down to where the new parameters sit.

MethodWhat is trainedChanges model internals?Inference cost
Prompt tuningA few soft prompt embeddings prepended to the inputNoConsumes context tokens
Prefix tuningTrainable vectors fed to attention in every layerIndirectly, via attention inputsPrefix must be present every call
AdaptersSmall bottleneck networks inserted into each blockYes, as extra layersAdded depth; not mergeable
LoRALow-rank factors beside existing weight matricesYes, as a weight correctionNone once merged

Prompt tuning is the lightest of all — the model is entirely frozen and only a handful of continuous embeddings are learned to steer it. It is also the least expressive, since it can only influence the model through its input, and it has a cost that is easy to overlook: the soft prompt occupies context-window space on every single request, permanently. Prefix tuning goes further by injecting learned vectors into attention at every layer rather than just at the input, which performs better, but it too must be present at inference. Adapters insert small bottleneck networks inside each block; they are expressive and efficient to train, but because they are non-linear they cannot be folded back into the existing weights, so they add depth and latency for as long as they are used.

LoRA's advantage is that it is the only one of the four that changes the model's internal computation and disappears into the weights afterwards. That combination — expressive during training, invisible at inference — is why it became the default.

Choosing an approach

Work down the list. If a well-written prompt does the job, do not train anything. If the problem is missing knowledge rather than missing behaviour, retrieval is usually the better tool — see RAG, Fine-Tuning or Long Context. If you need the model to behave differently — a format, a tone, a domain's conventions, a skill it performs unreliably — fine-tune, and start with LoRA. Reach for QLoRA when the base model will not otherwise fit, and for full fine-tuning when the change is deep, the dataset is large, and you have the hardware to do it properly.

MediumFine-tuning

Why does full fine-tuning need far more memory than the model's weights alone?

MediumLoRAPEFT

What is the low-rank hypothesis behind LoRA, and how much does it save?

MediumLoRAPEFT

Why does LoRA add no inference cost while adapters do?

HardQLoRAQuantization

When is QLoRA the right choice over plain LoRA, and what does it cost you?