The last chapter made a specific claim: a network initialised from weights that have already learned something about the data trains better than one started at random. That was argued for unsupervised pre-training on your own unlabelled data.
Transfer learning pushes the same claim further. The useful starting point need not come from your data at all. If a model has been trained on a large, related problem, its learned features are a better place to begin than randomness — even though it never saw your task.
Reusing a trained model
The setup has two halves. A source task with abundant data, on which some model was trained. A target task — yours — typically with far less data. Transfer learning carries the representations learned on the first across to the second.
Why should that work? Because what a network learns is layered in generality. Early layers in a vision model learn edges, corners and textures; these are properties of images, not of the particular categories it was asked to classify. Only the later layers become specific to the original labels. A language model learns grammar, usage and world structure well before it learns anything about the corpus it happened to be trained on. The general part transfers; the specific part is what you replace or retrain.
The everyday analogue is apt. Someone who speaks Spanish learns Italian faster than someone starting with no second language — not because they know Italian, but because the structure of the problem is familiar. A network that has learned to distinguish cats from dogs adapts to lions and wolves much faster than a fresh one, for the same reason.
Two ways to adapt
Feature extraction freezes the pretrained weights entirely and uses the model as a fixed transformation from input to representation. You attach a small new head — often a single layer — and train only that. Nothing in the original model moves. This is fast, needs very little data, and cannot damage what the model already knows.
Fine-tuning continues training the pretrained weights themselves, together with whatever new layers you added, on your data. More of the model is free to move, so it can adapt more deeply to your task — and more can go wrong, since the learned features can be degraded by a small or unrepresentative dataset.
The two are ends of a range rather than a binary. A common middle course freezes the early layers, which hold the most general features, and fine-tunes only the later ones. The rule of thumb: the less data you have and the closer your task is to the original, the more you should freeze.
What it buys
Less data. The model does not need to relearn edges or grammar from your examples, so your examples can be spent entirely on what makes your task distinctive. Tasks with a few thousand labels — hopeless from scratch — become routine.
Less compute. Training a large model from random initialisation costs an amount of hardware time most teams cannot contemplate; the largest ones cost millions. Adapting one already trained is frequently a single-accelerator job finishing in hours.
Better results. Not merely cheaper but usually better, because the pretrained model absorbed regularities from a corpus you could never assemble. On a small dataset, a fine-tuned model will generally beat anything trained from scratch, and often by a wide margin.
This is why the approach is the default in both vision and language. The naming makes it explicit: a pretrained transformer is distributed precisely so that others begin from it.
When starting from scratch is right
Transfer learning being usually right is not the same as always right, and five situations genuinely call for random initialisation.
Your domain is unlike anything pretrained. Unusual sensor readings, specialised scientific instruments, data with no analogue in public corpora — there may be no model whose features mean anything for your inputs.
You have an enormous dataset of your own. The advantage of pretraining shrinks as your data grows. With a very large, high-quality, domain-specific corpus, training from scratch lets the model fit your distribution directly, without inheriting assumptions from someone else's.
You need a different architecture. Transfer learning means adopting the pretrained model's structure. Research into new architectures, or a task with structural requirements nothing standard satisfies, requires building from the ground up.
You need control over what the model absorbed. Models trained on broad web data carry its biases. Where behaviour must be tightly controlled and auditable, a curated dataset and a fresh model may be worth the cost.
Licensing forbids it. Sometimes the constraint is not technical at all. Terms of use may rule out the model you would have chosen.
Transfer can be worse than random initialisation. If the source and target tasks are genuinely unrelated, the pretrained weights encode structure that does not apply, and training must first undo it before it can learn anything useful — slower and often to a worse result than starting clean.
The trap is reasoning by metaphor rather than by representation. "DNA is the language of life, so a language model should transfer to genomic prediction" sounds appealing and is wrong: the features a model learns from English — syntax, morphology, discourse — have essentially nothing to do with molecular structure. The question is never whether the two domains can be described with the same words. It is whether the learned features of the source are informative about the target.
A worked choice
Suppose you need an assistant that answers technical questions about your company's software. You have support transcripts, manuals and FAQs — say ten million words. Two options.
Adapt a pretrained language model. It already handles language fluently and has broad general knowledge. Fine-tuning on your material teaches it your products while keeping everything else. A few thousand good examples and a modest training run produce something that handles phrasings it never saw, follow-up questions and the occasional off-topic remark — because the general competence came with the model.
Train from scratch on your ten million words. That sounds like a lot and is nowhere near enough to learn English. The model would acquire narrow, brittle patterns from your documents, struggle to form fluent answers, and fail on anything phrased unlike its training data. You would have spent real compute to get a worse system.
The first wins clearly, and for the general reason: the data you have is enough to teach a model your specialisation, and nowhere near enough to teach it the foundation your specialisation rests on. That asymmetry is what makes transfer learning the default — and recognising the rarer cases where it inverts is the actual skill.