Transfer Learning

Pre-training showed that a network starts better from learned weights than from random ones. Transfer learning is that idea taken to its conclusion: begin from a model trained on someone else's much larger problem, and adapt it to yours.

Deep Learning- Fundamentals to Advanced Concepts

The last chapter made a specific claim: a network initialised from weights that have already learned something about the data trains better than one started at random. That was argued for unsupervised pre-training on your own unlabelled data.

Transfer learning pushes the same claim further. The useful starting point need not come from your data at all. If a model has been trained on a large, related problem, its learned features are a better place to begin than randomness — even though it never saw your task.

Reusing a trained model

The setup has two halves. A source task with abundant data, on which some model was trained. A target task — yours — typically with far less data. Transfer learning carries the representations learned on the first across to the second.

Why should that work? Because what a network learns is layered in generality. Early layers in a vision model learn edges, corners and textures; these are properties of images, not of the particular categories it was asked to classify. Only the later layers become specific to the original labels. A language model learns grammar, usage and world structure well before it learns anything about the corpus it happened to be trained on. The general part transfers; the specific part is what you replace or retrain.

The everyday analogue is apt. Someone who speaks Spanish learns Italian faster than someone starting with no second language — not because they know Italian, but because the structure of the problem is familiar. A network that has learned to distinguish cats from dogs adapts to lions and wolves much faster than a fresh one, for the same reason.

Two ways to adapt

Feature extraction freezes the pretrained weights entirely and uses the model as a fixed transformation from input to representation. You attach a small new head — often a single layer — and train only that. Nothing in the original model moves. This is fast, needs very little data, and cannot damage what the model already knows.

Fine-tuning continues training the pretrained weights themselves, together with whatever new layers you added, on your data. More of the model is free to move, so it can adapt more deeply to your task — and more can go wrong, since the learned features can be degraded by a small or unrepresentative dataset.

The two are ends of a range rather than a binary. A common middle course freezes the early layers, which hold the most general features, and fine-tunes only the later ones. The rule of thumb: the less data you have and the closer your task is to the original, the more you should freeze.

A pretrained network with early layers marked general and later layers marked task-specific, shown twice: once with all layers frozen under a new head, once with later layers unfrozen for fine-tuning
Feature extraction freezes the pretrained model and trains only a new head. Fine-tuning unfreezes some or all of it. How much to unfreeze depends on how much data you have and how close the tasks are.

What it buys

Less data. The model does not need to relearn edges or grammar from your examples, so your examples can be spent entirely on what makes your task distinctive. Tasks with a few thousand labels — hopeless from scratch — become routine.

Less compute. Training a large model from random initialisation costs an amount of hardware time most teams cannot contemplate; the largest ones cost millions. Adapting one already trained is frequently a single-accelerator job finishing in hours.

Better results. Not merely cheaper but usually better, because the pretrained model absorbed regularities from a corpus you could never assemble. On a small dataset, a fine-tuned model will generally beat anything trained from scratch, and often by a wide margin.

This is why the approach is the default in both vision and language. The naming makes it explicit: a pretrained transformer is distributed precisely so that others begin from it.

When starting from scratch is right

Transfer learning being usually right is not the same as always right, and five situations genuinely call for random initialisation.

Your domain is unlike anything pretrained. Unusual sensor readings, specialised scientific instruments, data with no analogue in public corpora — there may be no model whose features mean anything for your inputs.

You have an enormous dataset of your own. The advantage of pretraining shrinks as your data grows. With a very large, high-quality, domain-specific corpus, training from scratch lets the model fit your distribution directly, without inheriting assumptions from someone else's.

You need a different architecture. Transfer learning means adopting the pretrained model's structure. Research into new architectures, or a task with structural requirements nothing standard satisfies, requires building from the ground up.

You need control over what the model absorbed. Models trained on broad web data carry its biases. Where behaviour must be tightly controlled and auditable, a curated dataset and a fresh model may be worth the cost.

Licensing forbids it. Sometimes the constraint is not technical at all. Terms of use may rule out the model you would have chosen.

Negative transfer: when the head start is a handicap

Transfer can be worse than random initialisation. If the source and target tasks are genuinely unrelated, the pretrained weights encode structure that does not apply, and training must first undo it before it can learn anything useful — slower and often to a worse result than starting clean.

The trap is reasoning by metaphor rather than by representation. "DNA is the language of life, so a language model should transfer to genomic prediction" sounds appealing and is wrong: the features a model learns from English — syntax, morphology, discourse — have essentially nothing to do with molecular structure. The question is never whether the two domains can be described with the same words. It is whether the learned features of the source are informative about the target.

A worked choice

Suppose you need an assistant that answers technical questions about your company's software. You have support transcripts, manuals and FAQs — say ten million words. Two options.

Adapt a pretrained language model. It already handles language fluently and has broad general knowledge. Fine-tuning on your material teaches it your products while keeping everything else. A few thousand good examples and a modest training run produce something that handles phrasings it never saw, follow-up questions and the occasional off-topic remark — because the general competence came with the model.

Train from scratch on your ten million words. That sounds like a lot and is nowhere near enough to learn English. The model would acquire narrow, brittle patterns from your documents, struggle to form fluent answers, and fail on anything phrased unlike its training data. You would have spent real compute to get a worse system.

The first wins clearly, and for the general reason: the data you have is enough to teach a model your specialisation, and nowhere near enough to teach it the foundation your specialisation rests on. That asymmetry is what makes transfer learning the default — and recognising the rarer cases where it inverts is the actual skill.

MediumTransfer learning

Why do early layers transfer well while later layers often do not?

EasyTransfer learningFine-tuning

What is the difference between feature extraction and fine-tuning, and how do you choose?

HardTransfer learning

What is negative transfer, and how would you avoid walking into it?

MediumTransfer learning

You have a very large, high-quality dataset in a specialised domain. Why might training from scratch now be competitive?