The toolbox
| Technique | What it changes | Effect | Cost |
|---|---|---|---|
| More data | the data | lowers variance | collecting it |
| L2 (weight decay) | the loss | keeps weights small, smoother functions | one hyperparameter |
| L1 | the loss | sparse weights, many exactly zero | one hyperparameter |
| Noise at inputs | the data | robustness, behaves like L2 | noise level |
| Dataset augmentation | the data | builds in known invariances | needs a valid transformation |
| Parameter sharing | the architecture | far fewer free parameters | needs a structure (e.g. images) |
| Label smoothing | the targets | less over-confidence | smoothing amount |
| Ensembles / bagging | the model count | lowers variance | times the compute |
| Dropout | the training process | no co-adaptation, cheap ensemble | longer training |
| Early stopping | how long we train | limits effective complexity | needs validation data |
They attack the same problem, high variance, from different sides, so they can be combined. A common combination for a large network is a little weight decay, dropout, data augmentation and early stopping.
How to apply them
- First check which problem you have. High training error means the model is underfitting (bias). More regularization will make that worse, so make the model bigger or train longer.
- If there is a large gap between training and validation error, you are overfitting: get more data if you can, then add regularization one tool at a time.
- Tune on the validation set. Every strength above, , the dropout rate, the noise level and the patience, is a hyperparameter, and should be chosen on validation data. Keep the test set for the end.
- Change one thing at a time, so you know which change helped.
The loss landscape
It helps to picture a regularizer as reshaping the landscape that training explores.
- L2 adds a bowl centred on zero to the loss surface. Directions in which the data loss is almost flat, where the weights could wander off to huge values without changing the fit, are now tilted back towards the origin. The optimum moves to a point with smaller weights.
- Dropout and noise make each step use a slightly different loss, so training cannot settle into a narrow dip that fits one exact configuration. It has to find a region that works for many random perturbations.
- Early stopping does not change the surface at all. It limits how far along the path you walk.
A popular view, supported by some experiments but not settled, is that solutions in wide, flat regions of the loss surface tend to generalise better than solutions in sharp, narrow ones. A flat minimum is one where small changes to the weights barely change the loss, so the fit does not depend on fine details of the training sample. Many of the techniques above can be read as nudging training towards flat regions. Treat this as a helpful intuition, not a theorem.
What comes next
We have so far been fixing the loss and the data to make training generalise. The next module turns to other choices that decide whether training works at all, such as where the weights start and which activation function the neurons use.
Combine L2, dropout and early stopping and see what each does to the held-out loss.