Regularization Summary and the Loss Landscape

All the regularization tools side by side, how to combine them, and a picture of what they do to the loss surface.

Deep Learning- Fundamentals to Advanced Concepts

The toolbox

TechniqueWhat it changesEffectCost
More datathe datalowers variancecollecting it
L2 (weight decay)the losskeeps weights small, smoother functionsone hyperparameter λ\lambda
L1the losssparse weights, many exactly zeroone hyperparameter
Noise at inputsthe datarobustness, behaves like L2noise level
Dataset augmentationthe databuilds in known invariancesneeds a valid transformation
Parameter sharingthe architecturefar fewer free parametersneeds a structure (e.g. images)
Label smoothingthe targetsless over-confidencesmoothing amount
Ensembles / baggingthe model countlowers variancekk times the compute
Dropoutthe training processno co-adaptation, cheap ensemblelonger training
Early stoppinghow long we trainlimits effective complexityneeds validation data

They attack the same problem, high variance, from different sides, so they can be combined. A common combination for a large network is a little weight decay, dropout, data augmentation and early stopping.

How to apply them

  1. First check which problem you have. High training error means the model is underfitting (bias). More regularization will make that worse, so make the model bigger or train longer.
  2. If there is a large gap between training and validation error, you are overfitting: get more data if you can, then add regularization one tool at a time.
  3. Tune on the validation set. Every strength above, λ\lambda, the dropout rate, the noise level and the patience, is a hyperparameter, and should be chosen on validation data. Keep the test set for the end.
  4. Change one thing at a time, so you know which change helped.

The loss landscape

It helps to picture a regularizer as reshaping the landscape that training explores.

  • L2 adds a bowl centred on zero to the loss surface. Directions in which the data loss is almost flat, where the weights could wander off to huge values without changing the fit, are now tilted back towards the origin. The optimum moves to a point with smaller weights.
  • Dropout and noise make each step use a slightly different loss, so training cannot settle into a narrow dip that fits one exact configuration. It has to find a region that works for many random perturbations.
  • Early stopping does not change the surface at all. It limits how far along the path you walk.

A popular view, supported by some experiments but not settled, is that solutions in wide, flat regions of the loss surface tend to generalise better than solutions in sharp, narrow ones. A flat minimum is one where small changes to the weights barely change the loss, so the fit does not depend on fine details of the training sample. Many of the techniques above can be read as nudging training towards flat regions. Treat this as a helpful intuition, not a theorem.

What comes next

We have so far been fixing the loss and the data to make training generalise. The next module turns to other choices that decide whether training works at all, such as where the weights start and which activation function the neurons use.

Try it yourself
Network lab →

Combine L2, dropout and early stopping and see what each does to the held-out loss.

MediumDiagnosis

You see training error 0.30 and validation error 0.31. Should you add more regularization?

EasyRegularization

Name three regularization techniques that work by changing the training data or targets and not the model.

MediumRegularization

Why is it good to combine several regularization techniques?