Noise, Data Augmentation and Parameter Sharing

Regularize by changing the data rather than the loss: add noise to inputs, create new examples, share weights, and soften the labels.

Deep Learning- Fundamentals to Advanced Concepts

L2 regularization changes the loss. Several other techniques get a regularizing effect by changing what the model sees or how it is built.

Noise at the inputs

During training, add a small random perturbation to each input: x~=x+ε\tilde x = x + \varepsilon, with ε\varepsilon random noise of variance σ2\sigma^2 in each coordinate. The model is pushed to give nearly the same answer for slightly different inputs, which makes it smoother and less sensitive to details of any single training point.

For a linear model the effect can be worked out exactly. The expected squared error under the noise is

E[(wT(x+ε)−y)2]=(wTx−y)2+σ2 ∥w∥2\mathbb{E}\bigl[(w^{T}(x + \varepsilon) - y)^2\bigr] = (w^{T}x - y)^2 + \sigma^2\,\lVert w \rVert^2

The cross term vanishes because the noise has zero mean, and the noise contributes the extra σ2∥w∥2\sigma^2\lVert w\rVert^2. That is precisely L2 regularization with λ=σ2\lambda = \sigma^2. Training on noisy inputs is weight decay in disguise, at least for a linear model.

We can check this numerically. Fit three weights to 30 examples in three ways: plain least squares, ridge regression with λ=σ2\lambda = \sigma^2, and plain least squares on 3000 noisy copies of the data.

import numpy as np

rng = np.random.default_rng(3)
n, k, sigma = 30, 3, 0.3
X = rng.normal(0, 1, (n, k))
y = X @ np.array([1.0, -2.0, 0.5]) + rng.normal(0, 0.5, n)

w_ols = np.linalg.lstsq(X, y, rcond=None)[0]
w_ridge = np.linalg.solve(X.T @ X / n + sigma ** 2 * np.eye(k), X.T @ y / n)
copies = 3000
X_noisy = np.vstack([X + rng.normal(0, sigma, X.shape) for _ in range(copies)])
w_noise = np.linalg.lstsq(X_noisy, np.tile(y, copies), rcond=None)[0]
print("plain least squares      ", np.round(w_ols, 3))
print("ridge, lambda = sigma^2  ", np.round(w_ridge, 3))
print("trained on noisy copies  ", np.round(w_noise, 3))

Output:

plain least squares       [ 0.969 -1.929  0.61 ]
ridge, lambda = sigma^2   [ 0.955 -1.78   0.574]
trained on noisy copies   [ 0.952 -1.778  0.575]

Training on noisy copies gives the same shrunken weights as ridge regression, to within about 0.004 (the small difference is sampling noise in the copies). Both are pulled towards zero compared with plain least squares. For neural networks there is no exact formula, but the same intuition holds: noise in the inputs discourages the network from relying on tiny, brittle features.

Dataset augmentation

The best cure for overfitting is more data. When you cannot collect more, you can manufacture some: apply a transformation that changes the input but does not change the label, and add the result to the training set.

  • Images: flips, small shifts, rotations, crops, changes in brightness.
  • Audio: small time shifts, added background noise.

The transformations must match the task. Flipping a photo of a cat is fine. Rotating a handwritten "6" by 180 degrees turns it into a "9". Flipping left to right would be wrong for recognising the letter "b". Augmentation injects knowledge about what should not matter, and it effectively enlarges the dataset, which reduces variance.

Parameter sharing

Another way to reduce variance is to use fewer free parameters. In parameter sharing, the same weights are used in many places. The convolutional layers you met in the simulator are the main example: one small filter is slid over the whole image, so a few numbers do the work of a huge fully connected layer, and a pattern learned in one place is automatically recognised everywhere. Sharing builds in the assumption that the same pattern means the same thing wherever it appears.

Try it yourself
Convolutions simulator →

See how one filter is reused at every position of the image.

Noise at the outputs: label smoothing

For classification, the training label is a one-hot vector such as (1,0,0,0)(1, 0, 0, 0). To make the model's cross-entropy loss really small it has to push the probability of the true class towards exactly 1 and the others towards 0. That drives the scores to extreme values and makes the model over-confident.

Label smoothing replaces the hard target by a slightly softer one. With KK classes and a small α\alpha:

t=(1−α)⋅one-hot+αKt = (1 - \alpha)\cdot\text{one-hot} + \frac{\alpha}{K}

With K=4K = 4 and α=0.1\alpha = 0.1 the target (1,0,0,0)(1, 0, 0, 0) becomes (0.925,0.025,0.025,0.025)(0.925, 0.025, 0.025, 0.025). Think of it as saying: "the label might be wrong a tiny fraction of the time."

Compare a well-calibrated prediction with an over-confident one under both kinds of target:

import numpy as np

def cross_entropy(target, pred):
    return -np.sum(target * np.log(pred))

alpha, K = 0.1, 4
hard = np.array([1.0, 0, 0, 0])
smooth = (1 - alpha) * hard + alpha / K
for name, pred in [("calibrated (0.925, 0.025, ...)", np.array([0.925, 0.025, 0.025, 0.025])), ("over-confident (0.97, 0.01, ...)", np.array([0.97, 0.01, 0.01, 0.01]))]:
    print(f"{name}: hard-label loss {cross_entropy(hard, pred):.4f}   smoothed-label loss {cross_entropy(smooth, pred):.4f}")
print("smoothed target", smooth)

Output:

calibrated (0.925, 0.025, ...): hard-label loss 0.0780   smoothed-label loss 0.3488
over-confident (0.97, 0.01, ...): hard-label loss 0.0305   smoothed-label loss 0.3736
smoothed target [0.925 0.025 0.025 0.025]

With hard labels, the over-confident prediction looks better (0.0305 against 0.0780), so training keeps pushing towards more confidence. With smoothed labels, the calibrated prediction is now better (0.3488 against 0.3736): being too confident is penalised. The result is a network whose scores stay moderate and which tends to be better calibrated.

HardNoiseL2

Why is adding Gaussian noise to the inputs of a linear model equivalent to L2 regularization?

MediumAugmentation

Give one example of a label-preserving augmentation that would be wrong for some task.

MediumLabel smoothing

What problem does label smoothing address?