Bias and Variance

Split a model's test error into three parts: its systematic mistake, its sensitivity to the training sample, and unavoidable noise.

Deep Learning- Fundamentals to Advanced Concepts

Why is the test error U-shaped? The answer is an identity that splits the error into parts with clear meanings.

Two imaginary experiments

Suppose we could repeat the whole training process many times, each time with a fresh training set drawn from the same source, and see what prediction the model makes at a fixed input xx. Each run gives a slightly different prediction f^(x)\hat f(x), because each training set has different noise.

Two things can be said about this collection of predictions:

  • Bias. Where is their average, compared with the truth? Bias is the gap between the average prediction and the true function f(x)f(x). A large bias means the model is systematically wrong, even on average. This happens when the model is too simple for the pattern.
  • Variance. How much do the predictions spread around their own average? A large variance means the prediction depends strongly on which training set we happened to get. This happens when the model is so flexible that it follows the noise.

The decomposition

For squared error, the expected error at a new point splits exactly:

E[(f^(x)−y)2]  =  (E[f^(x)]−f(x))2⏟bias2  +  E[(f^(x)−E[f^(x)])2]⏟variance  +  σ2⏟noise\mathbb{E}\bigl[(\hat f(x) - y)^2\bigr] \;=\; \underbrace{\bigl(\mathbb{E}[\hat f(x)] - f(x)\bigr)^2}_{\text{bias}^2} \;+\; \underbrace{\mathbb{E}\Bigl[\bigl(\hat f(x) - \mathbb{E}[\hat f(x)]\bigr)^2\Bigr]}_{\text{variance}} \;+\; \underbrace{\sigma^2}_{\text{noise}}

where y=f(x)+noisey = f(x) + \text{noise} and σ2\sigma^2 is the noise variance. The third term is the irreducible error: no model can predict random noise.

Measuring them

We can measure all three in our polynomial experiment, because we know the true function. For each degree, fit 1000 training sets, record the predictions on a grid of test points, and compute the bias and the variance of those predictions. We also measure the actual test error directly, to check that the parts add up.

import numpy as np
from numpy.polynomial import legendre

def true_f(x):
    return np.sin(2 * np.pi * x)

def features(x, degree):
    return legendre.legvander(2 * x - 1, degree)

def fit(x, y, degree, lam=0.0):
    A = features(x, degree)
    return np.linalg.solve(A.T @ A + (lam + 1e-9) * np.eye(degree + 1), A.T @ y)

x_train = np.linspace(0, 1, 15)
x_test = np.linspace(0.005, 0.995, 200)
noise = 0.2
rng = np.random.default_rng(1)

for degree in (1, 3, 5, 9, 12):
    preds, measured = [], []
    for _ in range(1000):
        y = true_f(x_train) + rng.normal(0, noise, 15)
        w = fit(x_train, y, degree)
        p = features(x_test, degree) @ w
        preds.append(p)
        measured.append(np.mean((p - (true_f(x_test) + rng.normal(0, noise, 200))) ** 2))
    preds = np.array(preds)
    bias_sq = np.mean((preds.mean(axis=0) - true_f(x_test)) ** 2)
    variance = np.mean(preds.var(axis=0))
    print(f"degree {degree:2d}: bias^2 {bias_sq:.4f}  variance {variance:.4f}  noise {noise**2:.4f}  sum {bias_sq + variance + noise**2:.4f}  measured test error {np.mean(measured):.4f}")

Output:

degree  1: bias^2 0.2054  variance 0.0048  noise 0.0400  sum 0.2502  measured test error 0.2499
degree  3: bias^2 0.0056  variance 0.0089  noise 0.0400  sum 0.0545  measured test error 0.0543
degree  5: bias^2 0.0000  variance 0.0135  noise 0.0400  sum 0.0535  measured test error 0.0535
degree  9: bias^2 0.0000  variance 0.0283  noise 0.0400  sum 0.0683  measured test error 0.0681
degree 12: bias^2 0.0002  variance 0.2175  noise 0.0400  sum 0.2577  measured test error 0.2569

The last two columns agree, so the decomposition works. Reading the table:

  • Degree 1: nearly all the error is bias squared (0.205). The straight line cannot follow a sine wave however much data it sees, and it hardly changes between training sets (variance 0.005).
  • Degrees 3 to 5: bias has almost vanished and variance is still small. This is the sweet spot.
  • Degree 12: bias is essentially zero, yet the variance has exploded to 0.22. On average the model is right, but any single fit is wildly off.
Four curves against polynomial degree: training error falling, test error U-shaped, bias squared falling and variance rising
As the model gets more flexible, bias falls and variance rises. Their sum plus the noise is the U-shaped test error.

The trade-off

Making a model more flexible reduces bias and increases variance. Making it simpler does the opposite. The best model balances the two, and the total is the U.

You can diagnose which side you are on:

SymptomDiagnosisThings to try
Training error high, test error high and close to ithigh bias (underfitting)a bigger or more flexible model, better features, train longer
Training error low, test error much higherhigh variance (overfitting)more data, a simpler model, regularization (the rest of this module)

More training data reduces variance without increasing bias, which is why it is the most reliable fix for overfitting when it is available. When more data is not available, regularization is the next tool.

Try it yourself
Overfitting lab →

See bias, variance and test error for every degree on one chart.

EasyBias-variance

State the bias-variance decomposition of expected squared error.

MediumBias-variance

A very flexible model has almost zero bias on average but a large test error. How is that possible?

MediumBias-variance

Why does adding more training data reduce overfitting?