Ensembles: Bagging and Model Averaging

Average many models that make different mistakes, and the variance drops. Bagging trains the models on resampled data.

Deep Learning- Fundamentals to Advanced Concepts

If one model has high variance, train several and average them. Each model's mistakes are partly random, and averaging cancels part of the randomness.

Why averaging helps

Suppose each of kk models makes an error with variance σ2\sigma^2, and the errors of any two models have correlation ρ\rho. Then the error of their average has variance

ρ σ2+1−ρk σ2\rho\,\sigma^2 + \frac{1 - \rho}{k}\,\sigma^2

Two extreme cases:

  • Independent errors (ρ=0\rho = 0): the variance drops to σ2/k\sigma^2/k. Averaging 25 models divides it by 25.
  • Identical errors (ρ=1\rho = 1): the variance stays σ2\sigma^2. Averaging identical models gains nothing.

We can confirm the formula by simulation, with σ2=1\sigma^2 = 1:

import numpy as np

rng = np.random.default_rng(5)
for rho in (0.0, 0.5):
    for models in (1, 5, 25):
        common = rng.normal(0, 1, 200000)
        errors = np.sqrt(rho) * common[:, None] + np.sqrt(1 - rho) * rng.normal(0, 1, (200000, models))
        print(f"correlation {rho}: {models:2d} models -> variance of the average {errors.mean(axis=1).var():.3f}   formula {rho + (1 - rho) / models:.3f}")

Output:

correlation 0.0:  1 models -> variance of the average 0.996   formula 1.000
correlation 0.0:  5 models -> variance of the average 0.200   formula 0.200
correlation 0.0: 25 models -> variance of the average 0.040   formula 0.040
correlation 0.5:  1 models -> variance of the average 0.998   formula 1.000
correlation 0.5:  5 models -> variance of the average 0.601   formula 0.600
correlation 0.5: 25 models -> variance of the average 0.520   formula 0.520

With correlation 0.5, going from 5 to 25 models barely helps (0.60 to 0.52): the shared part of the error cannot be averaged away. So an ensemble works best when its members are as different as possible while each one is still good.

Bagging

How do you get different models from one dataset? Bagging (bootstrap aggregating) trains each model on a different bootstrap sample: a random sample of the same size as the training set, drawn with replacement. Some examples appear two or three times, others not at all. (On average about 63% of the distinct examples appear in each sample.) The models see different data, so they make different mistakes.

The recipe:

  1. For b=1,…,kb = 1, \dots, k: draw a bootstrap sample and train model bb on it.
  2. To predict, average the models' outputs (or take a majority vote for classification).

Bagging helps most for unstable models, whose predictions change a lot when the data changes slightly. Deep decision trees are the classic case. We try it on noisy data from y=sin⁡(2πx)y = \sin(2\pi x) with 60 training points:

import numpy as np
from sklearn.tree import DecisionTreeRegressor

def true_f(x):
    return np.sin(2 * np.pi * x)

rng = np.random.default_rng(9)
single, bagged = [], []
for _ in range(100):
    x = rng.uniform(0, 1, 60); y = true_f(x) + rng.normal(0, 0.3, 60)
    x_new = rng.uniform(0, 1, 500); y_new = true_f(x_new) + rng.normal(0, 0.3, 500)
    tree = DecisionTreeRegressor().fit(x[:, None], y)
    single.append(np.mean((tree.predict(x_new[:, None]) - y_new) ** 2))
    votes = []
    for _ in range(25):
        idx = rng.integers(0, 60, 60)                        # bootstrap sample: draw with replacement
        votes.append(DecisionTreeRegressor().fit(x[idx, None], y[idx]).predict(x_new[:, None]))
    bagged.append(np.mean((np.mean(votes, axis=0) - y_new) ** 2))
print(f"one deep tree: {np.mean(single):.4f}   average of 25 bagged trees: {np.mean(bagged):.4f}   (noise alone: 0.0900)")

Output:

one deep tree: 0.1825   average of 25 bagged trees: 0.1377   (noise alone: 0.0900)

A single fully grown tree memorises the noise and gets a test error of 0.1825. Averaging 25 trees trained on bootstrap samples reduces it to 0.1377, closing roughly half of the gap to the noise floor of 0.09, with no change to how each individual tree is built.

Cost

An ensemble of kk models costs kk times as much to train and to run. For large neural networks that is often too expensive, which is one reason the next lesson's trick is so popular: dropout behaves like a huge ensemble that shares its weights.

Bagging is one kind of ensemble. Others average models with different architectures or different random starting weights. In all cases the goal is the same: reduce variance by averaging models that err differently.

MediumEnsembles

Why does averaging k independent models reduce variance by a factor of k but averaging identical models does not?

EasyBagging

What is a bootstrap sample?

HardBagging

Why does bagging help deep decision trees more than it would help a straight-line fit?