If one model has high variance, train several and average them. Each model's mistakes are partly random, and averaging cancels part of the randomness.
Why averaging helps
Suppose each of models makes an error with variance , and the errors of any two models have correlation . Then the error of their average has variance
Two extreme cases:
- Independent errors (): the variance drops to . Averaging 25 models divides it by 25.
- Identical errors (): the variance stays . Averaging identical models gains nothing.
We can confirm the formula by simulation, with :
import numpy as np
rng = np.random.default_rng(5)
for rho in (0.0, 0.5):
for models in (1, 5, 25):
common = rng.normal(0, 1, 200000)
errors = np.sqrt(rho) * common[:, None] + np.sqrt(1 - rho) * rng.normal(0, 1, (200000, models))
print(f"correlation {rho}: {models:2d} models -> variance of the average {errors.mean(axis=1).var():.3f} formula {rho + (1 - rho) / models:.3f}")
Output:
correlation 0.0: 1 models -> variance of the average 0.996 formula 1.000
correlation 0.0: 5 models -> variance of the average 0.200 formula 0.200
correlation 0.0: 25 models -> variance of the average 0.040 formula 0.040
correlation 0.5: 1 models -> variance of the average 0.998 formula 1.000
correlation 0.5: 5 models -> variance of the average 0.601 formula 0.600
correlation 0.5: 25 models -> variance of the average 0.520 formula 0.520
With correlation 0.5, going from 5 to 25 models barely helps (0.60 to 0.52): the shared part of the error cannot be averaged away. So an ensemble works best when its members are as different as possible while each one is still good.
Bagging
How do you get different models from one dataset? Bagging (bootstrap aggregating) trains each model on a different bootstrap sample: a random sample of the same size as the training set, drawn with replacement. Some examples appear two or three times, others not at all. (On average about 63% of the distinct examples appear in each sample.) The models see different data, so they make different mistakes.
The recipe:
- For : draw a bootstrap sample and train model on it.
- To predict, average the models' outputs (or take a majority vote for classification).
Bagging helps most for unstable models, whose predictions change a lot when the data changes slightly. Deep decision trees are the classic case. We try it on noisy data from with 60 training points:
import numpy as np
from sklearn.tree import DecisionTreeRegressor
def true_f(x):
return np.sin(2 * np.pi * x)
rng = np.random.default_rng(9)
single, bagged = [], []
for _ in range(100):
x = rng.uniform(0, 1, 60); y = true_f(x) + rng.normal(0, 0.3, 60)
x_new = rng.uniform(0, 1, 500); y_new = true_f(x_new) + rng.normal(0, 0.3, 500)
tree = DecisionTreeRegressor().fit(x[:, None], y)
single.append(np.mean((tree.predict(x_new[:, None]) - y_new) ** 2))
votes = []
for _ in range(25):
idx = rng.integers(0, 60, 60) # bootstrap sample: draw with replacement
votes.append(DecisionTreeRegressor().fit(x[idx, None], y[idx]).predict(x_new[:, None]))
bagged.append(np.mean((np.mean(votes, axis=0) - y_new) ** 2))
print(f"one deep tree: {np.mean(single):.4f} average of 25 bagged trees: {np.mean(bagged):.4f} (noise alone: 0.0900)")
Output:
one deep tree: 0.1825 average of 25 bagged trees: 0.1377 (noise alone: 0.0900)
A single fully grown tree memorises the noise and gets a test error of 0.1825. Averaging 25 trees trained on bootstrap samples reduces it to 0.1377, closing roughly half of the gap to the noise floor of 0.09, with no change to how each individual tree is built.
Cost
An ensemble of models costs times as much to train and to run. For large neural networks that is often too expensive, which is one reason the next lesson's trick is so popular: dropout behaves like a huge ensemble that shares its weights.
Bagging is one kind of ensemble. Others average models with different architectures or different random starting weights. In all cases the goal is the same: reduce variance by averaging models that err differently.