Every chapter so far has built one model. This module's first two chapters build ensembles: many models combined into one. The idea is old and familiar: ask several independent experts and take the majority view, and you are usually better off than trusting any single one. To see when and why it works, start with what makes a single model wrong.
Bias and variance
Imagine retraining the same kind of model on many different training sets drawn from the same source, and look at its prediction at one fixed input . For squared error, the expected error at splits exactly into three parts:
where is the true underlying value and the noise in the labels.
- Bias: how far the average model is from the truth. A straight line fitted to a curved pattern is wrong on average however much data it gets. High bias means underfitting.
- Variance: how much the model changes from one training set to another. A degree-12 polynomial or a fully grown tree changes drastically when a few points change. High variance means overfitting.
- Noise: the irreducible error in the labels themselves. No model can remove it.
Simple models tend to have high bias and low variance; flexible models the reverse. Most of the course's techniques move a model along this trade-off: regularisation (ridge, lasso, the SVM's ) and tree pruning trade a little bias for less variance; adding features or kernels does the opposite.
Averaging reduces variance
Suppose we had models, each trained on an independent training set, each with variance at a point. Their average has variance , while its bias is the same as one model's. Averaging cuts variance without adding bias.
Two catches. We only have one training set, not independent ones. And models built from related data are correlated: if each pair has correlation , the variance of the average is
As grows, the second term vanishes but the first, , remains. The more similar the models, the less averaging can help. Bagging addresses the first catch; random forests the second.
Bagging
Bootstrap aggregating (bagging), proposed by Leo Breiman in 1996, manufactures many training sets from one.
- For : draw a bootstrap sample, points chosen from the training set with replacement. Some points appear several times, others not at all.
- Train a model on each bootstrap sample.
- Combine: average the predictions (regression) or take a majority vote (classification).
Each bootstrap sample contains on average about 63.2% of the distinct original points: a given point is missed by all draws with probability . The left-out 36.8%, the out-of-bag points, give each model a free validation set: average each point's predictions from the models that did not see it, and you have an error estimate without cross-validation.
Bagging helps most with models that have low bias and high variance, and deep decision trees are the perfect example. It does little for stable, high-bias models such as linear regression: averaging many similar lines gives the same line.
Random forests
Bagged trees are still strongly correlated. If one feature is a very strong predictor, nearly every tree puts it at the root, and the trees look alike. Random forests (Breiman, 2001) add a second source of randomness:
At each split, consider only a random subset of features (commonly for classification), and choose the best split among those.
Now the strong feature is unavailable at many splits, so different trees grow differently. Each tree becomes slightly worse on its own, but the trees become much less correlated, so falls and the average improves. Random forests are grown deep and unpruned, need little tuning (the number of trees and ), rarely overfit as trees are added, and give a useful measure of feature importance (the total impurity reduction each feature achieves across the forest). They remain one of the most reliable off-the-shelf methods for tabular data.