So far least squares has been justified by geometry: it is the closest point in the span of the features. Module 5 offers another justification. If we write down a probabilistic story for how the labels were generated, maximum likelihood should tell us how to fit it. For the most natural story, it gives least squares exactly, and the probabilistic view then lets us ask how good the fitted weights are.
A probabilistic model for the labels
Assume each label is a linear function of its features plus random noise:
The features are taken as given; the randomness is all in the noise. Equivalently, given , the label is Gaussian with mean and variance .
The log-likelihood of the observed labels is
The first term does not depend on . So maximising the likelihood over is the same as minimising the sum of squared errors:
Least squares is maximum likelihood under Gaussian noise. The squared error is not an arbitrary choice: it is what the Gaussian noise assumption implies. Change the assumption and the loss changes with it. Noise with heavier tails, the Laplace distribution, leads to minimising the sum of absolute errors instead, which is much less sensitive to outliers.
How good is ŵ?
The estimate depends on the noise in the labels, so it is random: a fresh data set with new noise would give a different . Treat it like any estimator and ask how close it lands to the true .
Substituting :
- Unbiased: the noise has mean zero, so . On average, least squares hits the truth.
- Its spread: the covariance of is , and the expected squared distance from the truth is
where are the eigenvalues of .
That formula explains a familiar failure. If two features are nearly duplicates (say, height in centimetres and height in inches with a little rounding noise), has an eigenvalue close to zero, its reciprocal is huge, and the estimated weights swing wildly from one data set to the next, often as large positive and negative weights that nearly cancel. The predictions may still be fine on the training data, but the weights are meaningless and the predictions on new data are fragile. This is multicollinearity, and the next chapter's ridge regression is built to tame it.
Estimating error on new data
The quantity we care about is the expected squared error on new data, the generalisation error. Training error underestimates it, more and more as the model gets more flexible, so it cannot choose between a degree-3 and a degree-9 polynomial, or between penalty strengths. We need an estimate from data the model was not fitted to.
A validation set
The simplest approach: split the labelled data. Fit each candidate model on the training part, measure its error on the held-out validation part, and pick the best. Then (optionally) refit the chosen model on all the data. The drawback is waste: the validation points are not used for fitting, and with little data the estimate from a small validation set is itself noisy.
k-fold cross-validation
Cross-validation reuses every point for both purposes.
- Split the data into equal folds (commonly or ).
- For each fold in turn, fit the model on the other folds and measure its error on the held-out fold.
- Average the errors. That average is the cross-validation estimate of generalisation error.
Run this for every candidate (each degree, each penalty strength) and choose the one with the lowest estimate. Every point is used for validation exactly once and for training times.
With , each fold is a single point: leave-one-out cross-validation. It wastes almost nothing and, for least squares and ridge regression, can even be computed without refitting times (there is a closed-form shortcut using the diagonal of the "hat" matrix). Its estimate has low bias but can have high variance, which is one reason 5 or 10 folds are the common default.
One rule matters more than any detail: the test data must not influence any choice. If the same data is used to choose the model and to report its accuracy, the reported accuracy is optimistic. Keep a final test set untouched until every decision has been made.