This course began with a quote, comprehension is compression, and a promise that the algorithms would turn out to be members of one family rather than a list to memorise. This last chapter looks back over the whole course to make good on that promise, then points to what comes next.
What each method compresses
| Method | Learns | The compressed description |
|---|---|---|
| PCA | directions of greatest variance | coefficients per point plus directions |
| Kernel PCA | non-linear directions, via a kernel | coefficients per point |
| k-means | cluster centres | one centre index per point |
| Gaussian mixture + EM | weighted Gaussians | a density with components |
| Linear / ridge / lasso regression | weights | numbers (lasso: few non-zero) |
| k-nearest neighbours | nothing (stores the data) | no compression at all |
| Decision tree | a sequence of questions | a short list of splits |
| Naive Bayes | per-class feature statistics | probabilities or means and variances |
| Perceptron / logistic regression | a hyperplane | weights |
| SVM | the widest-margin boundary | the support vectors and their 's |
| Random forest / boosting | many trees or stumps | an ensemble of small models |
| Neural network | features and a classifier together | its weights |
Asking "where is the compression?" explains many behaviours: kNN cannot generalise beyond its stored examples, the SVM's predictions are cheap because it keeps only support vectors, and lasso is interpretable because it keeps only a few weights.
The ideas that tie it together
- Projection. PCA projects points onto directions; least squares projects the label vector onto the span of the features. Both are "find the closest point in a subspace".
- Likelihood and priors. Maximum likelihood gave the sample mean, least squares (under Gaussian noise) and logistic regression's cross-entropy. Adding a prior turned maximum likelihood into MAP estimation: ridge (Gaussian prior), lasso (Laplace prior) and Laplace smoothing (Beta prior).
- Kernels. Whenever an algorithm needs only dot products (PCA, regression, the SVM dual, the perceptron), a kernel makes it non-linear without computing the feature space.
- Convexity. Jensen's inequality powered EM; convex losses (hinge, logistic, exponential) make classification tractable where the 0-1 loss is NP-hard; convex problems (least squares, logistic regression, SVMs) have a single global optimum.
- Iterative improvement. Lloyd's algorithm, EM, gradient descent, the perceptron and boosting each improve a solution step by step, and each came with an argument for why the step helps.
- Bias, variance and regularisation. Almost every design choice moves a model along the bias-variance trade-off: polynomial degree, in kNN, tree depth, , , the kernel width, ensembling.
Choosing a method in practice
There is no universally best algorithm. A sensible order of attack on a new tabular problem:
- Start simple: a regularised linear or logistic regression. It is fast, hard to overfit badly, and its weights explain what drives the predictions. It is also the baseline every fancier model must beat.
- If the relationship looks non-linear, try gradient-boosted trees or a random forest. On tabular data they are usually the strongest off-the-shelf choice.
- For a moderate number of points with a clear notion of similarity, a kernel SVM or kNN can work well.
- For images, audio and text, use neural networks, usually starting from a pretrained model.
- Always: hold out test data that plays no part in any choice, use cross-validation to tune hyperparameters, standardise features for distance- and penalty-based methods, and look at the errors the model makes.
For unsupervised problems: PCA for compression and visualisation, k-means for quick round clusters, Gaussian mixtures when clusters differ in size and shape or soft assignments matter, and kernel or spectral methods when the structure is curved.
What this course did not cover
Machine learning is far larger than any one course. Natural next topics include:
- Sequential learning, mapped in the first module: online learning, multi-armed bandits and reinforcement learning, where feedback arrives one decision at a time.
- Deep learning: convolutional networks for images, recurrent networks and transformers for sequences, and the training techniques (normalisation, dropout, adaptive optimisers) that make them work. Our Deep Learning course builds these from the single neuron up.
- Large language models, which apply the same principles (a softmax output, cross-entropy loss, gradient descent, regularisation) at enormous scale. See Large Language Models: From Transformers to Frontier Models.
- Probabilistic graphical models, Gaussian processes and Bayesian inference beyond conjugate priors.
- Learning theory: generalisation bounds, VC dimension and why margins and regularisation provably help.
- Ranking, structured prediction and recommendation, where the output is an ordering or a structure rather than a single label.
Keep the Machine Learning Lab bookmarked: every method in this course is there to experiment with, and the daily word games' machine learning topic is a quick way to keep the vocabulary fresh: try Guess the Term or the Crossword.