Ask an engineer to compute income tax and they will write a procedure: take the salary, subtract the allowed deductions, apply the slab rates, add the cess. Every step is known in advance, and the same input always produces the same output. Nothing is learned. Now ask the same engineer to write a procedure that decides whether a photograph contains a tree. There is no slab table for trees. Trees come in thousands of shapes, seen from every angle, in every light, partly hidden behind buildings. Nobody can write the rules down, yet every five-year-old applies them without effort.
Machine learning is the answer to the second kind of problem. Instead of writing the rules, we give the computer examples and an algorithm that extracts the rules from the examples.
Where it is used
Wherever a task is natural for people but hard to specify, or wherever a process generates more data than anyone can read, machine learning tends to appear.
- Seeing: recognising objects, faces and handwriting in images (computer vision).
- Hearing and reading: understanding speech and written text, which people express in endlessly different ways.
- Recommending: suggesting the next product on a shopping site, the next film on a streaming service, or people you may know on a social network.
- Forecasting: tomorrow's rainfall from pressure, humidity and past records; next week's demand for a product; whether a team is likely to win its next match.
- Science and engineering: finding patterns in biological, chemical and materials experiments, where a single run can produce millions of measurements.
- Finance: estimating risk, flagging unusual transactions, scoring applications.
Three things machine learning is not
It helps to define the subject by what it rules out.
It is not a fixed procedure. A tax calculator maps inputs to outputs by rules a person wrote. A learning algorithm also maps inputs to outputs, but the mapping is shaped by data. Change the data and the same algorithm produces a different model. Data is the ingredient that separates machine learning from ordinary programming.
It is not memorisation. Suppose you are shown 2,000 photographs, half with trees and half without, and you memorise which is which. You will answer perfectly on those 2,000 photos and learn nothing about trees. Shown a new photograph, you have no basis for an answer. The same is true of a student who memorises the solutions to every assignment: if the exam repeats the assignments they score full marks, but that tells us nothing about whether they understand the subject. What we want from a learning algorithm is generalisation: good performance on data it has not seen. Almost every idea in this course (regularisation, validation, margins, ensembles) exists to protect generalisation.
It is not magic. The results can look magical, but the machinery is mathematics, and a modest amount of it. That is good news: if it is mathematics, it can be understood, checked and improved.
The three ingredients
Take a simple example. You measure the weight and height of 100 people and plot them. The points form a rising cloud: heavier people tend to be taller. You want to predict a new person's height from their weight alone. Three questions immediately arise, and each needs a different branch of mathematics.
- What shape is the relationship? You might assume height rises roughly in a straight line with weight. Assuming a structure (here, a line) is what makes learning possible at all, because it narrows an infinite space of possible answers down to a family you can search. Lines, planes and projections are the language of linear algebra.
- How do we cope with noise? No line passes through all 100 points. Measurements are imprecise, people differ for reasons the data does not record, and some values may be missing. The mathematical language of uncertainty is probability.
- Which line is best? There are infinitely many lines. To pick one, we need a precise measure of how well a line fits, and a method to find the line that scores best. Turning data into a decision this way is optimisation.
Every algorithm in this course is a particular choice of these three ingredients. PCA assumes the data lies near a low-dimensional subspace and minimises the error of projecting onto it. Linear regression assumes a line plus Gaussian noise and maximises the likelihood. A support vector machine assumes a linear boundary and maximises the margin. Once you see the ingredients, the algorithms stop looking like a list of unrelated tricks.
A running theme: comprehension is compression
The computer scientist and philosopher Gregory Chaitin summed up a deep idea in three words: comprehension is compression. If you have really understood a body of data, you can describe it more briefly than by listing it. A physicist who understands falling bodies does not keep a table of every drop ever measured; she keeps one equation and a value for g.
We will use this as a lens throughout the course. When you meet a new algorithm, ask: where is the compression happening? PCA compresses each point to a few coefficients. K-means compresses a data set to K centres. A linear classifier compresses a training set of millions of examples into one weight vector. An SVM keeps only its support vectors. Asking the question often explains why an algorithm behaves the way it does.