Before choosing an algorithm, you have to recognise what kind of problem you are facing. The single most useful question is: what feedback does the learner get? The answer sorts almost every machine learning problem into one of three paradigms.
Supervised learning: data with answers
In supervised learning every training example comes with the answer we want the model to produce, its label. A set of emails, each marked "spam" or "not spam" by a person, is supervised data. So is a table of houses with their sale prices, or a set of X-ray images each marked by a radiologist.
The task is to learn the mapping from the input (the features) to the label, so that it can be applied to new inputs whose labels are unknown. Supervised problems split by what kind of label is predicted.
- Classification: the label comes from a finite set. Spam or not spam is binary classification (two labels). Recognising a handwritten digit is classification with ten labels, 0 to 9.
- Regression: the label is a real number. Tomorrow's rainfall in millimetres, a house price, a patient's blood-sugar level.
- Ranking: the output is an order. A streaming service has room to show five films; it must not only find films you will like but put the most likely one first, because a perfect recommendation in position 100 is useless.
- Structured prediction: the output has internal structure, such as a parse tree for a sentence or a sequence of tags.
This course concentrates on classification and regression, which are the foundation for the others.
Unsupervised learning: data without answers
In unsupervised learning there are no labels, only data points, and the task is to discover something useful about them. That sounds vague, and it is: without a label, "useful" has to be defined by us.
Take a folder of photographs of animals. You could group them by species, by the number of animals in the picture, by the colour of their fur, or by whether there is an animal in the foreground at all. Each grouping uses a different notion of similarity, and none is "correct" until you decide what you want. Once a notion of similarity is fixed, the problem becomes well defined.
The two unsupervised problems at the heart of this course are:
- Clustering: dividing the data into groups of similar points. A phone that groups your photos by the person in them, without being told who anyone is, is clustering faces.
- Representation learning: finding a better way to describe each data point, so that a later task becomes easier. An image stored as a million raw pixel values is a poor description for recognising trees; a short list of well-chosen features may be far better. Learning that description without labels is unsupervised representation learning, and Module 2 starts with it.
A third, density estimation (modelling the probability distribution the data came from), appears in Module 5 as the engine behind Gaussian mixture models.
Sequential learning: feedback one step at a time
In the first two paradigms, all the data arrives at once. In sequential learning the learner makes a decision, receives feedback, updates itself and decides again, round after round. How much the feedback reveals defines three sub-types.
- Online learning (full information): each morning you decide whether to buy a stock, guided by the advice of several experts. In the evening the market closes and you learn the outcome, which tells you not only whether you were right but whether every expert was right. Algorithms for this setting learn which experts to trust.
- Multi-armed bandits (partial information): a doctor chooses one of four treatments for a patient and a week later sees the outcome, but only for the treatment that was given. What the other three would have done is never observed. Learning with this partial feedback, while still treating patients well, is harder.
- Reinforcement learning: a robot must cross a cluttered room. At each position (a state) it chooses a move (an action), and it learns only from the consequences, such as bumping into a wall. The goal is a good policy, a mapping from states to actions, learned entirely from experience.
Sequential learning is outside the scope of this course, but knowing it exists stops you forcing a sequential problem into a supervised mould.
Placing familiar problems
| Problem | Paradigm | Type |
|---|---|---|
| Spam filter | Supervised | Binary classification |
| Rainfall forecast | Supervised | Regression |
| Film recommendations in order | Supervised | Ranking |
| People-you-may-know suggestions | Supervised | Link prediction (will an edge appear between two people?) |
| Separating a singer's voice from the instruments | Unsupervised | Source separation |
| Grouping phone photos by person | Unsupervised | Clustering |
| Daily stock decisions guided by experts | Sequential | Online learning |
| Robot crossing a room | Sequential | Reinforcement learning |
Notice that the same broad goal (say, recommendations) can be posed in different ways. A large part of practical machine learning is choosing the formulation, and the algorithms in this course will make more sense once you can recognise which formulation each one solves.
The road ahead
The course follows the order in which the ideas build on each other. It begins with unsupervised learning (representation learning, then clustering, then probabilistic estimation), because these develop the tools (projection, eigenvectors, likelihood, convexity) that the rest needs. It then moves to supervised learning: regression, basic classifiers, generative and discriminative models, and finally the more advanced methods (support vector machines, ensembles and neural networks).