This course has met a whole zoo of classifiers: the perceptron, logistic regression, the SVM, AdaBoost, even least squares pressed into service on ±1 labels. They look like separate inventions with separate motivations. This chapter puts them on one axis and shows that they are variations on a single recipe: a linear (or kernel, or tree-based) score, a convex loss on its margin, and usually a penalty.
The margin
Every one of these classifiers computes a real-valued score and predicts . With labels , define the margin of a point:
- : correct; : wrong.
- large: the classifier is confident (rightly if , badly wrong if ).
Each method's training loss depends on a point only through its margin, so each is a function of one variable we can draw.
The losses side by side
| Loss | Formula in terms of | Used by |
|---|---|---|
| 0-1 | the true goal | |
| Perceptron | perceptron | |
| Hinge | SVM | |
| Logistic | logistic regression | |
| Exponential | AdaBoost | |
| Squared | least squares on ±1 labels |
(The logistic loss is the cross-entropy of Module 9 rewritten for ±1 labels; with base-2 logs it passes through 1 at , like the others.)
Why not minimise the 0-1 loss?
The 0-1 loss is a step: flat at 1 for every mistake, flat at 0 for every correct point. Its gradient is zero almost everywhere, so gradient methods get no signal, and minimising it over linear classifiers is NP-hard on non-separable data (Module 7). So every practical method replaces it with a surrogate loss: something convex, so that optimisation is reliable, and ideally an upper bound on the 0-1 loss, so that driving the surrogate down drives the training error down too. Hinge, logistic and exponential losses all satisfy both.
What the shapes reveal
Correct, comfortable points (). Hinge loss is exactly zero: once a point is outside the margin, the SVM ignores it entirely. That is the source of its sparsity (only support vectors matter). Logistic and exponential losses never quite reach zero, so every point keeps a small influence, which is why logistic regression's solution depends on all the data.
Correct but close points (). Hinge, logistic and exponential losses all still charge something, pushing for a margin. The perceptron loss charges nothing: any correct point is fine, however close. That is why the perceptron settles for the first separating boundary it finds.
Confident mistakes (). Hinge, logistic and perceptron losses grow linearly; the exponential loss grows exponentially. That is why AdaBoost is so sensitive to mislabelled points (one wildly wrong point can dominate the loss) while logistic regression and the SVM are more robust.
Points that are too correct (). The squared loss rises again: a point with costs as much as one with . Least squares on ±1 labels therefore pulls its boundary towards easy, far-away points, which is the failure seen in Module 7. It is the only loss in the table that is not decreasing in the margin, and it is a poor choice for classification.
The common recipe
Almost every method in this course is an instance of
a loss on each training point's margin plus a regulariser that limits complexity.
| Method | Loss | Regulariser | Model |
|---|---|---|---|
| Ridge regression | squared | linear | |
| Logistic regression | logistic | (usually) | linear |
| Soft-margin SVM | hinge | linear or kernel | |
| AdaBoost | exponential | number of rounds | sum of stumps |
| Gradient boosting | any | rounds, depth, learning rate | sum of trees |
Choosing a method is largely choosing a loss (how to treat near and wrong points), a regulariser (what "simple" means) and a model family (what shapes are possible). The final chapters take the model family one step further: instead of fixing the features, learn them.