Learning means reducing errors, so first we need a precise way to count them.
Counting errors for AND
Fix and take the AND function. The perceptron fires when . Try a few values of and count how many of the four inputs are misclassified.
. The sums for the four inputs are . All are negative, so the output is 0 everywhere. Three are right, but AND says should give 1. 1 error.
. The sums are . Input now fires but should not. 1 error.
. The sums are . Input fires wrongly and input does not fire when it should. 2 errors.
. The sums are . Now and fire wrongly. 2 errors.
For the line is . The point is above the line but lies in the negative half-space, where . Which side is positive comes from the inequality, not from the picture.
The error surface
With fixed, the number of errors is a function of . Plot and on the floor and the error as the height, and you get an error surface (also called a loss surface).
Three things to see in it:
- The error is always a whole number (0, 1, 2), so the surface is made of flat plateaus.
- The dark region is where the error is 0: every weight pair there gives a perfect classifier.
- Inside a plateau, nudging the weights a little changes nothing, and at the edges the error jumps.
Explore the landscape and see where the error reaches zero.
A limit on the error
With the input always gives , so it is always classified as 0, which is correct for AND. That point never contributes an error, so the maximum is at most 3.
It cannot reach 3 either. To get and both wrong they must fire, which needs and . But then has sum and fires correctly. So the most you can get is 2 errors.
Why this matters
Plotting error for every weight pair works for two weights. With ten, you cannot draw it, let alone search it by eye. We need an algorithm that finds a zero-error point without seeing the whole surface.
And because the surface is flat almost everywhere, its slope is zero in most places, so "roll downhill" would get no signal. That is why the perceptron uses a different kind of update in the next lesson, and why smooth neurons, introduced later in the course, make gradient descent possible.