Errors and Error Surfaces

Count mistakes as a function of the weights, and meet the error surface that will return throughout the course.

Deep Learning- Fundamentals to Advanced Concepts

Learning means reducing errors, so first we need a precise way to count them.

Counting errors for AND

Fix w0=−1w_0 = -1 and take the AND function. The perceptron fires when −1+w1x1+w2x2≥0-1 + w_1 x_1 + w_2 x_2 \ge 0. Try a few values of (w1,w2)(w_1, w_2) and count how many of the four inputs are misclassified.

(w1,w2)=(−1,−1)(w_1, w_2) = (-1, -1). The sums for the four inputs are −1,−2,−2,−3-1, -2, -2, -3. All are negative, so the output is 0 everywhere. Three are right, but AND says (1,1)(1, 1) should give 1. 1 error.

(w1,w2)=(1.5,0)(w_1, w_2) = (1.5, 0). The sums are −1,−1,0.5,0.5-1, -1, 0.5, 0.5. Input (1,0)(1, 0) now fires but should not. 1 error.

(w1,w2)=(10,−10)(w_1, w_2) = (10, -10). The sums are −1,−11,9,−1-1, -11, 9, -1. Input (1,0)(1, 0) fires wrongly and input (1,1)(1, 1) does not fire when it should. 2 errors.

(w1,w2)=(1,1)(w_1, w_2) = (1, 1). The sums are −1,0,0,1-1, 0, 0, 1. Now (0,1)(0,1) and (1,0)(1,0) fire wrongly. 2 errors.

Half-spaces, not above or below

For (w1,w2)=(−1,−1)(w_1, w_2) = (-1, -1) the line is −1−x1−x2=0-1 - x_1 - x_2 = 0. The point (1,1)(1, 1) is above the line but lies in the negative half-space, where −1−x1−x2<0-1 - x_1 - x_2 < 0. Which side is positive comes from the inequality, not from the picture.

The error surface

With w0w_0 fixed, the number of errors is a function of (w1,w2)(w_1, w_2). Plot w1w_1 and w2w_2 on the floor and the error as the height, and you get an error surface (also called a loss surface).

Colour map of the number of errors for every pair of weights w1 and w2, showing flat coloured regions with sharp edges
The error is a whole number, so the surface is flat plateaus with sudden steps. Dark blue is the region we want.

Three things to see in it:

  1. The error is always a whole number (0, 1, 2), so the surface is made of flat plateaus.
  2. The dark region is where the error is 0: every weight pair there gives a perfect classifier.
  3. Inside a plateau, nudging the weights a little changes nothing, and at the edges the error jumps.
Try it yourself
Error surfaces in 3D →

Explore the landscape and see where the error reaches zero.

A limit on the error

With w0=−1w_0 = -1 the input (0,0)(0, 0) always gives −1<0-1 < 0, so it is always classified as 0, which is correct for AND. That point never contributes an error, so the maximum is at most 3.

It cannot reach 3 either. To get (0,1)(0,1) and (1,0)(1,0) both wrong they must fire, which needs w2≥1w_2 \ge 1 and w1≥1w_1 \ge 1. But then (1,1)(1, 1) has sum −1+w1+w2≥1-1 + w_1 + w_2 \ge 1 and fires correctly. So the most you can get is 2 errors.

Why this matters

Plotting error for every weight pair works for two weights. With ten, you cannot draw it, let alone search it by eye. We need an algorithm that finds a zero-error point without seeing the whole surface.

And because the surface is flat almost everywhere, its slope is zero in most places, so "roll downhill" would get no signal. That is why the perceptron uses a different kind of update in the next lesson, and why smooth neurons, introduced later in the course, make gradient descent possible.

EasyErrors

With w0 = -1, w1 = -1, w2 = -1, how many errors does AND get, and on which input?

HardErrorsGeometry

With w0 = -1, what is the maximum number of errors AND can make, and why?

MediumOptimization

Why is gradient descent a poor fit for the error surface of a perceptron?