The last two lessons ended with three complaints about the MP neuron: binary inputs only, a threshold set by hand, and every input counting equally. The perceptron (Frank Rosenblatt, 1958) addresses all three.
What changes
- Inputs can be real numbers. A salinity reading or a rating out of 10 is fine.
- Each input has a weight. A large weight means that input matters more.
- The weights can be learned from examples, by an algorithm we meet in "The Perceptron Learning Algorithm".
The rule looks almost like the MP neuron's:
The only change is where there used to be . A small input with a big weight can still push the sum over the line.
Rosenblatt's original design had more parts. This course uses the cleaner version analysed by Minsky and Papert in 1969, which is also what people usually mean today.
Hiding the threshold: the bias trick
Move to the left-hand side:
Now invent an extra input that is always 1, and give it the weight . The inequality becomes
The threshold has disappeared into the weights. From now on a perceptron fires when , and is called the bias.
Why call a bias?
The sum starts from before any input is considered, so it acts like a prior: how keen the neuron is to fire when it has seen nothing.
Take a film example with three Boolean inputs: is Matt Damon in it, is it a thriller, is Christopher Nolan the director, and suppose every weight is 1.
- A picky viewer who goes only when all three hold has . The sum is at least 0 only if all three are 1.
- A film buff who will watch almost anything has . Even with every input off, the neuron fires.
A very negative bias is a high bar to clear. A bias near zero is a low one.
What the weights add
Suppose you love Nolan's films. Only one input is on, "director is Nolan", but its weight is large, because your history says it matters. That single weight can carry the sum over the line even though the other inputs are off. That is the notion of importance, written as a number.
Why study Boolean functions at all?
Real problems are not truth tables, but many of them can be turned into one. The film question can be encoded as many yes/no features (actor? genre? director?) with a yes/no answer (did I like it?). Every film you have watched becomes one row of a truth table.
So what we learn about Boolean functions carries over to real classification: which patterns one perceptron can capture and which it cannot. Continuous inputs are handled the same way, and the later lessons on sigmoid neurons and deep networks build on this.
See a perceptron with weights and a bias. Watch what each weight and the bias do to the line.