From MP Neuron to Perceptron

Weights, real-valued inputs and a learnable bias: what Rosenblatt's perceptron added, and why Boolean examples still matter.

Deep Learning- Fundamentals to Advanced Concepts

The last two lessons ended with three complaints about the MP neuron: binary inputs only, a threshold set by hand, and every input counting equally. The perceptron (Frank Rosenblatt, 1958) addresses all three.

What changes

Perceptron diagram: three inputs with weights feed a sum, then a threshold, then one output
Inputs are multiplied by weights before being summed.
  1. Inputs can be real numbers. A salinity reading or a rating out of 10 is fine.
  2. Each input has a weight. A large weight means that input matters more.
  3. The weights can be learned from examples, by an algorithm we meet in "The Perceptron Learning Algorithm".

The rule looks almost like the MP neuron's:

y={1if ∑i=1nwixi≥θ0otherwisey = \begin{cases} 1 & \text{if } \sum_{i=1}^{n} w_i x_i \ge \theta \\ 0 & \text{otherwise} \end{cases}

The only change is wixiw_i x_i where there used to be xix_i. A small input with a big weight can still push the sum over the line.

Which perceptron?

Rosenblatt's original design had more parts. This course uses the cleaner version analysed by Minsky and Papert in 1969, which is also what people usually mean today.

Hiding the threshold: the bias trick

Move θ\theta to the left-hand side:

∑i=1nwixi−θ≥0\sum_{i=1}^{n} w_i x_i - \theta \ge 0

Now invent an extra input x0x_0 that is always 1, and give it the weight w0=−θw_0 = -\theta. The inequality becomes

∑i=0nwixi≥0\sum_{i=0}^{n} w_i x_i \ge 0

The threshold has disappeared into the weights. From now on a perceptron fires when ∑i=0nwixi≥0\sum_{i=0}^{n} w_i x_i \ge 0, and w0w_0 is called the bias.

Why call w0w_0 a bias?

The sum starts from w0w_0 before any input is considered, so it acts like a prior: how keen the neuron is to fire when it has seen nothing.

Take a film example with three Boolean inputs: is Matt Damon in it, is it a thriller, is Christopher Nolan the director, and suppose every weight is 1.

  • A picky viewer who goes only when all three hold has w0=−3w_0 = -3. The sum −3+x1+x2+x3-3 + x_1 + x_2 + x_3 is at least 0 only if all three are 1.
  • A film buff who will watch almost anything has w0=0w_0 = 0. Even with every input off, the neuron fires.

A very negative bias is a high bar to clear. A bias near zero is a low one.

What the weights add

Suppose you love Nolan's films. Only one input is on, "director is Nolan", but its weight is large, because your history says it matters. That single weight can carry the sum over the line even though the other inputs are off. That is the notion of importance, written as a number.

Why study Boolean functions at all?

Real problems are not truth tables, but many of them can be turned into one. The film question can be encoded as many yes/no features (actor? genre? director?) with a yes/no answer (did I like it?). Every film you have watched becomes one row of a truth table.

So what we learn about Boolean functions carries over to real classification: which patterns one perceptron can capture and which it cannot. Continuous inputs are handled the same way, and the later lessons on sigmoid neurons and deep networks build on this.

Try it yourself
Perceptron learning simulator →

See a perceptron with weights and a bias. Watch what each weight and the bias do to the line.

EasyPerceptron

Write the perceptron's rule with the bias absorbed, and say what x0 and w0 are.

MediumBias

A viewer goes to a film only if at least 2 of 3 equally weighted conditions hold. What are the weights and bias?

MediumBias

Why is the bias called a prior or baseline disposition?