The claim
The early perceptron era was full of big promises. Imagine any decision that depends on some inputs and gives an answer , through a rule that nobody knows. The claim was: show the machine enough examples of pairs and it will learn a rule that gives answers close to . Translation, medical risk, hiring: all the same problem in different clothes.
The embarrassment
In 1969 Marvin Minsky and Seymour Papert published the book Perceptrons. They analysed exactly what a single perceptron can and cannot do. Their best-known example is XOR ("exclusive or"), which outputs 1 when exactly one of its two inputs is 1.
| A | B | XOR |
|---|---|---|
| 0 | 0 | 0 |
| 0 | 1 | 1 |
| 1 | 0 | 1 |
| 1 | 1 | 0 |
Plot the four cases and the problem is visible. A single perceptron separates the two output classes with a straight line. For XOR, no straight line can do it.
Try to draw one line that separates the classes. Then see what a second layer changes.
If a machine cannot learn a function this small, how could it learn translation? Faith in the whole approach dropped.
What they actually said
The story usually gets simplified, and the simplification hurt the field. Minsky and Papert showed limits of a single layer of perceptrons. They knew that stacking layers removes the XOR problem. What they doubted was whether anyone would find a good way to train a multi-layer network.
That doubt was reasonable at the time. Nobody had a widely known training method for deep networks, and it took years to find one (Lesson 4). But readers heard "perceptrons cannot work", not "one layer is not enough".
The first winter
Other things were going wrong at the same time. Computers were weak, datasets were small, and promised results were not arriving. In 1973 a critical British government report (the Lighthill report) argued that AI research had failed to deliver on its promises. Funding in the US and the UK shrank through the 1970s. This period is called the first AI winter.
Two schools of thought existed. Symbolic AI tried to write rules by hand. Connectionism tried to learn from examples with networks of simple units. During the winter, funding and attention favoured the symbolic school. Neural networks survived only in small groups.
Hype, disappointment, withdrawal. It happened to AI in the 1970s, again in the late 1980s when expert systems disappointed, and arguably in smaller waves since. Knowing the pattern makes claims about any new technology easier to judge.