Hype, XOR and the First Winter

Why a simple logic function embarrassed the perceptron in 1969, what Minsky and Papert really said, and how funding dried up.

A Brief History of Deep Learning

The claim

The early perceptron era was full of big promises. Imagine any decision that depends on some inputs xx and gives an answer yy, through a rule ff that nobody knows. The claim was: show the machine enough examples of (x,y)(x, y) pairs and it will learn a rule that gives answers close to f(x)f(x). Translation, medical risk, hiring: all the same problem in different clothes.

The embarrassment

In 1969 Marvin Minsky and Seymour Papert published the book Perceptrons. They analysed exactly what a single perceptron can and cannot do. Their best-known example is XOR ("exclusive or"), which outputs 1 when exactly one of its two inputs is 1.

ABXOR
000
011
101
110

Plot the four cases and the problem is visible. A single perceptron separates the two output classes with a straight line. For XOR, no straight line can do it.

Four points on a square with opposite corners sharing a class, so no single line separates them
The two filled dots are on opposite corners. No single line puts them on one side.
Try it yourself
XOR problem simulator →

Try to draw one line that separates the classes. Then see what a second layer changes.

If a machine cannot learn a function this small, how could it learn translation? Faith in the whole approach dropped.

What they actually said

The story usually gets simplified, and the simplification hurt the field. Minsky and Papert showed limits of a single layer of perceptrons. They knew that stacking layers removes the XOR problem. What they doubted was whether anyone would find a good way to train a multi-layer network.

That doubt was reasonable at the time. Nobody had a widely known training method for deep networks, and it took years to find one (Lesson 4). But readers heard "perceptrons cannot work", not "one layer is not enough".

The first winter

Other things were going wrong at the same time. Computers were weak, datasets were small, and promised results were not arriving. In 1973 a critical British government report (the Lighthill report) argued that AI research had failed to deliver on its promises. Funding in the US and the UK shrank through the 1970s. This period is called the first AI winter.

Two schools of thought existed. Symbolic AI tried to write rules by hand. Connectionism tried to learn from examples with networks of simple units. During the winter, funding and attention favoured the symbolic school. Neural networks survived only in small groups.

A pattern to remember

Hype, disappointment, withdrawal. It happened to AI in the 1970s, again in the late 1980s when expert systems disappointed, and arguably in smaller waves since. Knowing the pattern makes claims about any new technology easier to judge.

MediumXORPerceptron

Why can a single perceptron not learn XOR?

MediumHistory

Minsky and Papert are often quoted as proving that neural networks cannot work. What is wrong with that summary?