Inspired by the biological neuron, researchers asked whether a machine could compute like one.
The first model and the birth of a field
In 1943 the neuroscientist Warren McCulloch and the logician Walter Pitts proposed the first highly simplified computational model of a neuron. The McCulloch-Pitts (MP) neuron tried to capture how a neuron combines several inputs into a yes-or-no decision.
That sparked a new field. The term artificial intelligence was formally introduced at a conference in 1956. Shortly after, in 1958, Frank Rosenblatt built on the MP neuron to propose the perceptron.
The hype was enormous. A 1958 New York Times article, reporting on the US Navy's funding, described the perceptron as the embryo of a computer that would be able to "walk, talk, see, write, reproduce itself and be conscious of its existence".
Reality check and winter
In 1969 Marvin Minsky and Seymour Papert published a book analysing exactly what a single perceptron can and cannot do. Its best-known result is that a single perceptron cannot compute simple non-linear functions such as XOR.
They were often misquoted as saying that no neural network could solve this. What they actually showed is that a single layer cannot. Funding nonetheless disappeared, and the field entered what is called the AI winter of connectionism, nearly two decades in which neural network research was close to dead.
Revival
In 1986 David Rumelhart, Geoffrey Hinton and Ronald Williams popularised backpropagation, which finally gave a practical way to train multi-layer networks. But limited computers and small datasets kept progress slow for years. In 2006 Hinton's work on unsupervised pre-training helped spark the modern deep learning revival.
The map of this course
| Module | What it covers |
|---|---|
| 02 Perceptrons and Networks | MP neuron, perceptron, learning rule, why one neuron is not enough |
| 03 Sigmoid Neurons and Gradient Descent | smooth neurons, loss, gradient descent, representation power |
| 04 Feedforward Networks and Backpropagation | how to train multi-layer networks |
| 05 Optimization | faster, steadier versions of gradient descent |
| 06 Bias, Variance and Regularization | why models overfit and how to stop it |
| 07 Initialization, Pre-training and Activations | starting training well, and the choice of activation function |