2006: deep networks can be trained after all
In 2006 Geoffrey Hinton and colleagues showed a way to train deep networks layer by layer. Each layer first learns, without labels, to model the layer below it (a model called a deep belief network). That gives the whole stack a sensible starting point, after which ordinary backpropagation fine-tunes it.
The technique was soon replaced by simpler ones, but the message mattered more: deep networks are trainable. Researchers began to use the name deep learning, partly as a fresh start after the bad years.
Three things arrive together
1. Better ingredients. The vanishing gradient problem was tamed by simple changes:
- ReLU activation, which outputs the input if positive and zero otherwise. Its slope is 1 for positive inputs, so signals do not shrink layer after layer.
- Smarter weight initialisation.
- Dropout (Hinton and colleagues), which randomly switches off neurons while training so the network cannot lean too hard on any one of them.
2. Compute. Graphics processors (GPUs) were built for drawing pixels, which means doing the same arithmetic on millions of numbers at once. That is exactly what neural networks need. NVIDIA's CUDA (2007) made them programmable, and work around 2009 showed large speed-ups for training networks.
3. Data. ImageNet, released in 2009 by a team led by Fei-Fei Li, collected millions of labelled images. A yearly contest built on it, ILSVRC, asked programs to classify images into 1000 categories using about 1.2 million training images.
Algorithms, data and compute. Each was available separately before 2012. The revival happened when all three were good enough at the same time.
2012: AlexNet
In the 2012 contest, a CNN by Alex Krizhevsky, Ilya Sutskever and Hinton, nicknamed AlexNet, reached a top-5 error rate of about 15.3%. The best non-neural entry scored about 26.2%. It was trained on two GPUs, used ReLU and dropout, and had eight learned layers.
A gap that large, on a hard public benchmark, ended most arguments. Within two years nearly every competitive entry was a deep network, and the same methods spread to speech recognition, where deep networks replaced older models, and then to language, translation and more.
Build and train a small network yourself, and see how its structure affects how well it learns.