Why Deep Networks Stayed Hard

Backpropagation and universal approximation existed by 1989, yet deep networks barely worked for almost two decades. The reasons, and the bright spots.

A Brief History of Deep Learning

By 1989 the theory looked complete. Backpropagation could train multi-layer networks, and the approximation theorem said such networks were powerful enough. So why did the next fifteen years belong to other methods?

Problem 1: signals fade going backwards

Backpropagation multiplies many small numbers together, one per layer. A sigmoid neuron's slope is never more than 0.25. After ten layers the signal reaching the first layer is at most 0.2510≈10−60.25^{10} \approx 10^{-6}, about a millionth of what it started as.

The early layers therefore learn extremely slowly, or not at all. This is the vanishing gradient problem. Sepp Hochreiter analysed it in his 1991 thesis, and Yoshua Bengio and colleagues studied it in the 1990s.

Problem 2: not enough of everything else

  • Data. Labelled datasets with millions of examples did not exist.
  • Compute. Training a network of useful size took days or weeks on the processors of the time.
  • Know-how. Choosing starting weights, activations and learning rates was trial and error, and a bad choice just gave an untrained network with no clue why.

Meanwhile other methods, such as support vector machines and random forests, were easier to train and often beat neural networks on the problems of the day. For much of the 1990s and 2000s they were the standard tools.

Wrong idea, right name

Many people blamed "neural networks" for this. The real culprits were specific and fixable: small data, slow hardware and fragile training. Each was fixed in turn, which is the story of the next lesson.

The survivors

A few teams kept going and produced results that held up.

Convolutional neural networks (CNNs). Instead of connecting every pixel to every neuron, a CNN slides small shared filters across an image, so it needs far fewer weights and respects the fact that an edge is an edge wherever it appears. Kunihiko Fukushima's Neocognitron (1980) was an early form. Yann LeCun trained CNNs with backpropagation in 1989 to read handwritten digits on postal codes, and his LeNet-5 (1998) was used to read digits on bank cheques.

Try it yourself
Convolutions simulator →

Slide a filter over an image and see how it detects edges and patterns.

Long Short-Term Memory (LSTM). Hochreiter and Jürgen Schmidhuber (1997) designed a recurrent unit with gates that control what to remember and what to forget, so signals over long sequences do not vanish. LSTMs later powered speech recognition and machine translation.

MediumTrainingGradients

Why do gradients vanish in deep networks with sigmoid activations?

MediumCNN

Why do convolutional networks need fewer weights than fully connected networks on images?