Gradient Descent and Backpropagation

An idea from 1847 astronomy, rediscovered several times, finally made deep networks trainable in the 1980s.

A Brief History of Deep Learning

The problem: how do you train many layers?

A single perceptron has a simple rule: when it is wrong, adjust its weights. With several layers, a mistake at the output could be caused by any weight in any layer. Which ones should change, and by how much?

The answer has two ingredients, and both are older than neural networks.

Ingredient 1: gradient descent (1847)

Augustin-Louis Cauchy described gradient descent while working on astronomy, to find the orbit of a heavenly body. The idea works for any quantity you want to make small:

  1. Measure how wrong you are. Call it the loss, a number that depends on the weights.
  2. Work out, for each weight, which direction makes the loss go up. This is the gradient.
  3. Move each weight a small step in the opposite direction.
  4. Repeat.

Picture standing on a hilly landscape in fog, feeling the slope under your feet, and always stepping downhill. The loss is the height and the weights are your position.

Try it yourself
Error surfaces in 3D →

See the landscape the weights walk on, and watch gradient descent go downhill.

Ingredient 2: backpropagation

The gradient of the loss with respect to a weight deep inside a network can be computed with the chain rule of calculus: the effect of a weight on the loss is the effect of the weight on its neuron, times the effect of that neuron on the next layer, and so on to the output.

Backpropagation organises this so that the work for all weights is done in one sweep, from the output back to the input, reusing the pieces instead of recomputing them. The cost is about the same as running the network forward once, which is what made it practical.

Who invented it?

Nobody alone. Variations appeared through the 1960s and 70s in different fields: Seppo Linnainmaa described the general method (reverse-mode automatic differentiation) in 1970, and Paul Werbos proposed using it for neural networks in the 1970s. With no internet to spread ideas, people reinvented it without knowing. In 1986 David Rumelhart, Geoffrey Hinton and Ronald Williams showed clearly that it could train multi-layer networks to learn useful internal representations, and that paper made it famous.

Try it yourself
Backpropagation flow →

Follow the signal forward, then watch the error travel backwards and update each weight.

Why it matters now

Almost every modern network, including the large language models behind chat assistants, is still trained with backpropagation plus a descendant of gradient descent. The 1986 idea is alive inside systems with billions of weights.

Notice the link to Lesson 3. Backpropagation is the training method that Minsky and Papert doubted anyone would find. Once it existed, XOR and far harder problems were within reach in principle.

EasyTraining

In one sentence each, what do gradient descent and backpropagation do?

MediumTrainingMath

Why does the gradient need the chain rule in a multi-layer network?