From One Neuron to a Network

Gradient descent stays the same when we scale up. What changes is how hard it is to compute the gradient, and why we need backpropagation.

Deep Learning- Fundamentals to Advanced Concepts

In the last module we trained one sigmoid neuron with two parameters, ww and bb. Real networks have thousands, millions or billions of parameters spread over many layers. Does the method still work?

The same algorithm, more parameters

Yes. Gradient descent does not care how many parameters there are. Collect all of them, every weight and every bias in every layer, into one long vector θ\theta. The loss L(θ)L(\theta) is then a function of that vector, and the update is exactly the rule from before:

θ←θ−η ∇L(θ)\theta \leftarrow \theta - \eta\,\nabla L(\theta)

The gradient ∇L\nabla L is a vector of the same length as θ\theta, with one partial derivative for each parameter: how much the loss changes if that single parameter is nudged.

To get a feel for the numbers, take a network for small images with 784 inputs, one hidden layer of 100 neurons and 10 outputs:

LayerWeightsBiases
Input to hidden784×100=78,400784 \times 100 = 78{,}400100
Hidden to output100×10=1,000100 \times 10 = 1{,}00010
Total79,510

Nearly 80,000 parameters, and therefore 80,000 partial derivatives at every step. That is a small network by today's standards.

Where the difficulty is

For the single neuron we derived the two partial derivatives by hand. The loss depended on ww through one sigmoid, so the work was short.

In a network the loss depends on an early-layer weight through everything that comes after it. A nudge to a first-layer weight changes the first layer's outputs, which change the second layer's inputs, and so on all the way to the output and then the loss. The loss is a nested composite function of that weight.

Two obvious ways to get all these derivatives fail:

  • Derive each by hand. There are tens of thousands, each a long chain of terms, and the formulas change whenever the architecture changes.
  • Estimate each numerically. Nudge one parameter slightly, recompute the loss, and see how it changed. That costs a full forward pass per parameter, about 80,000 passes for each single gradient descent step in the network above.

We want something systematic and cheap.

Backpropagation

Backpropagation computes the derivatives with respect to every parameter in one sweep from the output back to the input, reusing work as it goes. The cost of the whole backward sweep is comparable to the cost of a couple of forward passes, no matter how many parameters there are. That is what makes training large networks possible.

The idea rests on the chain rule from calculus, applied in a smart order. The plan for this module:

  1. Fix the notation for a feedforward network and its forward pass.
  2. Build intuition for the chain rule on a very thin network.
  3. Find the gradient at the output layer (softmax with cross-entropy).
  4. Find the gradients of the weights and biases in a layer.
  5. Find how the gradient passes back through a hidden layer.
  6. Combine everything into the algorithm, and run it.
Try it yourself
Backpropagation flow simulator →

Watch a signal go forward through a small network, then see the error travel backwards.

EasyNetworks

How many parameters does a network with 10 inputs, 5 hidden neurons and 3 outputs have?

MediumBackpropagation

Why is estimating the gradient by nudging each parameter and recomputing the loss too expensive for large networks?