Backpropagation Intuition: The Chain Rule on a Thin Network
Use a network with one neuron per layer to see how the chain rule gives every gradient, why we work backwards, and why gradients fade.
Deep Learning- Fundamentals to Advanced Concepts
Before the general formulas, look at the smallest possible deep network. It has one neuron in every layer, so every weight is a single number and there are no matrices to confuse us.
Every factor is easy: the derivative of a sigmoid is σ(1−σ), the derivative of a product wh with respect to h is w, and with respect to w it is h.
Real numbers
Take x=1, y=1 and weights w1=0.5, w2=−0.3, w3=0.8.
Forward pass:h1=0.6225, h2=0.4535, y^=0.5897, loss =0.0842.
The factors:
Factor
Value
y^−y
−0.4103
y^(1−y^)
0.2420
w3
0.80
h2(1−h2)
0.2478
w2
−0.30
h1(1−h1)
0.2350
x
1.0
The gradients: multiplying the factors gives
∂w3∂L=−0.0450,∂w2∂L=−0.01225,∂w1∂L=0.00139
To check the derivation, we nudge each weight by a tiny amount, recompute the loss and divide by the nudge. The numerical estimates are −0.04501, −0.01225 and 0.001388, which agree with the chain-rule values.
Observation 1: work backwards and reuse
Look at how the three gradients share pieces. The first part of the product for w1, namely (y^−y)y^(1−y^)w3h2(1−h2)w2h1(1−h1), contains the first part of the product for w2 as a prefix.
So instead of multiplying the whole chain again for every weight, we go backwards and keep a running product:
Start with ∂y^∂L.
Multiply by y^(1−y^) to get ∂a3∂L. Now ∂w3∂L=∂a3∂Lh2.
Multiply by w3 to get ∂h2∂L, then by h2(1−h2) to get ∂a2∂L. Now ∂w2∂L=∂a2∂Lh1.
Multiply by w2, then by h1(1−h1), to get ∂a1∂L. Now ∂w1∂L=∂a1∂Lx.
Each layer's gradient is built from the next layer's, using a couple of multiplications. The total work is proportional to the number of layers. That is backpropagation: the error signal flows backwards and each weight reads off its own gradient on the way.
Look at the three results: 0.045, 0.012, 0.0014. The earlier the weight, the smaller its gradient, shrinking by a factor of about 4 to 9 at each step here.
The reason is in the factors. Every step back multiplies by a weight and by a sigmoid derivative σ(1−σ), which is never more than 0.25. Multiply enough numbers smaller than 1 and the product collapses toward zero. The early layers then learn very slowly.
This is the vanishing gradient problem, which you met in the history course when we asked why deep networks stayed hard to train. Now you can see it happen in numbers, and the later remedies (other activations, better initialisation) will make sense.
MediumBackpropagation
In the thin network, which quantities does dL/dw2 need, and which does it reuse from the gradient of w3?
MediumVanishing gradients
Why do gradients for early layers tend to be small when every activation is a sigmoid?
MediumBackpropagation
Why is computing gradients from the output backwards better than starting from each weight?