How the gradient at one layer's pre-activation turns into the gradient at the layer before it: multiply by the transposed weights, then by the activation's slope.
Deep Learning- Fundamentals to Advanced Concepts
We can now compute the weight and bias gradients of layer k once we know δk. The remaining question is how to get δk−1, the same quantity for the layer before, from δk. There are two links to cross: from ak back to hk−1, and from hk−1 back to ak−1.
Link 1: back through the weights
The activation hk−1,j feeds every neuron of layer k, through the weights Wk,ij for each i:
ak,i=bk,i+j∑Wk,ijhk−1,j
So a change in hk−1,j changes all the ak,i at once, and the loss feels the sum of those effects. The chain rule now has one term for each i:
In matrix form this sum is a matrix-vector product with the transposed weights:
∂hk−1∂L=WkTδk
The intuition: the error travels back along the same connections it came through, each scaled by the weight on that connection. A neuron that contributed strongly to a large error receives a large share.
Link 2: back through the activation
The activation is applied to each entry on its own: hk−1,j=g(ak−1,j). Changing ak−1,j affects only hk−1,j, so there is a single term:
δk−1,j=∂ak−1,j∂L=∂hk−1,j∂Lg′(ak−1,j)
In vector form, with ⊙ for element-by-element multiplication:
δk−1=(WkTδk)⊙g′(ak−1)
The slope g′ depends on the activation function:
Activation
g′ in terms of the stored values
Sigmoid
h(1−h)
Tanh
1−h2
ReLU
1 if a>0, otherwise 0
Because the forward pass stored hk−1, we get g′ for free.
From delta at layer k: read off the weight and bias gradients, then pass the error one layer back.
Continuing the worked example
For the two-layer network we had δ2=(0.376,−0.376) and
A numerical check again agrees with every entry. The whole gradient of this network, all 12 parameters, came from a few matrix operations.
About the dimensions
The derivative of a vector with respect to another vector is a matrix, with one entry for each pair of components. Here that matrix is Wk for the link through the weights, and a diagonal matrix of slopes for the element-wise activation. Multiplying by these matrices, or by their transposes going backwards, is what the two boxed formulas do.
EasyBackpropagation
Given W2 = [[1, 2], [0, -1]] and delta2 = (0.5, 1), compute W2 transposed times delta2.
MediumBackpropagation
Why does the backward pass use the transpose of the weight matrix?
MediumActivation
Why is g-prime applied element by element and not through a matrix multiplication?