Given the gradient at a layer's pre-activation, the weight gradient is the product of that gradient with the layer's input, and the bias gradient is the gradient itself.
Suppose we already know how the loss changes with the pre-activation of layer k:
δk=∂ak∂L
This is a vector with one entry per neuron in the layer. We will get the gradients for that layer's weights and biases directly from δk.
One weight at a time
Write out one entry of the weighted sum:
ak,i=bk,i+j∑Wk,ijhk−1,j
The weight Wk,ij connects neuron j of the previous layer to neuron i of this layer. It appears only in ak,i, and nowhere else in that layer's sums. So the loss depends on it through that one route, and the chain rule has just one term:
∂Wk,ij∂L=∂ak,i∂L∂Wk,ij∂ak,i=δk,ihk−1,j
because ∂ak,i/∂Wk,ij=hk−1,j, the input that the weight multiplies.
The same argument for the bias, which has coefficient 1 in the sum:
∂bk,i∂L=δk,i
Doing this for every pair (i,j) builds a matrix whose entry in row i, column j is δk,ihk−1,j. That is the outer product of the two vectors:
∇WkL=δkhk−1T∇bkL=δk
The outer product of a column vector of length nk with a row vector of length nk−1 is exactly an nk×nk−1 matrix, the same shape as Wk, as it must be.
Numbers from the worked network
For the last layer of the example we have
δ2=(0.376,−0.376),h1=(0.4256,0.6225)
so
∇W2L=(0.376−0.376)(0.42560.6225)=(0.160−0.1600.234−0.234)
and ∇b2L=(0.376,−0.376). A numerical check, nudging each weight and recomputing the loss, matches these values to about ten decimal places.
The gradient for a weight is (the error signal at its output end) times (the activation at its input end).
- A weight connected to a silent input (hk−1,j=0) has a zero gradient: it had no say in the result, so there is nothing to correct.
- A weight feeding a neuron with a large error signal is changed more.
- Every weight in a row shares the same δk,i and every weight in a column shares the same hk−1,j. That is why the whole matrix comes from two vectors.
In the worked example the second row is exactly the negative of the first, because the softmax gradient (0.376,−0.376) has opposite signs for the two classes.
Try it yourself
Backpropagation flow simulator →
Follow how each weight receives a share of the error.
EasyOuter product
Compute the weight gradient for delta = (2, -1) and input h = (1, 0, 3).
EasyBias
Why is the gradient of a bias equal to delta itself?
MediumOuter product
Why does the weight gradient have the same shape as the weight matrix?