Gradients of Weights and Biases: The Outer Product

Given the gradient at a layer's pre-activation, the weight gradient is the product of that gradient with the layer's input, and the bias gradient is the gradient itself.

Deep Learning- Fundamentals to Advanced Concepts

Suppose we already know how the loss changes with the pre-activation of layer kk:

δk=∂L∂ak\delta_k = \frac{\partial L}{\partial a_k}

This is a vector with one entry per neuron in the layer. We will get the gradients for that layer's weights and biases directly from δk\delta_k.

One weight at a time

Write out one entry of the weighted sum:

ak,i=bk,i+∑jWk,ij  hk−1,ja_{k,i} = b_{k,i} + \sum_{j} W_{k,ij}\; h_{k-1,j}

The weight Wk,ijW_{k,ij} connects neuron jj of the previous layer to neuron ii of this layer. It appears only in ak,ia_{k,i}, and nowhere else in that layer's sums. So the loss depends on it through that one route, and the chain rule has just one term:

∂L∂Wk,ij=∂L∂ak,i  ∂ak,i∂Wk,ij=δk,i  hk−1,j\frac{\partial L}{\partial W_{k,ij}} = \frac{\partial L}{\partial a_{k,i}}\;\frac{\partial a_{k,i}}{\partial W_{k,ij}} = \delta_{k,i}\; h_{k-1,j}

because ∂ak,i/∂Wk,ij=hk−1,j\partial a_{k,i} / \partial W_{k,ij} = h_{k-1,j}, the input that the weight multiplies.

The same argument for the bias, which has coefficient 1 in the sum:

∂L∂bk,i=δk,i\frac{\partial L}{\partial b_{k,i}} = \delta_{k,i}

The matrix form: an outer product

Doing this for every pair (i,j)(i, j) builds a matrix whose entry in row ii, column jj is δk,i hk−1,j\delta_{k,i}\,h_{k-1,j}. That is the outer product of the two vectors:

  ∇WkL=δk hk−1T    ∇bkL=δk  \boxed{\;\nabla_{W_k} L = \delta_k\, h_{k-1}^{T}\;}\qquad\qquad \boxed{\;\nabla_{b_k} L = \delta_k\;}

The outer product of a column vector of length nkn_k with a row vector of length nk−1n_{k-1} is exactly an nk×nk−1n_k \times n_{k-1} matrix, the same shape as WkW_k, as it must be.

Numbers from the worked network

For the last layer of the example we have

δ2=(0.376,  −0.376),h1=(0.4256,  0.6225)\delta_2 = (0.376,\; -0.376), \qquad h_1 = (0.4256,\; 0.6225)

so

∇W2L=(0.376−0.376)(0.42560.6225)=(0.1600.234−0.160−0.234)\nabla_{W_2} L = \begin{pmatrix} 0.376 \\ -0.376 \end{pmatrix}\begin{pmatrix} 0.4256 & 0.6225 \end{pmatrix} = \begin{pmatrix} 0.160 & 0.234 \\ -0.160 & -0.234 \end{pmatrix}

and ∇b2L=(0.376,  −0.376)\nabla_{b_2} L = (0.376,\; -0.376). A numerical check, nudging each weight and recomputing the loss, matches these values to about ten decimal places.

Reading the formula

The gradient for a weight is (the error signal at its output end) times (the activation at its input end).

  • A weight connected to a silent input (hk−1,j=0h_{k-1,j} = 0) has a zero gradient: it had no say in the result, so there is nothing to correct.
  • A weight feeding a neuron with a large error signal is changed more.
  • Every weight in a row shares the same δk,i\delta_{k,i} and every weight in a column shares the same hk−1,jh_{k-1,j}. That is why the whole matrix comes from two vectors.

In the worked example the second row is exactly the negative of the first, because the softmax gradient (0.376,−0.376)(0.376, -0.376) has opposite signs for the two classes.

Try it yourself
Backpropagation flow simulator →

Follow how each weight receives a share of the error.

EasyOuter product

Compute the weight gradient for delta = (2, -1) and input h = (1, 0, 3).

EasyBias

Why is the gradient of a bias equal to delta itself?

MediumOuter product

Why does the weight gradient have the same shape as the weight matrix?