Sending the Gradient Back Through a Hidden Layer

How the gradient at one layer's pre-activation turns into the gradient at the layer before it: multiply by the transposed weights, then by the activation's slope.

Deep Learning- Fundamentals to Advanced Concepts

We can now compute the weight and bias gradients of layer kk once we know δk\delta_k. The remaining question is how to get δk−1\delta_{k-1}, the same quantity for the layer before, from δk\delta_k. There are two links to cross: from aka_k back to hk−1h_{k-1}, and from hk−1h_{k-1} back to ak−1a_{k-1}.

The activation hk−1,jh_{k-1,j} feeds every neuron of layer kk, through the weights Wk,ijW_{k,ij} for each ii:

ak,i=bk,i+∑jWk,ij hk−1,ja_{k,i} = b_{k,i} + \sum_j W_{k,ij}\,h_{k-1,j}

So a change in hk−1,jh_{k-1,j} changes all the ak,ia_{k,i} at once, and the loss feels the sum of those effects. The chain rule now has one term for each ii:

∂L∂hk−1,j=∑i∂L∂ak,i ∂ak,i∂hk−1,j=∑iδk,i Wk,ij\frac{\partial L}{\partial h_{k-1,j}} = \sum_i \frac{\partial L}{\partial a_{k,i}}\,\frac{\partial a_{k,i}}{\partial h_{k-1,j}} = \sum_i \delta_{k,i}\,W_{k,ij}

In matrix form this sum is a matrix-vector product with the transposed weights:

  ∂L∂hk−1=WkT δk  \boxed{\;\frac{\partial L}{\partial h_{k-1}} = W_k^{T}\,\delta_k\;}

The intuition: the error travels back along the same connections it came through, each scaled by the weight on that connection. A neuron that contributed strongly to a large error receives a large share.

The activation is applied to each entry on its own: hk−1,j=g(ak−1,j)h_{k-1,j} = g(a_{k-1,j}). Changing ak−1,ja_{k-1,j} affects only hk−1,jh_{k-1,j}, so there is a single term:

δk−1,j=∂L∂ak−1,j=∂L∂hk−1,j  g′(ak−1,j)\delta_{k-1,j} = \frac{\partial L}{\partial a_{k-1,j}} = \frac{\partial L}{\partial h_{k-1,j}}\; g'(a_{k-1,j})

In vector form, with ⊙\odot for element-by-element multiplication:

  δk−1=(WkT δk)⊙g′(ak−1)  \boxed{\;\delta_{k-1} = \bigl(W_k^{T}\,\delta_k\bigr) \odot g'(a_{k-1})\;}

The slope g′g' depends on the activation function:

Activationg′g' in terms of the stored values
Sigmoidh (1−h)h\,(1 - h)
Tanh1−h21 - h^2
ReLU1 if a>0a > 0, otherwise 0

Because the forward pass stored hk−1h_{k-1}, we get g′g' for free.

Diagram showing delta at layer k feeding the weight gradient, the bias gradient, and the gradient to the previous hidden layer
From delta at layer k: read off the weight and bias gradients, then pass the error one layer back.

Continuing the worked example

For the two-layer network we had δ2=(0.376,−0.376)\delta_2 = (0.376, -0.376) and

W2=(0.5−0.4−0.30.8),h1=(0.4256,  0.6225)W_2 = \begin{pmatrix} 0.5 & -0.4 \\ -0.3 & 0.8 \end{pmatrix},\qquad h_1 = (0.4256,\; 0.6225)

Link 1:

W2Tδ2=(0.5−0.3−0.40.8)(0.376−0.376)=(0.3008−0.4512)W_2^{T}\delta_2 = \begin{pmatrix} 0.5 & -0.3 \\ -0.4 & 0.8 \end{pmatrix}\begin{pmatrix} 0.376 \\ -0.376 \end{pmatrix} = \begin{pmatrix} 0.3008 \\ -0.4512 \end{pmatrix}

Link 2: the sigmoid slope is h1(1−h1)=(0.2445,  0.2350)h_1(1 - h_1) = (0.2445,\; 0.2350), so

δ1=(0.3008×0.2445,  −0.4512×0.2350)=(0.0735,  −0.1060)\delta_1 = (0.3008 \times 0.2445,\; -0.4512 \times 0.2350) = (0.0735,\; -0.1060)

Then the first layer's gradients follow from the previous lesson:

∇W1L=δ1 xT=(0.07350.1471−0.1060−0.2121),∇b1L=δ1\nabla_{W_1} L = \delta_1\,x^{T} = \begin{pmatrix} 0.0735 & 0.1471 \\ -0.1060 & -0.2121 \end{pmatrix},\qquad \nabla_{b_1} L = \delta_1

A numerical check again agrees with every entry. The whole gradient of this network, all 12 parameters, came from a few matrix operations.

About the dimensions

The derivative of a vector with respect to another vector is a matrix, with one entry for each pair of components. Here that matrix is WkW_k for the link through the weights, and a diagonal matrix of slopes for the element-wise activation. Multiplying by these matrices, or by their transposes going backwards, is what the two boxed formulas do.

EasyBackpropagation

Given W2 = [[1, 2], [0, -1]] and delta2 = (0.5, 1), compute W2 transposed times delta2.

MediumBackpropagation

Why does the backward pass use the transpose of the weight matrix?

MediumActivation

Why is g-prime applied element by element and not through a matrix multiplication?