Backpropagation is easiest to follow once the network's structure and notation are fixed. We will use them for the rest of the module.
Structure
A feedforward network passes information in one direction, from the input to the output, with no loops. It has:
- an input layer, the vector ,
- hidden layers,
- an output layer.
There are layers of weights. Each layer does the same two-step job.
Step 1, the weighted sum (called the pre-activation):
Step 2, the activation function applied to each entry:
We start from . The last layer is different: it applies the output function instead of :
The output function is softmax for classification and linear for regression, as in the previous module. Finally the loss compares the prediction with the true answer.
Shapes
If layer has neurons, then
- , and are vectors of length ,
- is an matrix: one row for each neuron in layer , one column for each neuron in layer .
Row of holds the weights going into neuron of layer . The individual entries are : from neuron in layer to neuron in layer .
Activation functions
The hidden activation is applied to each entry of the vector separately. Two classic choices:
- Sigmoid: , with outputs between 0 and 1.
- Tanh: , with outputs between and 1.
The derivatives we will need later are tidy: for the sigmoid and for tanh . A third choice, ReLU, , is very common in modern networks, and we will use it later in the course.
Why "deep"?
Substitute the layers into each other and you see a nested function:
A deep network is one with many hidden layers, so this expression is a long composition. The long composition is what makes the derivatives interesting.
A worked forward pass
Take a network with 2 inputs, one hidden layer of 2 sigmoid neurons, and 2 softmax outputs. The parameters are
and the input is , with true class 2.
Layer 1. . .
Layer 2. .
Output. Softmax gives .
Loss. The true class is 2, so the cross-entropy loss is .
We will reuse this exact example to see every backward-pass formula produce real numbers.
In code
import numpy as np
def sigmoid(a):
return 1 / (1 + np.exp(-a))
def softmax(a):
e = np.exp(a - a.max()) # subtracting the max avoids overflow
return e / e.sum()
def forward(x, Ws, bs):
"""Returns the pre-activations a, the activations h (h[0] = x) and y_hat."""
hs, as_ = [x], []
h = x
for k in range(len(Ws) - 1): # hidden layers
a = bs[k] + Ws[k] @ h
h = sigmoid(a)
as_.append(a)
hs.append(h)
a = bs[-1] + Ws[-1] @ h # output layer
as_.append(a)
return as_, hs, softmax(a)
Run on the example above this gives , as computed by hand. The function stores every and along the way, because the backward pass will need them.
Build a small network, change its size, and watch the values flow forward.