Feedforward Networks: Notation and the Forward Pass

The layer-by-layer structure of a feedforward network, the symbols we will use, and a worked forward pass with real numbers.

Deep Learning- Fundamentals to Advanced Concepts

Backpropagation is easiest to follow once the network's structure and notation are fixed. We will use them for the rest of the module.

Structure

A feedforward network passes information in one direction, from the input to the output, with no loops. It has:

  • an input layer, the vector xx,
  • L−1L - 1 hidden layers,
  • an output layer.

There are LL layers of weights. Each layer does the same two-step job.

Step 1, the weighted sum (called the pre-activation):

ak=bk+Wk hk−1a_k = b_k + W_k\,h_{k-1}

Step 2, the activation function gg applied to each entry:

hk=g(ak)h_k = g(a_k)

We start from h0=xh_0 = x. The last layer is different: it applies the output function OO instead of gg:

y^=O(aL)\hat y = O(a_L)

The output function is softmax for classification and linear for regression, as in the previous module. Finally the loss L(y^,y)L(\hat y, y) compares the prediction with the true answer.

A chain from h0 equals x through a1, h1, a2, h2 to a(L), the output y-hat and the loss, with the operation on each arrow
Weighted sum, then activation, layer after layer. The last layer uses the output function O.

Shapes

If layer kk has nkn_k neurons, then

  • hkh_k, aka_k and bkb_k are vectors of length nkn_k,
  • WkW_k is an nk×nk−1n_k \times n_{k-1} matrix: one row for each neuron in layer kk, one column for each neuron in layer k−1k-1.

Row ii of WkW_k holds the weights going into neuron ii of layer kk. The individual entries are Wk,ijW_{k,ij}: from neuron jj in layer k−1k-1 to neuron ii in layer kk.

Activation functions

The hidden activation gg is applied to each entry of the vector separately. Two classic choices:

  • Sigmoid: g(a)=11+e−ag(a) = \dfrac{1}{1 + e^{-a}}, with outputs between 0 and 1.
  • Tanh: g(a)=ea−e−aea+e−ag(a) = \dfrac{e^{a} - e^{-a}}{e^{a} + e^{-a}}, with outputs between −1-1 and 1.

The derivatives we will need later are tidy: for the sigmoid g′(a)=g(a)(1−g(a))g'(a) = g(a)\bigl(1 - g(a)\bigr) and for tanh g′(a)=1−g(a)2g'(a) = 1 - g(a)^2. A third choice, ReLU, g(a)=max⁡(0,a)g(a) = \max(0, a), is very common in modern networks, and we will use it later in the course.

Why "deep"?

Substitute the layers into each other and you see a nested function:

y^=O(bL+WL g(bL−1+WL−1 g(⋯g(b1+W1x)⋯ )))\hat y = O\Bigl(b_L + W_L\, g\bigl(b_{L-1} + W_{L-1}\, g(\cdots g(b_1 + W_1 x)\cdots)\bigr)\Bigr)

A deep network is one with many hidden layers, so this expression is a long composition. The long composition is what makes the derivatives interesting.

A worked forward pass

Take a network with 2 inputs, one hidden layer of 2 sigmoid neurons, and 2 softmax outputs. The parameters are

W1=(0.2−0.30.40.1),b1=(0.1−0.1),W2=(0.5−0.4−0.30.8),b2=(00.1)W_1 = \begin{pmatrix} 0.2 & -0.3 \\ 0.4 & 0.1 \end{pmatrix},\quad b_1 = \begin{pmatrix} 0.1 \\ -0.1 \end{pmatrix},\quad W_2 = \begin{pmatrix} 0.5 & -0.4 \\ -0.3 & 0.8 \end{pmatrix},\quad b_2 = \begin{pmatrix} 0 \\ 0.1 \end{pmatrix}

and the input is x=(1,2)x = (1, 2), with true class 2.

Layer 1. a1=b1+W1x=(0.1+0.2−0.6,  −0.1+0.4+0.2)=(−0.3,  0.5)a_1 = b_1 + W_1 x = (0.1 + 0.2 - 0.6,\; -0.1 + 0.4 + 0.2) = (-0.3,\; 0.5). h1=σ(a1)=(0.4256,  0.6225)h_1 = \sigma(a_1) = (0.4256,\; 0.6225).

Layer 2. a2=b2+W2h1=(0.5(0.4256)−0.4(0.6225),  −0.3(0.4256)+0.8(0.6225)+0.1)=(−0.0362,  0.4703)a_2 = b_2 + W_2 h_1 = (0.5(0.4256) - 0.4(0.6225),\; -0.3(0.4256) + 0.8(0.6225) + 0.1) = (-0.0362,\; 0.4703).

Output. Softmax gives y^=(0.376,  0.624)\hat y = (0.376,\; 0.624).

Loss. The true class is 2, so the cross-entropy loss is −log⁡0.624=0.4716-\log 0.624 = 0.4716.

We will reuse this exact example to see every backward-pass formula produce real numbers.

In code

import numpy as np

def sigmoid(a):
    return 1 / (1 + np.exp(-a))

def softmax(a):
    e = np.exp(a - a.max())        # subtracting the max avoids overflow
    return e / e.sum()

def forward(x, Ws, bs):
    """Returns the pre-activations a, the activations h (h[0] = x) and y_hat."""
    hs, as_ = [x], []
    h = x
    for k in range(len(Ws) - 1):           # hidden layers
        a = bs[k] + Ws[k] @ h
        h = sigmoid(a)
        as_.append(a)
        hs.append(h)
    a = bs[-1] + Ws[-1] @ h                # output layer
    as_.append(a)
    return as_, hs, softmax(a)

Run on the example above this gives y^=(0.376,0.624)\hat y = (0.376, 0.624), as computed by hand. The function stores every aka_k and hkh_k along the way, because the backward pass will need them.

Try it yourself
Neural network sandbox →

Build a small network, change its size, and watch the values flow forward.

EasyNotation

For a network with layers of 784, 100 and 10 neurons, what are the shapes of W1 and W2?

EasyNotation

What is the difference between a_k and h_k?

MediumBackpropagation

Why does the forward pass store every a and h instead of only the final output?