A short reference for what this module and the previous one built.
The vocabulary of derivatives
| Term | What it is | Example |
|---|---|---|
| Derivative | slope of a function of one variable | |
| Partial derivative | slope with respect to one variable, others held constant | |
| Gradient | vector of all partial derivatives | |
| Hessian | matrix of second partial derivatives | needed for second-order Taylor terms |
| Chain rule | derivative of a composite function is a product of derivatives | the heart of backpropagation |
A derivative evaluated at a point is a number. A formula for the derivative becomes a number only once you substitute the current parameter values.
Gradient descent
- Choose a loss that measures how wrong the model is.
- Start with random parameters .
- Repeat: compute and update .
The reason the minus sign works: the Taylor series says a small step in direction changes the loss by about , and the most negative value comes from pointing opposite to the gradient.
The four backpropagation formulas
For softmax with cross-entropy and a feedforward network with activation :
| Step | Formula |
|---|---|
| Output gradient | |
| Weights and biases of layer | |
| Through the weights | |
| Through the activation |
Checklist when training misbehaves
- Loss goes up or explodes. The learning rate is probably too large. Try a smaller .
- Loss barely moves. It may be on a plateau (saturated sigmoids), or may be too small. Early layers with sigmoid activations can be very slow because of vanishing gradients.
- Loss refuses to fall and you suspect a bug. Run a gradient check against numerical gradients on a tiny network.
- Shapes do not match. Every gradient should have the same shape as the parameter it belongs to.
- All neurons behave the same. Check that the weights were initialised randomly.
Where this leaves us
We now have, for a network of any depth: a forward pass, a way to get every gradient, and gradient descent. What remains are ways to make training faster and more reliable, and architectures built for particular kinds of data.