Recap: Derivatives and Gradient Descent in Networks

The vocabulary, the four formulas and a debugging checklist, in one place.

Deep Learning- Fundamentals to Advanced Concepts

A short reference for what this module and the previous one built.

The vocabulary of derivatives

TermWhat it isExample
Derivativeslope of a function of one variableddww2=2w\dfrac{d}{dw} w^2 = 2w
Partial derivativeslope with respect to one variable, others held constant∂∂w(w2+b2)=2w\dfrac{\partial}{\partial w}(w^2 + b^2) = 2w
Gradientvector of all partial derivatives∇L=(2w,2b)\nabla L = (2w, 2b)
Hessianmatrix of second partial derivativesneeded for second-order Taylor terms
Chain rulederivative of a composite function is a product of derivativesthe heart of backpropagation

A derivative evaluated at a point is a number. A formula for the derivative becomes a number only once you substitute the current parameter values.

Gradient descent

  1. Choose a loss that measures how wrong the model is.
  2. Start with random parameters θ\theta.
  3. Repeat: compute ∇L(θ)\nabla L(\theta) and update θ←θ−η ∇L(θ)\theta \leftarrow \theta - \eta\,\nabla L(\theta).

The reason the minus sign works: the Taylor series says a small step in direction uu changes the loss by about η uT∇L\eta\,u^T\nabla L, and the most negative value comes from uu pointing opposite to the gradient.

The four backpropagation formulas

For softmax with cross-entropy and a feedforward network with activation gg:

StepFormula
Output gradientδL=y^−eℓ\delta_L = \hat y - e_\ell
Weights and biases of layer kk∇WkL=δk hk−1T,∇bkL=δk\nabla_{W_k}L = \delta_k\,h_{k-1}^{T}, \quad \nabla_{b_k}L = \delta_k
Through the weights∂L∂hk−1=WkTδk\dfrac{\partial L}{\partial h_{k-1}} = W_k^{T}\delta_k
Through the activationδk−1=∂L∂hk−1⊙g′(ak−1)\delta_{k-1} = \dfrac{\partial L}{\partial h_{k-1}} \odot g'(a_{k-1})

Checklist when training misbehaves

  • Loss goes up or explodes. The learning rate is probably too large. Try a smaller η\eta.
  • Loss barely moves. It may be on a plateau (saturated sigmoids), or η\eta may be too small. Early layers with sigmoid activations can be very slow because of vanishing gradients.
  • Loss refuses to fall and you suspect a bug. Run a gradient check against numerical gradients on a tiny network.
  • Shapes do not match. Every gradient should have the same shape as the parameter it belongs to.
  • All neurons behave the same. Check that the weights were initialised randomly.

Where this leaves us

We now have, for a network of any depth: a forward pass, a way to get every gradient, and gradient descent. What remains are ways to make training faster and more reliable, and architectures built for particular kinds of data.

MediumRecap

Which part of backpropagation uses the chain rule, and which part uses the outer product?

EasyRecap

Name three checks you would run if a network's loss is not decreasing.