Why Move Against the Gradient

Use the Taylor series to find the best direction for a small step, and arrive at the gradient descent update rule.

Deep Learning- Fundamentals to Advanced Concepts

Now we put the Taylor series to work. The question: standing at parameters θ=(w,b)\theta = (w, b), which small move lowers the loss the most?

The move

Write the parameters as a vector θ=(w,b)\theta = (w, b) and propose a change η u\eta\,u:

θnew=θ+η u\theta_{\text{new}} = \theta + \eta\,u

Here uu is a direction (a vector of length 1) and η\eta is a small positive number that controls the step size. We keep η\eta small because the Taylor approximation only holds close by.

We do not yet know which uu to use. The Taylor series will tell us.

Gradients in one minute

A function of two variables has a slope in each direction.

For L(w,b)=w2+b2L(w, b) = w^2 + b^2:

∂L∂w=2w,∂L∂b=2b\frac{\partial L}{\partial w} = 2w, \qquad \frac{\partial L}{\partial b} = 2b

These are partial derivatives: the derivative with respect to one variable while treating the other as a constant. (That is why the b2b^2 term disappears when differentiating with respect to ww.)

Collect them into a vector and you get the gradient:

∇L=(∂L∂w,  ∂L∂b)=(2w,  2b)\nabla L = \left( \frac{\partial L}{\partial w},\; \frac{\partial L}{\partial b} \right) = (2w,\; 2b)

The formula becomes a vector of numbers when you evaluate it at a point. At (w,b)=(1,−2)(w, b) = (1, -2) the gradient is (2,−4)(2, -4).

Going one level up

The derivative of the gradient, a matrix of all second partial derivatives, is called the Hessian. It is the quadratic term in the Taylor series. We do not need it in this lesson.

What Taylor says about the move

Apply the linear Taylor approximation to the new point:

L(θ+ηu)≈L(θ)+η  uT∇L(θ)L(\theta + \eta u) \approx L(\theta) + \eta\; u^{T} \nabla L(\theta)

The terms with η2\eta^2, η3\eta^3 and higher are far smaller when η\eta is small, so we drop them.

A move is good only if the loss goes down:

L(θ+ηu)−L(θ)<0⟺uT∇L(θ)<0L(\theta + \eta u) - L(\theta) < 0 \quad\Longleftrightarrow\quad u^{T} \nabla L(\theta) < 0

(We can ignore η\eta here because it is positive.) So any uu with uT∇L<0u^{T}\nabla L < 0 is an improvement. Which one is best?

The best direction

The dot product of two vectors is

uT∇L=∥u∥ ∥∇L∥cos⁡βu^{T} \nabla L = \|u\|\,\|\nabla L\|\cos\beta

where β\beta is the angle between uu and the gradient. With ∥u∥=1\|u\| = 1, the length ∥∇L∥\|\nabla L\| is fixed by where we are standing, so all we can change is cos⁡β\cos\beta, which lies between −1-1 and 11.

So uT∇Lu^{T}\nabla L is bounded:

−∥∇L∥  ≤  uT∇L  ≤  ∥∇L∥-\|\nabla L\| \;\le\; u^{T}\nabla L \;\le\; \|\nabla L\|

It is most negative when cos⁡β=−1\cos\beta = -1, that is, when β=180∘\beta = 180^\circ. So the best uu points in exactly the opposite direction to the gradient.

That is where the famous phrase comes from: move in the direction opposite to the gradient. The gradient points uphill, where the loss rises fastest, so the opposite direction is the steepest way downhill.

The update rule

Folding the vector length into η\eta, the rule is

θt+1=θt−η ∇L(θt)\theta_{t+1} = \theta_t - \eta\,\nabla L(\theta_t)

or, written out for our two parameters:

wt+1=wt−η∂L∂w∣wt, bt,bt+1=bt−η∂L∂b∣wt, btw_{t+1} = w_t - \eta \frac{\partial L}{\partial w}\Big|_{w_t,\,b_t}, \qquad b_{t+1} = b_t - \eta \frac{\partial L}{\partial b}\Big|_{w_t,\,b_t}

The bar means "evaluated at the current values." The derivative formula is turned into a number by plugging in the current wtw_t and btb_t.

The whole algorithm is a short loop:

  1. Start with random ww and bb.
  2. Repeat for a fixed number of iterations (or until the loss stops improving):
    1. Compute ∂L/∂w\partial L / \partial w and ∂L/∂b\partial L / \partial b at the current point.
    2. Update ww and bb using the rule above.

The update is gradient descent. The quantity η\eta is the learning rate.

Try it yourself
Error surfaces in 3D →

See the landscape that gradient descent walks across.

Watching it go wrong: the learning rate

Take the simplest loss, L(w)=w2L(w) = w^2. Its derivative is 2w2w, so the update is

w←w−η⋅2w=(1−2η) ww \leftarrow w - \eta \cdot 2w = (1 - 2\eta)\,w

Start at w=1w = 1 and see what different learning rates do:

η\etaSequence of wwWhat happens
0.11, 0.8, 0.64, 0.51, 0.41, ...steady march towards the minimum at 0
1.01, −1, 1, −1, 1, ...bounces forever between two points
1.11, −1.2, 1.44, −1.73, 2.07, ...diverges: each step overshoots further

The guarantee from the Taylor series holds only for a small step. Too large a step jumps over the minimum and can make things worse.

Two more things to watch for:

  • The starting point matters. A loss with several valleys sends you to whichever valley is downhill from where you start. A local minimum is not necessarily the lowest point.
  • Direction is guaranteed, not distance. Gradient descent promises that a small enough step reduces the loss. It does not promise the best possible final answer.
Try it yourself
Optimizer playground →

Run plain gradient descent on the valley and try a learning rate above 0.2.

EasyGradient descent

Compute the gradient of L(w, b) = w^2 + b^2 at (3, -1), then one gradient descent step with learning rate 0.1.

MediumGradient descent

Why does the best direction to move turn out to be 180 degrees from the gradient?

MediumLearning rate

With L(w) = w squared and learning rate 1.1, why does gradient descent diverge?

MediumLearning rateTaylor series

Why do we keep the learning rate small in the derivation?