Nesterov Accelerated Gradient

Look before you leap: measure the gradient where momentum is about to carry you, so you can brake before overshooting.

Deep Learning- Fundamentals to Advanced Concepts

Momentum computes the gradient at the current point, then adds the old velocity. But we already know that the old velocity is going to move us by about γ vt−1\gamma\,v_{t-1} regardless. Nesterov accelerated gradient (NAG) uses that knowledge.

The idea

First take the step that momentum is going to take anyway and arrive at a look-ahead point

θ~t=θt−γ vt−1\tilde\theta_t = \theta_t - \gamma\,v_{t-1}

Then measure the gradient there, and use it to correct:

vt=γ vt−1+η ∇L(θ~t),θt+1=θt−vtv_t = \gamma\,v_{t-1} + \eta\,\nabla L(\tilde\theta_t), \qquad \theta_{t+1} = \theta_t - v_t

The only difference from momentum is where the gradient is evaluated.

Why does that help? Suppose momentum is about to overshoot the minimum. At the look-ahead point the loss is already rising again, so the gradient there points back toward the minimum, and the update brakes early. Ordinary momentum only finds out after it has overshot, when the gradient at the new position pushes back.

In code

def nag_step(theta, v, grad, eta=0.05, gamma=0.9):
    v = gamma * v + eta * grad(theta - gamma * v)   # gradient at the look-ahead point
    return theta - v, v

Compare with the momentum code in the last lesson: only the argument of grad changes.

Does it help?

Same rate and same γ=0.9\gamma = 0.9 for momentum and NAG. Steps until the loss stays below the threshold:

ProblemPlain GDMomentumNAG
Narrow valley (η=0.05\eta = 0.05)8410249
Sigmoid toy problem (η=1\eta = 1)18314883

In both cases NAG settles faster than momentum, and faster than plain gradient descent. It brakes before the overshoot, so there is less swinging back and forth.

As always with optimisers, treat these numbers as an illustration of a mechanism on two tiny problems, not a general ranking.

Try it yourself
Optimizer playground →

Race momentum against Nesterov on the same surface.

EasyNAG

What is the only difference between momentum and Nesterov accelerated gradient?

MediumNAG

Why does looking ahead help avoid overshooting?