Adam and AdaMax

Combine momentum's memory of the gradient with RMSProp's per-parameter scaling, fix the start with bias correction, and see the max-norm variant.

Deep Learning- Fundamentals to Advanced Concepts

Momentum smooths the direction of the update. RMSProp adapts the scale for each parameter. Adam (Kingma and Ba, 2015) does both.

Two running averages

Keep an exponentially decaying average of the gradient and of its square:

mt=β1 mt−1+(1−β1) gt,vt=β2 vt−1+(1−β2) gt2m_t = \beta_1\,m_{t-1} + (1 - \beta_1)\,g_t, \qquad v_t = \beta_2\,v_{t-1} + (1 - \beta_2)\,g_t^2
  • mtm_t is the first moment: a smoothed gradient, playing the role of momentum's velocity.
  • vtv_t is the second moment: the recent size of the squared gradient, exactly RMSProp's EtE_t.

The update moves along the smoothed gradient and divides by the recent size:

θt+1=θt−η m^tv^t+ϵ\theta_{t+1} = \theta_t - \eta\,\frac{\hat m_t}{\sqrt{\hat v_t} + \epsilon}

The hats are a correction we explain next. The usual defaults are β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, ϵ=10−8\epsilon = 10^{-8} and a rate η\eta around 0.0010.001 for deep networks.

Bias correction

Both averages start at 0, so early on they are biased towards zero. After one step m1=(1−β1)g1=0.1 g1m_1 = (1 - \beta_1)g_1 = 0.1\,g_1, which is ten times too small. And v1=0.001 g12v_1 = 0.001\,g_1^2 is a thousand times too small. The correction divides by the fraction of the average that has actually been filled in:

m^t=mt1−β1 t,v^t=vt1−β2 t\hat m_t = \frac{m_t}{1 - \beta_1^{\,t}}, \qquad \hat v_t = \frac{v_t}{1 - \beta_2^{\,t}}

At t=1t = 1 this turns 0.1 g10.1\,g_1 back into g1g_1 and 0.001 g120.001\,g_1^2 back into g12g_1^2. As tt grows, βt→0\beta^t \to 0 and the correction disappears.

Why it matters, with real numbers. On the sigmoid toy problem from (−2,−2)(-2, -2) with η=0.3\eta = 0.3, the first gradient is (−0.0055,−0.0077)(-0.0055, -0.0077):

m1m_1v1v_1Update size
Without correction(−0.00055,−0.00077)(-0.00055, -0.00077)(3×10−8, 6×10−8)(3\times10^{-8},\, 6\times10^{-8})0.95 in each coordinate
With correctionm^1=g1\hat m_1 = g_1v^1=g12\hat v_1 = g_1^20.30 in each coordinate

The uncorrected first step would be about 3.2 times too large, because the tiny vv is shrunk by more than the tiny mm. With the correction the first step is exactly η\eta. On the next steps the corrected updates are about (0.29,0.30)(0.29, 0.30) and (0.28,0.30)(0.28, 0.30).

Adam in practice

Steps until the loss stays below the threshold, with the rates shown:

ProblemPlain GDAdam
Sigmoid toy problem183 (η=1\eta = 1)81 (η=0.3\eta = 0.3)
Narrow valley84 (η=0.05\eta = 0.05)94 (η=0.1\eta = 0.1)

On the toy problem Adam is more than twice as fast. On the small valley, tuned plain gradient descent is slightly faster. What Adam buys is less tuning of the scale: on the toy problem it settles within 74 to 234 steps for every rate from 0.03 to 0.3, while plain gradient descent needs 606 to 1815 steps for rates of 0.3 and 0.1 and does not converge in 3000 steps at 0.03. But like RMSProp, with a very large rate it fails (it never settles at η=3\eta = 3 and η=10\eta = 10, where plain gradient descent still does).

AdaMax

Adam's second moment is built from squares, which corresponds to the ℓ2\ell_2 norm. The AdaMax variant (from the same paper) uses the maximum norm (ℓ∞\ell_\infty) instead. Replace vtv_t by

ut=max⁡(β2 ut−1,  ∣gt∣),θt+1=θt−η1−β1 t  mtut+ϵu_t = \max\bigl(\beta_2\,u_{t-1},\; |g_t|\bigr), \qquad \theta_{t+1} = \theta_t - \frac{\eta}{1 - \beta_1^{\,t}}\;\frac{m_t}{u_t + \epsilon}

utu_t is the largest recent gradient magnitude, decaying slowly. Two consequences:

  • The step is always bounded by about η\eta, since ∣mt∣≤ut|m_t| \le u_t.
  • A max never starts near zero the way an average does, so no bias correction is needed for utu_t (only the factor for mtm_t).

On our problems AdaMax settles after 51 steps on the sigmoid toy problem (η=0.3\eta = 0.3) and 79 on the narrow valley (η=0.1\eta = 0.1).

In code

def adam_step(theta, m, v, t, grad, eta=0.1, beta1=0.9, beta2=0.999, eps=1e-8):
    g = grad(theta)
    m = beta1 * m + (1 - beta1) * g
    v = beta2 * v + (1 - beta2) * g ** 2
    m_hat = m / (1 - beta1 ** t)        # t counts from 1
    v_hat = v / (1 - beta2 ** t)
    return theta - eta * m_hat / (np.sqrt(v_hat) + eps), m, v
Try it yourself
Optimizer playground →

Compare Adam and AdaMax with the other methods, and change the learning rate.

MediumAdam

Why are the raw first and second moments biased towards zero early in training?

EasyAdam

Compute the bias-corrected first moment at t = 1 for beta1 = 0.9 and a gradient g.

EasyAdam

How does Adam combine the ideas of momentum and RMSProp?

MediumAdaMax

What does AdaMax use in place of Adam's second moment, and why is no bias correction needed for it?