Momentum smooths the direction of the update. RMSProp adapts the scale for each parameter. Adam (Kingma and Ba, 2015) does both.
Two running averages
Keep an exponentially decaying average of the gradient and of its square:
- is the first moment: a smoothed gradient, playing the role of momentum's velocity.
- is the second moment: the recent size of the squared gradient, exactly RMSProp's .
The update moves along the smoothed gradient and divides by the recent size:
The hats are a correction we explain next. The usual defaults are , , and a rate around for deep networks.
Bias correction
Both averages start at 0, so early on they are biased towards zero. After one step , which is ten times too small. And is a thousand times too small. The correction divides by the fraction of the average that has actually been filled in:
At this turns back into and back into . As grows, and the correction disappears.
Why it matters, with real numbers. On the sigmoid toy problem from with , the first gradient is :
| Update size | |||
|---|---|---|---|
| Without correction | 0.95 in each coordinate | ||
| With correction | 0.30 in each coordinate |
The uncorrected first step would be about 3.2 times too large, because the tiny is shrunk by more than the tiny . With the correction the first step is exactly . On the next steps the corrected updates are about and .
Adam in practice
Steps until the loss stays below the threshold, with the rates shown:
| Problem | Plain GD | Adam |
|---|---|---|
| Sigmoid toy problem | 183 () | 81 () |
| Narrow valley | 84 () | 94 () |
On the toy problem Adam is more than twice as fast. On the small valley, tuned plain gradient descent is slightly faster. What Adam buys is less tuning of the scale: on the toy problem it settles within 74 to 234 steps for every rate from 0.03 to 0.3, while plain gradient descent needs 606 to 1815 steps for rates of 0.3 and 0.1 and does not converge in 3000 steps at 0.03. But like RMSProp, with a very large rate it fails (it never settles at and , where plain gradient descent still does).
AdaMax
Adam's second moment is built from squares, which corresponds to the norm. The AdaMax variant (from the same paper) uses the maximum norm () instead. Replace by
is the largest recent gradient magnitude, decaying slowly. Two consequences:
- The step is always bounded by about , since .
- A max never starts near zero the way an average does, so no bias correction is needed for (only the factor for ).
On our problems AdaMax settles after 51 steps on the sigmoid toy problem () and 79 on the narrow valley ().
In code
def adam_step(theta, m, v, t, grad, eta=0.1, beta1=0.9, beta2=0.999, eps=1e-8):
g = grad(theta)
m = beta1 * m + (1 - beta1) * g
v = beta2 * v + (1 - beta2) * g ** 2
m_hat = m / (1 - beta1 ** t) # t counts from 1
v_hat = v / (1 - beta2 ** t)
return theta - eta * m_hat / (np.sqrt(v_hat) + eps), m, v
Compare Adam and AdaMax with the other methods, and change the learning rate.