RMSProp and AdaDelta

Fix AdaGrad's ever-shrinking rate by forgetting old gradients, and see one way to remove the learning rate altogether.

Deep Learning- Fundamentals to Advanced Concepts

AdaGrad's per-parameter scaling is a good idea, but its sum GtG_t never stops growing. Both methods in this lesson repair that by forgetting.

RMSProp

Replace the running sum of squared gradients by an exponentially decaying average:

Et=β Et−1+(1−β) gt2,θt+1=θt−ηEt+ϵ  gtE_t = \beta\,E_{t-1} + (1 - \beta)\,g_t^2, \qquad \theta_{t+1} = \theta_t - \frac{\eta}{\sqrt{E_t} + \epsilon}\;g_t

with β≈0.9\beta \approx 0.9. Old gradients fade away, so EtE_t tracks the recent size of the gradient instead of growing forever. The name comes from the denominator: Et\sqrt{E_t} is a root mean square (RMS) of recent gradients.

Because the step is η\eta times "gradient divided by its typical size", each parameter moves by roughly η\eta per step whatever the scale of its gradient. A parameter with tiny gradients and one with huge gradients both take steps of about the same size.

How it behaves

On the sigmoid toy problem, RMSProp with η=0.1\eta = 0.1 has the loss below 10−310^{-3} for good after roughly 100 to 160 steps (RMSProp keeps jittering close to the threshold, so the exact step changes with tiny numerical differences), against 183 for plain gradient descent at its best rate of 1.

There is a catch in that scale-free step. Near the minimum, the gradient becomes small, but EtE_t shrinks along with it, so the step stays about η\eta in size. The parameters then bounce around the minimum instead of settling. On the narrow valley with η=0.05\eta = 0.05:

  • the loss never stays below 10−410^{-4}: after several hundred steps it sits at 0.00340.0034,
  • the parameters end up alternating around (±0.025,±0.025)(\pm 0.025, \pm 0.025), which is η/2\eta/2 either side of the minimum.

The remedy is a schedule from the previous lesson. Decaying the rate as ηk=0.05/(1+0.01k)\eta_k = 0.05/(1 + 0.01k) brings the loss to 7.7×10−67.7 \times 10^{-6} after 2000 steps, instead of stalling at 3.4×10−33.4 \times 10^{-3}.

The learning rate still matters

The rate means something different here: it is approximately a distance per step. That makes it easier to choose when gradients vary widely in scale, but it still has to fit the problem. The table shows the steps needed on the sigmoid toy problem (blank means it never settled in 3000 steps, or diverged):

Rate η\etaPlain GDRMSProp
0.03none129
0.11815about 100 to 160
0.3606about 3000 (barely settles)
1183none
362none
1015none

Plain gradient descent needs a large rate here because the gradients on this problem are small. RMSProp is happy with small rates but, with η\eta of 1 or more, takes steps larger than the whole problem and never settles. There is no rate that is best everywhere.

AdaDelta

RMSProp's update is unit-inconsistent: the parameter change should have the units of the parameter, but η g/E\eta\,g/\sqrt{E} has the units of η\eta. AdaDelta (Zeiler, 2012) fixes this by replacing η\eta with a running RMS of past updates, so no learning rate is needed:

E[g2]t=ρ E[g2]t−1+(1−ρ) gt2E[g^2]_t = \rho\,E[g^2]_{t-1} + (1-\rho)\,g_t^2 Δθt=−E[Δθ2]t−1+ϵE[g2]t+ϵ  gt,E[Δθ2]t=ρ E[Δθ2]t−1+(1−ρ) Δθt2\Delta\theta_t = -\frac{\sqrt{E[\Delta\theta^2]_{t-1} + \epsilon}}{\sqrt{E[g^2]_t + \epsilon}}\;g_t, \qquad E[\Delta\theta^2]_t = \rho\,E[\Delta\theta^2]_{t-1} + (1-\rho)\,\Delta\theta_t^2

and θt+1=θt+Δθt\theta_{t+1} = \theta_t + \Delta\theta_t. Here ρ≈0.95\rho \approx 0.95.

So the method needs no initial learning rate. The catch is how it starts. At step 1 the average of past updates is 0, so the numerator is ϵ\sqrt\epsilon, and with ϵ=10−6\epsilon = 10^{-6} the first steps are tiny, about 0.0045 per parameter on the valley and about 0.0035 to 0.0039 on the toy problem, whatever the gradient. They grow as the running average of updates builds up. The result on these problems is a slow start:

ProblemLoss after 100 stepsSteps to stay below threshold
Narrow valley1.96never within 3000
Sigmoid toy problem0.40348

The constant ϵ\epsilon quietly plays the role of the learning rate at the start. A larger ϵ\epsilon gives bigger initial steps.

Summary so far

MethodWhat it keepsResult
AdaGradsum of squared gradientsrate shrinks forever
RMSPropdecaying average of squared gradientsrate adapts, but steps stay about η\eta
AdaDeltadecaying averages of gradients and of updatesno learning rate to tune, slow start
Try it yourself
Optimizer playground →

Try the preset where RMSProp jitters, then add a decay schedule.

EasyRMSProp

What is the one change from AdaGrad to RMSProp, and what problem does it fix?

HardRMSProp

Why can RMSProp with a constant learning rate fail to settle at the minimum?

MediumAdaDelta

Why does AdaDelta start slowly?