Scheduling the Learning Rate

Why one fixed rate is rarely best, and the common ways to change it over training: decay, line search, warm-up and cosine annealing.

Deep Learning- Fundamentals to Advanced Concepts

Until now the learning rate η\eta has been one fixed number. Yet the ideal step size changes during training. Early on we want large steps to cover ground quickly. Near the end we want small steps so the parameters can settle instead of bouncing around the minimum. A schedule changes η\eta as training goes on.

Decay schedules

The simplest idea is to start with a rate η0\eta_0 and shrink it. Three common forms, where kk is the step number:

ScheduleFormulaBehaviour
Step decayηk=η0 d⌊k/s⌋\eta_k = \eta_0\,d^{\lfloor k/s \rfloor}constant for ss steps, then multiplied by dd (for example halved every 20 steps)
Exponential decayηk=η0 e−λk\eta_k = \eta_0\,e^{-\lambda k}smooth, fast shrink
Inverse-time decayηk=η01+λk\eta_k = \dfrac{\eta_0}{1 + \lambda k}smooth, slower shrink

With η0=0.1\eta_0 = 0.1, halving every 20 steps, λ=0.03\lambda = 0.03 for the exponential and λ=0.05\lambda = 0.05 for inverse-time, the rates at a few steps are:

Step kkStep decayExponentialInverse-time
00.10.10.1
200.050.0550.05
500.0250.0220.029
1000.0030.0050.017
Four curves of learning rate against training step: step decay, exponential decay, inverse-time decay, and a warm-up followed by cosine
Different schedules for the same starting rate of 0.1 over 100 steps.

The decay speed is a hyperparameter like any other. Too fast and training stalls before it has finished learning. Too slow and the noise near the end is not reduced.

Line search: choosing the rate each step

Instead of a fixed plan, we can search for a good step size at every update. Take the gradient direction g=∇L(θ)g = \nabla L(\theta) and try a large step. If it does not reduce the loss enough, shrink it and try again. A standard test, the backtracking rule, accepts α\alpha when

L(θ−αg)≤L(θ)−c α ∥g∥2L(\theta - \alpha g) \le L(\theta) - c\,\alpha\,\|g\|^2

for a small constant cc such as 10−410^{-4}: the loss must drop by a fixed fraction of what the slope predicts. Start with α=1\alpha = 1 and halve until the test passes.

On the narrow valley, starting from (1,1)(1, 1), this rule picks different rates at different steps:

IterationRate chosenLoss after the step
10.1250.695
20.1250.313
30.250.209
40.250.192
50.1250.077
60.250.054
70.250.054
80.1250.019

The loss started at 5.5 and falls at every step without us choosing a rate by hand. The price is extra loss evaluations in every trial. For a large network each trial costs a forward pass over the data, and with mini-batches the loss is itself noisy, so line search is rarely used for training deep networks. Fixed schedules are cheaper.

Warm-up

Sometimes the start is the dangerous part. At the very beginning the weights are random, gradient estimates can be wild, and methods that keep running averages (like those in the coming lessons) have had little data to average. A large rate right away can throw training off.

Warm-up starts with a tiny rate and raises it linearly to the target over the first WW steps:

ηt=ηmax⁡⋅t+1W(t<W)\eta_t = \eta_{\max}\cdot\frac{t + 1}{W} \qquad (t < W)

Cosine annealing

After warm-up, we can lower the rate smoothly along a half cosine wave, from ηmax⁡\eta_{\max} down to ηmin⁡\eta_{\min} over TT steps:

ηt=ηmin⁡+12(ηmax⁡−ηmin⁡)(1+cos⁡π (t−W)T−W)\eta_t = \eta_{\min} + \tfrac12\left(\eta_{\max} - \eta_{\min}\right)\left(1 + \cos\frac{\pi\,(t - W)}{T - W}\right)

It falls slowly at first, fastest in the middle, and slowly again at the end, which lets the parameters settle gently. Some schemes restart the cosine cycle several times.

With ηmax⁡=0.1\eta_{\max} = 0.1, ηmin⁡=0\eta_{\min} = 0, W=10W = 10 and T=100T = 100, the rate at steps 0, 4, 9, 10, 32, 55 and 99 is

0.01,  0.05,  0.10,  0.10,  0.086,  0.05,  0.000.01,\; 0.05,\; 0.10,\; 0.10,\; 0.086,\; 0.05,\; 0.00

It ramps up, peaks at step 10, and then declines to zero by the end. A warm-up followed by cosine decay is a very common recipe for training large networks.

In code

import math

def step_decay(k, eta0=0.1, drop=0.5, every=20):
    return eta0 * drop ** (k // every)

def warmup_cosine(t, eta_max=0.1, eta_min=0.0, warmup=10, total=100):
    if t < warmup:
        return eta_max * (t + 1) / warmup
    progress = (t - warmup) / (total - warmup)
    return eta_min + 0.5 * (eta_max - eta_min) * (1 + math.cos(math.pi * progress))

print([round(warmup_cosine(t), 3) for t in (0, 4, 9, 10, 32, 55, 99)])

Inside a training loop, just set eta = warmup_cosine(step) before each update.

Try it yourself
Optimizer playground →

Pick a learning rate schedule and see how it calms RMSProp near the minimum.

EasySchedules

Why might you want a large learning rate early in training and a small one later?

MediumLine search

Why is line search seldom used to train large neural networks?

MediumWarm-up

What problem does warm-up address?