Reading Contour Plots

Contour maps show a loss surface from above. Learn to read them, and see why plain gradient descent zigzags in a narrow valley.

Deep Learning- Fundamentals to Advanced Concepts

Plain gradient descent works, but it can be slow or unstable. To understand why, and how the methods in this module fix it, we need a way to look at loss surfaces on a flat page. That tool is the contour plot.

From a surface to a map

A loss surface L(w,b)L(w, b) is a landscape: the height above each point (w,b)(w, b) is the loss. A contour plot is the same landscape seen from directly above, the way a hiking map shows a hill. Each closed curve, a contour line or level curve, joins all the points with the same loss value.

Reading one is like reading a hiking map:

  • Rings close together mean the ground changes quickly: the surface is steep.
  • Rings far apart mean the ground is nearly flat.
  • The smallest ring is at the bottom of the bowl, the minimum.
  • Moving along a contour, the loss does not change. Moving across contours, it does.

The gradient is perpendicular to the contour

Along a contour line the loss stays constant, so the slope in that direction is zero. The gradient, which points in the direction of steepest increase, therefore has no component along the contour. It must be perpendicular to it, pointing across the contour towards higher loss.

A gradient descent step is the opposite direction: straight across the contour lines, downhill.

A narrow valley

We will use one test surface again and again in this module:

L(w,b)=12(10 w2+b2)L(w, b) = \tfrac12\bigl(10\,w^2 + b^2\bigr)

It has its minimum at (0,0)(0, 0). The loss changes ten times faster along ww than along bb. The contours are ellipses stretched along the bb direction, a long narrow valley with steep walls.

Elongated elliptical contour lines with a zigzag gradient descent path crossing the valley and slowly descending
Gradient descent from (1, 1) with learning rate 0.18. The path bounces from wall to wall while creeping along the valley floor.

The gradient is (10w,  b)(10w,\; b), so gradient descent updates each coordinate separately:

w←(1−10η) w,b←(1−η) bw \leftarrow (1 - 10\eta)\,w, \qquad b \leftarrow (1 - \eta)\,b

Take η=0.18\eta = 0.18 from the start (1,1)(1, 1):

Stepwwbb
011
1−0.80.82
20.640.67
3−0.510.55
40.410.45
5−0.330.37

The steep coordinate ww flips sign at every step (it is multiplied by −0.8-0.8), which is the zigzag across the valley. The gentle coordinate bb shrinks by a factor of 0.82 and makes steady but slow progress.

Why you cannot just raise the rate

The steep direction sets a speed limit. ww shrinks only if ∣1−10η∣<1|1 - 10\eta| < 1, so we need η<0.2\eta < 0.2. At η=0.21\eta = 0.21 the factor is −1.1-1.1, the zigzag grows, and gradient descent diverges. With a smaller rate such as η=0.05\eta = 0.05 the path is smooth, but bb is multiplied by 0.95 each step and crawls.

For a bowl like this there is a best choice, η=2/(10+1)≈0.18\eta = 2/(10 + 1) \approx 0.18, which is where we started. Even then each step shrinks both coordinates by only a factor of (κ−1)/(κ+1)=9/11≈0.82(\kappa - 1)/(\kappa + 1) = 9/11 \approx 0.82, where κ=10\kappa = 10 is the ratio of the steepest to the gentlest curvature. On this surface that still finishes quickly: the loss stays below 10−410^{-4} after 25 steps (η=0.18\eta = 0.18), against 84 steps with η=0.05\eta = 0.05.

But the ratio κ\kappa can be huge for real networks. With κ=1000\kappa = 1000, the best factor is 999/1001≈0.998999/1001 \approx 0.998, and shrinking an error by a factor of 10410^4 takes thousands of steps. Steep in one direction and flat in another is the core difficulty that the methods in this module attack, in different ways.

Flat regions too

Contours also explain the slow start we saw for the sigmoid neuron. Near the starting point (−2,−2)(-2, -2) the rings are far apart: the surface is nearly flat, the gradient is tiny (≈0.005\approx 0.005), and gradient descent barely moves. A method that builds up speed on flat ground would escape sooner. That is the idea of momentum, in the next lesson.

Try it yourself
Optimizer playground →

Read the contour map, click to choose a start, and watch gradient descent cross the valley.

EasyContours

On a contour plot, what does it mean when contour lines are very close together?

MediumContoursGradient

Why is the gradient perpendicular to the contour line?

MediumLearning rate

For L = (10 w^2 + b^2)/2, what happens to w at each gradient descent step with learning rate 0.21?