Now we put the Taylor series to work. The question: standing at parameters , which small move lowers the loss the most?
The move
Write the parameters as a vector and propose a change :
Here is a direction (a vector of length 1) and is a small positive number that controls the step size. We keep small because the Taylor approximation only holds close by.
We do not yet know which to use. The Taylor series will tell us.
Gradients in one minute
A function of two variables has a slope in each direction.
For :
These are partial derivatives: the derivative with respect to one variable while treating the other as a constant. (That is why the term disappears when differentiating with respect to .)
Collect them into a vector and you get the gradient:
The formula becomes a vector of numbers when you evaluate it at a point. At the gradient is .
The derivative of the gradient, a matrix of all second partial derivatives, is called the Hessian. It is the quadratic term in the Taylor series. We do not need it in this lesson.
What Taylor says about the move
Apply the linear Taylor approximation to the new point:
The terms with , and higher are far smaller when is small, so we drop them.
A move is good only if the loss goes down:
(We can ignore here because it is positive.) So any with is an improvement. Which one is best?
The best direction
The dot product of two vectors is
where is the angle between and the gradient. With , the length is fixed by where we are standing, so all we can change is , which lies between and .
So is bounded:
It is most negative when , that is, when . So the best points in exactly the opposite direction to the gradient.
That is where the famous phrase comes from: move in the direction opposite to the gradient. The gradient points uphill, where the loss rises fastest, so the opposite direction is the steepest way downhill.
The update rule
Folding the vector length into , the rule is
or, written out for our two parameters:
The bar means "evaluated at the current values." The derivative formula is turned into a number by plugging in the current and .
The whole algorithm is a short loop:
- Start with random and .
- Repeat for a fixed number of iterations (or until the loss stops improving):
- Compute and at the current point.
- Update and using the rule above.
The update is gradient descent. The quantity is the learning rate.
See the landscape that gradient descent walks across.
Watching it go wrong: the learning rate
Take the simplest loss, . Its derivative is , so the update is
Start at and see what different learning rates do:
| Sequence of | What happens | |
|---|---|---|
| 0.1 | 1, 0.8, 0.64, 0.51, 0.41, ... | steady march towards the minimum at 0 |
| 1.0 | 1, −1, 1, −1, 1, ... | bounces forever between two points |
| 1.1 | 1, −1.2, 1.44, −1.73, 2.07, ... | diverges: each step overshoots further |
The guarantee from the Taylor series holds only for a small step. Too large a step jumps over the minimum and can make things worse.
Two more things to watch for:
- The starting point matters. A loss with several valleys sends you to whichever valley is downhill from where you start. A local minimum is not necessarily the lowest point.
- Direction is guaranteed, not distance. Gradient descent promises that a small enough step reduces the loss. It does not promise the best possible final answer.
Run plain gradient descent on the valley and try a learning rate above 0.2.