Gradient Descent for a Sigmoid Neuron

Work out the two partial derivatives of the squared-error loss, turn them into code, and watch the loss fall on the toy problem.

Deep Learning- Fundamentals to Advanced Concepts

The update rule needs two numbers at each step, ∂L/∂w\partial L/\partial w and ∂L/∂b\partial L/\partial b. We now derive them for our sigmoid neuron and then run the algorithm.

The loss for one example

The model is f(x)=σ(wx+b)f(x) = \sigma(wx + b). For a single training example (x,y)(x, y) the loss is

L=12 (f(x)−y)2L = \frac{1}{2}\,\bigl(f(x) - y\bigr)^2

Derivative with respect to ww

Use the chain rule. The derivative of a square is twice the thing times the derivative of the thing, and the 12\tfrac12 cancels the 2:

∂L∂w=(f(x)−y) ∂f(x)∂w\frac{\partial L}{\partial w} = \bigl(f(x) - y\bigr)\,\frac{\partial f(x)}{\partial w}

The true value yy is a constant, so it contributes nothing to the derivative. What remains is the derivative of the sigmoid. Write z=wx+bz = wx + b. The sigmoid is σ(z)=(1+e−z)−1\sigma(z) = (1 + e^{-z})^{-1}, so

σ′(z)=e−z(1+e−z)2=11+e−z⋅e−z1+e−z=σ(z)(1−σ(z))\sigma'(z) = \frac{e^{-z}}{(1 + e^{-z})^{2}} = \frac{1}{1 + e^{-z}}\cdot\frac{e^{-z}}{1 + e^{-z}} = \sigma(z)\bigl(1 - \sigma(z)\bigr)

This is a neat property: the derivative of the sigmoid is expressed using the sigmoid's own output. With f(x)=σ(z)f(x) = \sigma(z) and the chain rule factor ∂z/∂w=x\partial z/\partial w = x:

∂f(x)∂w=f(x)(1−f(x)) x\frac{\partial f(x)}{\partial w} = f(x)\bigl(1 - f(x)\bigr)\,x

Putting it together:

  ∂L∂w=(f(x)−y) f(x)(1−f(x)) x  \boxed{\;\frac{\partial L}{\partial w} = \bigl(f(x) - y\bigr)\,f(x)\bigl(1 - f(x)\bigr)\,x\;}

Derivative with respect to bb

The same steps apply, except ∂z/∂b=1\partial z / \partial b = 1, so the factor xx disappears:

  ∂L∂b=(f(x)−y) f(x)(1−f(x))  \boxed{\;\frac{\partial L}{\partial b} = \bigl(f(x) - y\bigr)\,f(x)\bigl(1 - f(x)\bigr)\;}

Many examples

With NN examples the loss is a sum, and the derivative of a sum is the sum of the derivatives:

∂L∂w=∑i=1N(f(xi)−yi) f(xi)(1−f(xi)) xi\frac{\partial L}{\partial w} = \sum_{i=1}^{N} \bigl(f(x_i) - y_i\bigr)\,f(x_i)\bigl(1 - f(x_i)\bigr)\,x_i ∂L∂b=∑i=1N(f(xi)−yi) f(xi)(1−f(xi))\frac{\partial L}{\partial b} = \sum_{i=1}^{N} \bigl(f(x_i) - y_i\bigr)\,f(x_i)\bigl(1 - f(x_i)\bigr)

Notice how much is shared: the two formulas differ only by the final xix_i.

In code

We train on the toy data from before, the points (0.5,0.2)(0.5, 0.2) and (2.5,0.9)(2.5, 0.9), starting from w=b=−2w = b = -2 with learning rate η=1\eta = 1:

import numpy as np

X = np.array([0.5, 2.5])
Y = np.array([0.2, 0.9])

def f(x, w, b):
    return 1 / (1 + np.exp(-(w * x + b)))

def loss(w, b):
    return 0.5 * np.sum((f(X, w, b) - Y) ** 2)

def gradients(w, b):
    pred = f(X, w, b)
    common = (pred - Y) * pred * (1 - pred)   # shared by both derivatives
    return np.sum(common * X), np.sum(common)

w, b, eta = -2.0, -2.0, 1.0
for epoch in range(1, 1001):
    dw, db = gradients(w, b)
    w -= eta * dw
    b -= eta * db
    if epoch in (1, 10, 100, 500, 1000):
        print(epoch, round(w, 3), round(b, 3), round(loss(w, b), 5))

Output:

1 -1.995 -1.992 0.41573
10 -1.942 -1.92 0.41484
100 1.003 -1.051 0.01774
500 1.785 -2.272 0.0
1000 1.792 -2.282 0.0

Reading the run

  • The loss only goes down. Unlike guesswork, no step takes us uphill.
  • The start is slow. From (−2,−2)(-2, -2) the gradient is tiny, about (−0.0055,−0.0077)(-0.0055, -0.0077), because both sigmoids are saturated there. The loss barely moves in the first ten epochs (0.4157 to 0.4148). Then the path reaches steeper ground and the loss drops fast. By epoch 100 it is 0.018.
  • It ends where we wanted. The final w≈1.79w \approx 1.79, b≈−2.28b \approx -2.28 is almost exactly the parameter pair we reached by hand earlier, (1.78,−2.27)(1.78, -2.27), and the loss is essentially 0.
The same error surface as before with the path of gradient descent drawn from the start point in the upper left corner towards the dark low-loss region
The path taken by gradient descent. Steps are small where the surface is flat and larger on steeper ground.
Try it yourself
Error surfaces in 3D →

Run gradient descent on a loss surface and watch the point roll towards the minimum.

Completing the cycle

The story now mirrors the perceptron's:

PerceptronSigmoid neuron
Modelstep on a weighted sumsigmoid on a weighted sum
Errormisclassified pointssquared-error loss
Learning algorithmperceptron learning algorithmgradient descent
Ran it and watched it convergeyesyes

Next, as we did for the perceptron, we ask what this kind of neuron can represent.

MediumSigmoidDerivative

Why is the derivative of the sigmoid convenient for computation?

EasyGradient descent

What is the only difference between the formulas for the derivative with respect to w and with respect to b?

HardGradient descentSaturation

Why did the loss decrease so slowly at the start of the run?