From Step to Sigmoid

Why the perceptron's all-or-nothing decision is too harsh for real problems, and the smooth sigmoid neuron that replaces it.

Deep Learning- Fundamentals to Advanced Concepts

So far we have worked almost entirely with Boolean functions, where inputs and outputs are 0 or 1. Real problems rarely look like that, so we now move to functions y=f(x)y = f(x) where both xx and yy are real numbers.

Real inputs, real outputs

Suppose an oil company must decide whether to build a drilling station at a location in the ocean. The decision depends on many measurements: salinity, density, pressure, temperature, marine diversity. Each is a real number, so the input is a vector x∈Rnx \in \mathbb{R}^n. The quantity we want to predict, the amount of oil that could be extracted, is also a real number, y∈Ry \in \mathbb{R}.

Other examples have the same shape:

  • Interest rate for a loan. Inputs are salary, family size, past loans, past defaults. The output is a rate, say between 5% and 20%.
  • Images. A 100 by 100 grey image is just 10,000 pixel values, a very long input vector, and the output might be a label or score.

For Boolean functions we asked whether a network could represent the function exactly. For real functions we relax that: can a network represent it approximately, so that its output y^\hat y is close to the true yy for the examples we have? That is the quest for the rest of this module.

The step function is too harsh

Recall how a perceptron decides: it fires when the weighted sum of its inputs passes a threshold. Take a movie decision based on a single input, the critics' rating on a 0 to 1 scale, with threshold 0.5.

  • A film rated 0.49 gives output 0: dislike.
  • A film rated 0.51 gives output 1: like.

A difference of 0.02 flips the decision completely. Nobody decides that way. We would expect 0.49 to be nearly as good as 0.51.

This has nothing to do with movies or with the threshold we picked. It is a property of the function itself, a step function: flat at 0, then a sudden jump to 1.

The sigmoid neuron

We replace the step with a smooth S-shaped curve. A very common choice is the logistic function:

σ(z)=11+e−z,z=wTx+b\sigma(z) = \frac{1}{1 + e^{-z}}, \qquad z = w^{T}x + b

Here zz is the same weighted sum as before. From now on we write the bias as bb (this was w0w_0, which equals −θ-\theta, in the perceptron). A neuron that applies σ\sigma to its weighted sum is a sigmoid neuron.

A step function that jumps from 0 to 1 at z equals 0, and a smooth S-shaped sigmoid curve passing through 0.5 at z equals 0
The step jumps. The sigmoid climbs gradually and flattens at both ends.

Check its behaviour by plugging in values:

  • As z→+∞z \to +\infty, e−z→0e^{-z} \to 0, so σ(z)→1\sigma(z) \to 1.
  • As z→−∞z \to -\infty, e−z→∞e^{-z} \to \infty, so σ(z)→0\sigma(z) \to 0.
  • At z=0z = 0, σ(0)=1/(1+1)=0.5\sigma(0) = 1/(1+1) = 0.5.

In between it rises smoothly. For example σ(2)≈0.88\sigma(2) \approx 0.88 and σ(−2)≈0.12\sigma(-2) \approx 0.12. At the extremes it saturates, barely changing however large ∣z∣|z| becomes.

With one input, z=wx+bz = wx + b, the curve is centred where z=0z = 0, that is at x=−b/wx = -b/w. A larger ∣w∣|w| makes it steeper, and as ww grows it approaches the step function. Changing bb slides it left or right.

Try it yourself
Sigmoid versus step simulator →

Change w and b and watch the sigmoid sharpen into a step or flatten out.

What we gain

A graded output. The output is any number between 0 and 1, not just 0 or 1.

A probability reading. Because the output lies between 0 and 1, we can read it as a probability. An output of 0.49 means roughly a 49% chance of liking the film, and 0.51 means 51%. That is exactly the gentle behaviour the step lacked. A domain expert can then choose how to act on the probability, for example investing only above 70%.

A smooth, differentiable curve. The step is not differentiable at its jump, and everywhere else its slope is zero. The sigmoid has a slope at every point. This matters enormously: for much of this course, calculus is the main tool, and the learning algorithm ahead needs derivatives.

The logistic function is one member of the sigmoid family of S-shaped functions. Another, tanh⁡\tanh, appears later in the course.

EasySigmoid

Compute the output of a sigmoid neuron with w = 2, b = -1 for the input x = 0.5.

MediumSigmoidStep

Why does the perceptron treat a rating of 0.49 and 0.51 so differently, and how does the sigmoid change that?

MediumSigmoid

Give two reasons the sigmoid is preferred over the step function for learning.