For perceptrons we asked what a network can represent. Now the same question for sigmoid neurons, this time with real numbers and approximate answers: given an unknown function , can a network of sigmoids produce output close to ?
A sigmoid can act like a step
Recall from the first lesson of this module that for a single input is centred where . Increasing makes the S steeper. For a very large it is almost vertical:
So a sigmoid neuron with weight and bias is a step at the position . We can place the step wherever we like by choosing , and make its edge as sharp as we like by choosing .
Increase the weight until the sigmoid turns into a step.
Two steps make a tower
Take two such steps, one at and one at , with , and subtract them:
- For : both steps are 0, so the difference is 0.
- For : the first step is 1 and the second is still 0, so the difference is 1.
- For : both are 1, so the difference is 0 again.
The result is a flat-topped tower of height 1 between and , zero elsewhere. Multiplying it by a number makes the tower height .
Many towers trace any curve
Now take a target function that we want to approximate. Cut the axis into small intervals. Over each interval build a tower as above, and set its height to the value of at the middle of the interval. Add all the towers.
The sum is a staircase that follows . The thinner the towers, the closer the staircase hugs the curve, so we can approximate to any desired accuracy by using enough of them.
The network that does this is small in structure:
- Input: .
- Hidden layer: two sigmoid neurons per tower, one for each edge.
- Output: a neuron that adds them, with weight on the left edge of tower and on the right edge.
The heights are the output weights, and the positions of the edges are set by the hidden biases.
What we have shown, and what we have not
This construction is the core of the universal approximation theorem: for any reasonable function there is a sigmoid network that approximates it as closely as you like, provided it is allowed enough neurons.
Be careful about what it says:
- It is about existence. It tells us a good network exists, and how to build one in principle. It does not say gradient descent will find it.
- It is expensive. Ten towers cost 20 hidden neurons. To double the resolution you double the towers. The networks we use in practice are far more economical, and later lessons explain how.
- The edges are not perfect. With finite the step has a slope, so each tower has slightly sloped sides. More sharpness means larger weights.