Real problems have many inputs. Let us see whether the tower idea survives with two.
The situation
Return to the oil example with two inputs, salinity and pressure . At each location oil was found (output 1) or not (output 0). Plot the locations on the – plane and suppose the two groups are mixed in a way no straight line can separate. The true function we wish to learn is 1 in some regions and 0 in others.
Approximating it calls for a two-dimensional tower: a block that is 1 inside a small rectangle and 0 everywhere else. If we can build one, we can place many side by side, with chosen heights and sizes, to match any shape.
One sigmoid in two dimensions
A sigmoid neuron with two inputs computes . Its parameters act like this:
- controls the slope along the direction. Raising it from 0 turns a flat sheet into a gentle slope, then a steep one, and finally a cliff.
- does the same along .
- shifts the line where the sigmoid crosses 0.5.
Set and make very large and you get a step along : a wall running across the plane at one value of . Set instead and the wall runs at one value of .
Building the tower
Step 1: a strip in . Take two steps along , at positions and , and subtract them. Just as in one dimension this gives a band that is 1 for and 0 elsewhere. In two dimensions it is an open tower: a long wall that runs on forever in the direction.
Step 2: a strip in . Do the same along , using two more sigmoids, to get a band that is 1 for , running forever in the direction.
Step 3: add the strips. The sum has three levels:
- 0 outside both strips,
- 1 inside exactly one strip (the arms of a cross),
- 2 only where the strips overlap, which is the rectangle we want.
This sum is a tower sitting on top of some unwanted scaffolding of height 1.
Step 4: keep only level 2. Pass the sum through one more sigmoid whose switch-over lies between 1 and 2, for example with large. Anything at level 0 or 1 is pushed to 0, and only level 2 becomes 1. What is left is a clean, closed tower over the rectangle.
A numerical check for a square on with large confirms it: inside the square the output is 1, and outside it is essentially 0 (below ). The switch-over at 1.5, rather than at exactly 1, keeps the level-1 scaffolding safely on the 0 side.
Rotate the tower surface and see how walls combine into a block.
Counting the cost
- In one dimension a tower needs 2 sigmoids.
- In two dimensions it needs 4 in the first layer (two per input), then an addition, then a final sigmoid.
- In dimensions, sigmoids in the first layer, plus the combination stages.
Then you need many towers to cover the space. If each axis is cut into pieces, there are about towers. That grows exponentially with the number of inputs. It is the same kind of explosion we saw for Boolean functions, and it is why this construction proves that approximation is possible, not that it is efficient.
The construction uses two layers of sigmoids between input and output. The universal approximation theorem itself holds for a single hidden layer, but its proof is more involved, and this illustration is enough to see why the claim is believable.
What it means for learning
Any function, however complicated, can in principle be approximated as precisely as you like by a network of sigmoid neurons with enough of them. That is the foundation of why deep networks are such general tools: we can write down a flexible family , and gradient descent searches for good parameters within it.