Sigmoid Towers in Two Dimensions

Extend the tower trick to two inputs with four sigmoids, a sum and a final sigmoid, and see how the cost grows with the number of inputs.

Deep Learning- Fundamentals to Advanced Concepts

Real problems have many inputs. Let us see whether the tower idea survives with two.

The situation

Return to the oil example with two inputs, salinity x1x_1 and pressure x2x_2. At each location oil was found (output 1) or not (output 0). Plot the locations on the x1x_1–x2x_2 plane and suppose the two groups are mixed in a way no straight line can separate. The true function we wish to learn is 1 in some regions and 0 in others.

Approximating it calls for a two-dimensional tower: a block that is 1 inside a small rectangle and 0 everywhere else. If we can build one, we can place many side by side, with chosen heights and sizes, to match any shape.

One sigmoid in two dimensions

A sigmoid neuron with two inputs computes σ(w1x1+w2x2+b)\sigma(w_1 x_1 + w_2 x_2 + b). Its parameters act like this:

  • w1w_1 controls the slope along the x1x_1 direction. Raising it from 0 turns a flat sheet into a gentle slope, then a steep one, and finally a cliff.
  • w2w_2 does the same along x2x_2.
  • bb shifts the line where the sigmoid crosses 0.5.

Set w2=0w_2 = 0 and make w1w_1 very large and you get a step along x1x_1: a wall running across the plane at one value of x1x_1. Set w1=0w_1 = 0 instead and the wall runs at one value of x2x_2.

Building the tower

Network diagram: two inputs feed four sigmoid neurons, which are subtracted in pairs, added, and passed through a final sigmoid to give a tower
Four steep sigmoids, two subtractions, one addition and a final sigmoid.

Step 1: a strip in x1x_1. Take two steps along x1x_1, at positions a1a_1 and c1c_1, and subtract them. Just as in one dimension this gives a band that is 1 for a1<x1<c1a_1 < x_1 < c_1 and 0 elsewhere. In two dimensions it is an open tower: a long wall that runs on forever in the x2x_2 direction.

Step 2: a strip in x2x_2. Do the same along x2x_2, using two more sigmoids, to get a band that is 1 for a2<x2<c2a_2 < x_2 < c_2, running forever in the x1x_1 direction.

Step 3: add the strips. The sum has three levels:

  • 0 outside both strips,
  • 1 inside exactly one strip (the arms of a cross),
  • 2 only where the strips overlap, which is the rectangle we want.

This sum is a tower sitting on top of some unwanted scaffolding of height 1.

Step 4: keep only level 2. Pass the sum through one more sigmoid whose switch-over lies between 1 and 2, for example σ(k (S−1.5))\sigma\bigl(k\,(S - 1.5)\bigr) with kk large. Anything at level 0 or 1 is pushed to 0, and only level 2 becomes 1. What is left is a clean, closed tower over the rectangle.

A numerical check for a square on [0.3,0.7]2[0.3, 0.7]^2 with large kk confirms it: inside the square the output is 1, and outside it is essentially 0 (below 10−4010^{-40}). The switch-over at 1.5, rather than at exactly 1, keeps the level-1 scaffolding safely on the 0 side.

Try it yourself
Tower functions in 3D →

Rotate the tower surface and see how walls combine into a block.

Counting the cost

  • In one dimension a tower needs 2 sigmoids.
  • In two dimensions it needs 4 in the first layer (two per input), then an addition, then a final sigmoid.
  • In nn dimensions, 2n2n sigmoids in the first layer, plus the combination stages.

Then you need many towers to cover the space. If each axis is cut into mm pieces, there are about mnm^n towers. That grows exponentially with the number of inputs. It is the same kind of explosion we saw for Boolean functions, and it is why this construction proves that approximation is possible, not that it is efficient.

The construction uses two layers of sigmoids between input and output. The universal approximation theorem itself holds for a single hidden layer, but its proof is more involved, and this illustration is enough to see why the claim is believable.

What it means for learning

Any function, however complicated, can in principle be approximated as precisely as you like by a network of sigmoid neurons with enough of them. That is the foundation of why deep networks are such general tools: we can write down a flexible family f^(x;θ)\hat f(x; \theta), and gradient descent searches for good parameters within it.

MediumUniversal approximation

Why does one sigmoid neuron with weight w2 = 0 and large w1 produce a wall?

MediumUniversal approximation

In the 2D tower, why is a final sigmoid needed after adding the two strips?

HardUniversal approximationDimensionality

Why does this construction become impractical for many inputs?