Gradient descent is only as good as its starting point. In a deep network, a poor start can make training crawl or fail entirely, no matter how clever the optimizer.
Two questions, one starting point
Researchers often ask whether a technique helps optimization (training reaches a low loss faster) or regularization (the final model generalises better). Good initialization is mainly about optimization: it decides whether gradient descent can get moving at all. We will meet an example (pre-training) in the next lesson that does both.
Bad starts
We already know one bad start: all weights equal (for example all zero). Every neuron in a layer computes the same thing and gets the same gradient, so they stay identical forever. Random initial weights break the symmetry.
But "random" is not enough. The scale of the random numbers matters, because each layer multiplies the signal by a weight matrix and then applies the activation. After many layers the effects compound:
- Weights too small: the signal shrinks layer by layer and vanishes. The gradients going backwards shrink the same way, so the early layers learn nothing.
- Weights too large: with a saturating activation such as tanh or sigmoid, the pre-activations are huge and every neuron sits at its extreme, where the slope is zero. The signal is stuck at and the gradients vanish again. With ReLU, the signal can instead explode.
A variance rule
Consider one neuron that sums inputs, each multiplied by a random weight of variance :
(assuming independent, zero-mean weights). To stop the signal from shrinking or growing from layer to layer we want to be about the same as the variance of the inputs. That requires , which gives the rule for saturating activations such as tanh:
This is usually called Xavier or Glorot initialization (Glorot and Bengio's version uses , which averages the forward and backward requirements).
ReLU sets roughly half of its inputs to zero, which halves the signal's second moment at each layer. To compensate we double the weight variance. This is He initialization:
Seeing it
We feed 1000 random inputs through a stack of 10 layers, each 100 neurons wide, and print the standard deviation of the activations at layers 1, 5 and 10:
import numpy as np, math
def layer_stds(weight_std, activation, layers=10, width=100, seed=0):
rng = np.random.default_rng(seed)
h = rng.normal(0, 1, (1000, width)) # 1000 random inputs
stds = []
for _ in range(layers):
W = rng.normal(0, weight_std(width), (width, width))
h = activation(h @ W)
stds.append(h.std())
return stds
relu = lambda z: np.maximum(0, z)
settings = [
("tanh, weight std 0.01 ", lambda n: 0.01, np.tanh),
("tanh, weight std 1 ", lambda n: 1.0, np.tanh),
("tanh, Xavier 1/sqrt(n) ", lambda n: 1 / math.sqrt(n), np.tanh),
("ReLU, weight std 1/sqrt(n)", lambda n: 1 / math.sqrt(n), relu),
("ReLU, He sqrt(2/n) ", lambda n: math.sqrt(2 / n), relu),
]
for name, std_fn, act in settings:
s = layer_stds(std_fn, act)
print(f"{name}: layer 1 {s[0]:.4f} layer 5 {s[4]:.4g} layer 10 {s[9]:.4g}")
Output:
tanh, weight std 0.01 : layer 1 0.0987 layer 5 9.801e-06 layer 10 1.045e-10
tanh, weight std 1 : layer 1 0.9587 layer 5 0.9584 layer 10 0.9565
tanh, Xavier 1/sqrt(n) : layer 1 0.6255 layer 5 0.3185 layer 10 0.2319
ReLU, weight std 1/sqrt(n): layer 1 0.5830 layer 5 0.1767 layer 10 0.03273
ReLU, He sqrt(2/n) : layer 1 0.8245 layer 5 0.9998 layer 10 1.047
Reading the results:
- Tanh with tiny weights (0.01): the activations fall from 0.1 to by layer 10. The signal has vanished.
- Tanh with large weights (1): the standard deviation sits at 0.96, close to the maximum possible value of 1 for tanh. Almost every neuron is saturated at , where the slope is nearly zero, so gradients cannot flow.
- Tanh with Xavier: the signal shrinks gently (0.63 to 0.23) but stays useful. The shrinkage is because tanh itself compresses values.
- ReLU with the Xavier scale: the signal decays (0.58 to 0.033), as the derivation predicted, since half the units are off.
- ReLU with He: the standard deviation stays near 1 through all ten layers (0.82, 1.00, 1.05).
The backward pass has the same problem: gradients are multiplied by weight matrices as they travel back, which is why a good initialization helps the gradients as much as the activations.
Practical rules
- Biases are usually initialised to zero. Symmetry is broken by the weights.
- Use Xavier/Glorot for tanh or sigmoid layers and He for ReLU-type layers. Deep learning libraries apply sensible defaults, but it helps to know what they are doing.
- If a deep network's loss does not move at all at the start, check the initial scale.
Switch to the deep network view and change the weight scale.