Initialization: Starting Training Well

Why the starting weights decide whether signals and gradients survive a deep network, and the variance rules that keep them healthy.

Deep Learning- Fundamentals to Advanced Concepts

Gradient descent is only as good as its starting point. In a deep network, a poor start can make training crawl or fail entirely, no matter how clever the optimizer.

Two questions, one starting point

Researchers often ask whether a technique helps optimization (training reaches a low loss faster) or regularization (the final model generalises better). Good initialization is mainly about optimization: it decides whether gradient descent can get moving at all. We will meet an example (pre-training) in the next lesson that does both.

Bad starts

We already know one bad start: all weights equal (for example all zero). Every neuron in a layer computes the same thing and gets the same gradient, so they stay identical forever. Random initial weights break the symmetry.

But "random" is not enough. The scale of the random numbers matters, because each layer multiplies the signal by a weight matrix and then applies the activation. After many layers the effects compound:

  • Weights too small: the signal shrinks layer by layer and vanishes. The gradients going backwards shrink the same way, so the early layers learn nothing.
  • Weights too large: with a saturating activation such as tanh or sigmoid, the pre-activations are huge and every neuron sits at its extreme, where the slope is zero. The signal is stuck at ±1\pm 1 and the gradients vanish again. With ReLU, the signal can instead explode.

A variance rule

Consider one neuron that sums nn inputs, each multiplied by a random weight of variance Var(w)\mathrm{Var}(w):

z=∑i=1nwi hi⇒Var(z)=n Var(w) E[h2]z = \sum_{i=1}^{n} w_i\,h_i \quad\Rightarrow\quad \mathrm{Var}(z) = n\,\mathrm{Var}(w)\,\mathbb{E}[h^2]

(assuming independent, zero-mean weights). To stop the signal from shrinking or growing from layer to layer we want Var(z)\mathrm{Var}(z) to be about the same as the variance of the inputs. That requires n Var(w)≈1n\,\mathrm{Var}(w) \approx 1, which gives the rule for saturating activations such as tanh:

Var(w)=1nin\mathrm{Var}(w) = \frac{1}{n_{\text{in}}}

This is usually called Xavier or Glorot initialization (Glorot and Bengio's version uses 2/(nin+nout)2/(n_{\text{in}} + n_{\text{out}}), which averages the forward and backward requirements).

ReLU sets roughly half of its inputs to zero, which halves the signal's second moment at each layer. To compensate we double the weight variance. This is He initialization:

Var(w)=2nin\mathrm{Var}(w) = \frac{2}{n_{\text{in}}}

Seeing it

We feed 1000 random inputs through a stack of 10 layers, each 100 neurons wide, and print the standard deviation of the activations at layers 1, 5 and 10:

import numpy as np, math

def layer_stds(weight_std, activation, layers=10, width=100, seed=0):
    rng = np.random.default_rng(seed)
    h = rng.normal(0, 1, (1000, width))            # 1000 random inputs
    stds = []
    for _ in range(layers):
        W = rng.normal(0, weight_std(width), (width, width))
        h = activation(h @ W)
        stds.append(h.std())
    return stds

relu = lambda z: np.maximum(0, z)
settings = [
    ("tanh,  weight std 0.01   ", lambda n: 0.01, np.tanh),
    ("tanh,  weight std 1      ", lambda n: 1.0, np.tanh),
    ("tanh,  Xavier 1/sqrt(n)  ", lambda n: 1 / math.sqrt(n), np.tanh),
    ("ReLU,  weight std 1/sqrt(n)", lambda n: 1 / math.sqrt(n), relu),
    ("ReLU,  He sqrt(2/n)      ", lambda n: math.sqrt(2 / n), relu),
]
for name, std_fn, act in settings:
    s = layer_stds(std_fn, act)
    print(f"{name}: layer 1 {s[0]:.4f}   layer 5 {s[4]:.4g}   layer 10 {s[9]:.4g}")

Output:

tanh,  weight std 0.01   : layer 1 0.0987   layer 5 9.801e-06   layer 10 1.045e-10
tanh,  weight std 1      : layer 1 0.9587   layer 5 0.9584   layer 10 0.9565
tanh,  Xavier 1/sqrt(n)  : layer 1 0.6255   layer 5 0.3185   layer 10 0.2319
ReLU,  weight std 1/sqrt(n): layer 1 0.5830   layer 5 0.1767   layer 10 0.03273
ReLU,  He sqrt(2/n)      : layer 1 0.8245   layer 5 0.9998   layer 10 1.047

Reading the results:

  • Tanh with tiny weights (0.01): the activations fall from 0.1 to 10−1010^{-10} by layer 10. The signal has vanished.
  • Tanh with large weights (1): the standard deviation sits at 0.96, close to the maximum possible value of 1 for tanh. Almost every neuron is saturated at ±1\pm 1, where the slope is nearly zero, so gradients cannot flow.
  • Tanh with Xavier: the signal shrinks gently (0.63 to 0.23) but stays useful. The shrinkage is because tanh itself compresses values.
  • ReLU with the Xavier scale: the signal decays (0.58 to 0.033), as the derivation predicted, since half the units are off.
  • ReLU with He: the standard deviation stays near 1 through all ten layers (0.82, 1.00, 1.05).

The backward pass has the same problem: gradients are multiplied by weight matrices as they travel back, which is why a good initialization helps the gradients as much as the activations.

Practical rules

  • Biases are usually initialised to zero. Symmetry is broken by the weights.
  • Use Xavier/Glorot for tanh or sigmoid layers and He for ReLU-type layers. Deep learning libraries apply sensible defaults, but it helps to know what they are doing.
  • If a deep network's loss does not move at all at the start, check the initial scale.
Try it yourself
Activations and initialization →

Switch to the deep network view and change the weight scale.

EasyInitialization

Why does initialising all weights to the same value make a layer useless?

MediumInitialization

Why does He initialization use 2/n where Xavier uses 1/n?

HardInitializationSaturation

With tanh and weights of standard deviation 1, the activations stayed large. Why is that still a problem?