Dropout

Randomly switch off hidden neurons during training. It prevents neurons from relying on each other and acts like averaging a huge number of networks.

Deep Learning- Fundamentals to Advanced Concepts

Bagging needs several whole networks, which is expensive. Dropout (Hinton and colleagues, 2012, developed by Srivastava et al.) gets much of the benefit of an ensemble from a single network.

The method

During training, at every update and for every hidden neuron independently, flip a biased coin:

  • with probability pp (the keep probability) the neuron works normally,
  • with probability 1−p1 - p its output is set to zero for this step.

A fresh random mask is drawn for every example or mini-batch. Typical keep probabilities are around 0.5 for hidden layers and higher, such as 0.8 or more, for the inputs. The figure shows a single step: the crossed neurons (and all connections through them) are simply absent.

A small network in which four hidden neurons are crossed out and their connections removed, leaving a thinner network
One training step with dropout. The next step uses a different random set of neurons.

Why it helps

No co-adaptation. Without dropout a neuron can rely on particular other neurons being present and correct their mistakes. With dropout any neuron may vanish at any step, so each neuron must learn to be useful on its own and in many different company. That tends to give more robust features.

An ensemble in disguise. Each mask defines a different thinned sub-network, and all of them share weights. A network with nn neurons has 2n2^n possible sub-networks, and training visits a random selection of them. At test time we use the full network, which approximates the average of all these sub-networks.

The scaling problem

If the full network is used at test time, each neuron's output is larger on average than it was during training (when a fraction 1−p1 - p of neurons were off and the next layer saw less input). The two must be matched. The usual fix is inverted dropout: scale the surviving outputs up during training by 1/p1/p, so their expected value is unchanged, and use the network untouched at test time.

h~=m⊙hp,mj∼Bernoulli(p)\tilde h = \frac{m \odot h}{p}, \qquad m_j \sim \text{Bernoulli}(p)

Check that the expectation is preserved: E[h~j]=p hjp=hj\mathbb{E}[\tilde h_j] = \dfrac{p\,h_j}{p} = h_j.

import numpy as np

rng = np.random.default_rng(1)
h = np.array([0.2, 0.9, 0.5, 0.7, 0.1])
keep = 0.8

def dropout(h, keep, rng):
    mask = rng.random(h.shape) < keep
    return mask * h / keep

print("one mask applied:", np.round(dropout(h, keep, rng), 3))
average = np.mean([dropout(h, keep, rng) for _ in range(200000)], axis=0)
print("original activations:          ", h)
print("average over 200000 dropouts:  ", np.round(average, 3))

Output:

one mask applied: [0.25  0.    0.625 0.    0.125]
original activations:           [0.2 0.9 0.5 0.7 0.1]
average over 200000 dropouts:   [0.2   0.901 0.5   0.699 0.1  ]

In the single draw, two neurons were dropped (0.9 and 0.7 became 0), and the survivors were scaled up by 1/0.8=1.251/0.8 = 1.25 (0.2 became 0.25). Averaged over many draws the activations match the originals, so nothing needs to be rescaled at test time.

In backpropagation

Dropout is just another element-wise multiplication in the forward pass, so the backward pass multiplies by the same mask:

∂L∂h=mp⊙∂L∂h~\frac{\partial L}{\partial h} = \frac{m}{p} \odot \frac{\partial L}{\partial \tilde h}

Dropped neurons receive no gradient in that step. In code for a network like the one in the backpropagation lesson, apply dropout to each hidden layer's output in the forward pass only during training, store the mask, and multiply the gradient by mask / keep when going backwards.

Practical points

  • Only during training. At test time all neurons are active, with no random masking.
  • Training takes longer, because each step's gradient is noisier, but the test error is often lower.
  • Dropout combines well with the other regularizers of this module. It is one of the most widely used ways to reduce overfitting in large fully connected networks.
  • The maxout units in a later lesson were designed to work well together with dropout.
Try it yourself
Network lab →

Train a network on a small noisy dataset with and without dropout.

MediumDropout

Why do we divide by the keep probability in inverted dropout?

MediumDropout

How does dropout prevent co-adaptation of neurons?

HardDropoutEnsembles

In what sense is dropout like training an ensemble?