Every activation we have seen is a fixed curve chosen in advance. Maxout (Goodfellow and colleagues, 2013) lets the network learn its activation function.
Definition
A maxout unit with pieces computes different linear functions of its input and outputs the largest:
Each piece has its own weights and bias, all learned by backpropagation. The unit is a normal neuron with weighted sums instead of one, followed by a max. The gradient flows through whichever piece is currently the largest.
It includes ReLU and much more
The maximum of straight lines is always a convex, piecewise-linear function, and the shape is determined by the pieces:
- With two pieces and (that is, and ) we get , which is ReLU.
- With the pieces and we get , the absolute value.
- With more pieces, it can approximate any convex function as closely as we like.
Here is the same idea in code:
import numpy as np
x = np.linspace(-2, 2, 401)
print("ReLU as maxout:", np.max([1 * x + 0, 0 * x + 0], axis=0)[[0, 100, 200, 300, 400]])
print("|x| as maxout: ", np.max([1 * x, -1 * x], axis=0)[[0, 100, 200, 300, 400]])
for pieces in (3, 5, 9):
points_ = np.linspace(-2, 2, pieces)
lines = [(2 * p, -p * p) for p in points_] # tangent line of x^2 at p
approx = np.max([a * x + b for a, b in lines], axis=0)
print(f"{pieces} pieces: largest error approximating x^2 on [-2, 2]: {np.max(np.abs(approx - x ** 2)):.4f}")
Output (values at for the first two lines):
ReLU as maxout: [0. 0. 0. 1. 2.]
|x| as maxout: [ 2. 1. -0. 1. 2.]
3 pieces: largest error approximating x^2 on [-2, 2]: 1.0000
5 pieces: largest error approximating x^2 on [-2, 2]: 0.2500
9 pieces: largest error approximating x^2 on [-2, 2]: 0.0625
The last three lines use tangent lines to the curve , which is convex. More pieces give a closer fit: halving the spacing between the tangent points divides the largest error by four (1, 0.25, 0.0625).
Advantages and costs
Advantages
- No saturation and no dead region. One of the pieces is always active, and each piece has a non-zero slope in general, so gradients flow. (A ReLU unit, in contrast, is flat for negative inputs.)
- Flexible. The network learns what shape of activation it needs, separately for each unit.
- Pairs well with dropout. Maxout was designed to be used with dropout, and the combination gave strong results when it was introduced.
Costs
- times as many parameters for each unit, because each piece has its own weights. A maxout layer with pieces has twice the weights of a ReLU layer with the same number of units.
- The functions it represents are convex in each unit, which is less general than it may sound, though a network of many units can still model non-convex functions.
Where it fits
Maxout is a useful conceptual bridge. It shows ReLU and leaky ReLU as special cases of a more general idea, and it shows one way to learn what the earlier lesson chose by hand. In practice, simple ReLU-type activations or GELU are used more often today because they are cheaper, but the idea of learnable piecewise-linear activations appears in several later designs.
Switch to the maxout view and add more linear pieces.