A network ends with a final output function, and training uses a loss that compares the output with the true answer. The right choice for the two depends on the kind of problem. Two kinds cover most cases: regression and classification.
Throughout, let denote the vector of numbers the network has computed just before the output function (its final weighted sum). The output function turns into the prediction .
Regression: predicting real numbers
Suppose the input describes a film and we want to predict three ratings at once: the IMDb rating, the critics' rating and the Rotten Tomatoes score. Now , a vector with three real numbers.
Loss. The natural choice is the mean squared error. For each example, square the difference for each of the three outputs, add them, and average over the examples:
Output function. Could we use the logistic (sigmoid) function here? No. A sigmoid squeezes every output into the range 0 to 1, but our ratings run from 0 to 10, or 0 to 100. No setting of the weights could ever reach a rating of 8.
So for regression the output function is linear: multiply by a weight matrix and add a bias, with no squashing:
The output is then unbounded. That seems dangerous, since the network might predict 10,000 when ratings only go to 10. The loss prevents it. If the true rating is 9.5 and the network outputs 1000, the term is enormous. Gradient descent sees the large loss and pushes the weights away from that configuration. The only way to keep the loss small is to stay in a sensible range.
A sigmoid output could not do this, however hard the loss pushed, because the output simply cannot leave .
Classification: predicting a class
Now suppose the input is an image and the output is one of four classes: apple, banana, mango, orange. The true answer is a probability distribution. If the picture is an apple, we know with certainty that it is an apple:
Our network should output a distribution too, such as , and we need a loss that measures the gap between the two. Two problems follow.
- How do we make sure the output is a probability distribution, with every value between 0 and 1 and the values summing to 1?
- What loss should compare two distributions?
The output function: softmax
The softmax takes the vector and produces
Take . The exponentials are , which sum to . So
All four values are positive and they sum to 1, as a distribution must. Two facts make this always work:
- is positive for any , even a negative one, so no output is negative.
- Dividing each by the sum of all of them makes the total exactly 1, so no value can exceed 1.
Why the exponent? Try dividing each raw value by the sum of the raw values instead. Since some of can be negative, you would get negative "probabilities". The exponential removes that problem. A larger also gives a larger share, so the ordering of the classes is preserved, and the largest value takes a disproportionately large share.
The loss: cross-entropy
For a true distribution and a predicted one over classes, the cross-entropy loss is
Our true has a 1 for the correct class and 0 for all others. So every term in the sum vanishes except one:
This is the negative log of the probability the network gave to the correct class. It is also called the negative log-likelihood.
With our example output :
- If the true class is the third one, the loss is .
- If the true class is the fourth, the loss is .
The better the network does on the true class, the smaller the loss. As the loss goes to 0, and as it grows without bound. Minimising the negative log-likelihood is the same as maximising the log-likelihood of the correct class.
The loss depends on the network's parameters because is computed from the input through every layer and finally the softmax. That is what lets us differentiate it and use gradient descent.
With just two classes you can use a single sigmoid output , the probability of class 1, with the loss . This is the two-class version of the same cross-entropy.
Summary
| Problem | Output function | Loss function |
|---|---|---|
| Regression (real-valued outputs) | linear | squared error |
| Classification (probabilities) | softmax | cross-entropy |
The rest of the course will mostly use softmax with cross-entropy, and everything we learn for it carries over to regression with a small change.
Explore how cross-entropy responds as the predicted probability of the correct class changes.