Output Functions and Loss Functions

Matching the output function and the loss to the problem: linear output with squared error for regression, softmax with cross-entropy for classification.

Deep Learning- Fundamentals to Advanced Concepts

A network ends with a final output function, and training uses a loss that compares the output with the true answer. The right choice for the two depends on the kind of problem. Two kinds cover most cases: regression and classification.

Throughout, let aa denote the vector of numbers the network has computed just before the output function (its final weighted sum). The output function turns aa into the prediction y^\hat y.

Regression: predicting real numbers

Suppose the input describes a film and we want to predict three ratings at once: the IMDb rating, the critics' rating and the Rotten Tomatoes score. Now y∈R3y \in \mathbb{R}^3, a vector with three real numbers.

Loss. The natural choice is the mean squared error. For each example, square the difference for each of the three outputs, add them, and average over the NN examples:

L=1N∑i=1N∑j=13(y^ij−yij)2L = \frac{1}{N} \sum_{i=1}^{N} \sum_{j=1}^{3} \bigl(\hat y_{ij} - y_{ij}\bigr)^2

Output function. Could we use the logistic (sigmoid) function here? No. A sigmoid squeezes every output into the range 0 to 1, but our ratings run from 0 to 10, or 0 to 100. No setting of the weights could ever reach a rating of 8.

So for regression the output function is linear: multiply by a weight matrix and add a bias, with no squashing:

y^=Wo a+bo\hat y = W_o\, a + b_o

The output is then unbounded. That seems dangerous, since the network might predict 10,000 when ratings only go to 10. The loss prevents it. If the true rating is 9.5 and the network outputs 1000, the term (9.5−1000)2(9.5 - 1000)^2 is enormous. Gradient descent sees the large loss and pushes the weights away from that configuration. The only way to keep the loss small is to stay in a sensible range.

A sigmoid output could not do this, however hard the loss pushed, because the output simply cannot leave (0,1)(0, 1).

Classification: predicting a class

Now suppose the input is an image and the output is one of four classes: apple, banana, mango, orange. The true answer is a probability distribution. If the picture is an apple, we know with certainty that it is an apple:

y=(1,0,0,0)y = (1, 0, 0, 0)

Our network should output a distribution too, such as (0.2,0.3,0.4,0.1)(0.2, 0.3, 0.4, 0.1), and we need a loss that measures the gap between the two. Two problems follow.

  1. How do we make sure the output is a probability distribution, with every value between 0 and 1 and the values summing to 1?
  2. What loss should compare two distributions?

The output function: softmax

The softmax takes the vector aa and produces

y^j=eaj∑keak\hat y_j = \frac{e^{a_j}}{\sum_{k} e^{a_k}}

Take a=(−1,  1,  2,  3)a = (-1,\; 1,\; 2,\; 3). The exponentials are 0.37,  2.72,  7.39,  20.090.37,\; 2.72,\; 7.39,\; 20.09, which sum to 30.5630.56. So

y^≈(0.012,  0.089,  0.242,  0.657)\hat y \approx (0.012,\; 0.089,\; 0.242,\; 0.657)

All four values are positive and they sum to 1, as a distribution must. Two facts make this always work:

  • eaje^{a_j} is positive for any aja_j, even a negative one, so no output is negative.
  • Dividing each by the sum of all of them makes the total exactly 1, so no value can exceed 1.

Why the exponent? Try dividing each raw value by the sum of the raw values instead. Since some of aa can be negative, you would get negative "probabilities". The exponential removes that problem. A larger aja_j also gives a larger share, so the ordering of the classes is preserved, and the largest value takes a disproportionately large share.

The loss: cross-entropy

For a true distribution yy and a predicted one y^\hat y over KK classes, the cross-entropy loss is

L=−∑c=1Kyclog⁡y^cL = -\sum_{c=1}^{K} y_c \log \hat y_c

Our true yy has a 1 for the correct class ℓ\ell and 0 for all others. So every term in the sum vanishes except one:

L=−log⁡y^ℓL = -\log \hat y_{\ell}

This is the negative log of the probability the network gave to the correct class. It is also called the negative log-likelihood.

With our example output (0.012,0.089,0.242,0.657)(0.012, 0.089, 0.242, 0.657):

  • If the true class is the third one, the loss is −log⁡0.242≈1.42-\log 0.242 \approx 1.42.
  • If the true class is the fourth, the loss is −log⁡0.657≈0.42-\log 0.657 \approx 0.42.

The better the network does on the true class, the smaller the loss. As y^ℓ→1\hat y_\ell \to 1 the loss goes to 0, and as y^ℓ→0\hat y_\ell \to 0 it grows without bound. Minimising the negative log-likelihood is the same as maximising the log-likelihood of the correct class.

The loss depends on the network's parameters because y^ℓ\hat y_\ell is computed from the input through every layer and finally the softmax. That is what lets us differentiate it and use gradient descent.

Two classes

With just two classes you can use a single sigmoid output y^\hat y, the probability of class 1, with the loss −[ylog⁡y^+(1−y)log⁡(1−y^)]-\bigl[y \log \hat y + (1 - y)\log(1 - \hat y)\bigr]. This is the two-class version of the same cross-entropy.

Summary

ProblemOutput functionLoss function
Regression (real-valued outputs)linearsquared error
Classification (probabilities)softmaxcross-entropy

The rest of the course will mostly use softmax with cross-entropy, and everything we learn for it carries over to regression with a small change.

Try it yourself
Entropy and loss simulator →

Explore how cross-entropy responds as the predicted probability of the correct class changes.

EasyRegression

Why is a sigmoid a poor output function for predicting film ratings from 0 to 10?

MediumSoftmax

Compute the softmax of the two values (0, 0) and of (2, 0).

MediumCross-entropy

The network gives probability 0.9 to the true class. What is the cross-entropy loss, and what if it gave 0.5?

MediumCross-entropy

Why does the cross-entropy sum collapse to a single term for a one-hot true label?