Derive the starting point of backpropagation for softmax with cross-entropy: the gradient is simply the prediction minus the target.
Deep Learning- Fundamentals to Advanced Concepts
Backpropagation needs a place to start, the gradient of the loss with respect to the last pre-activation aL. For classification we use softmax with cross-entropy, and the result is unusually clean.
Setup
Let a=aL be the vector of K numbers before the output function. Then
y^j=∑keakeaj,L=−logy^ℓ
where ℓ is the index of the true class. We want ∂aj∂L for every j. The chain passes through y^, so we proceed in two steps.
Step 1: the loss with respect to the prediction
The loss depends on y^ only through the entry y^ℓ, so
∂y^i∂L=⎩⎨⎧−y^ℓ10i=ℓi=ℓ
Step 2: the softmax derivative
Each output y^i depends on all the entries of a, because every ak appears in the denominator. Differentiating gives
∂aj∂y^i=y^i(1[i=j]−y^j)
where 1[i=j] is 1 if i=j and 0 otherwise. In words:
for i=j: y^i(1−y^i),
for i=j: −y^iy^j.
A quick derivation of the first case. Write S=∑keak, so y^i=eai/S. Then ∂y^i/∂ai=eai/S−eaieai/S2=y^i−y^i2. For i=j the numerator does not depend on aj, so only the denominator contributes, giving −eaieaj/S2=−y^iy^j.
Combine with the chain rule
The loss reaches aj through every y^i:
∂aj∂L=i∑∂y^i∂L∂aj∂y^i
Only the term i=ℓ survives, because every other ∂L/∂y^i is 0:
∂aj∂L=−y^ℓ1⋅y^ℓ(1[ℓ=j]−y^j)=y^j−1[ℓ=j]
The y^ℓ factors cancel, leaving a very clean result:
∇aLL=y^−eℓ
where eℓ is the one-hot vector for the true class (1 in position ℓ, 0 elsewhere). The gradient is simply the prediction minus the target.
What it means
For a wrong classj, the gradient is y^j>0. Gradient descent subtracts it, pushing that class's score down, more strongly the more probability the network wrongly gave it.
For the true class, the gradient is y^ℓ−1<0. Subtracting it pushes the score up.
The entries always sum to 0 (the predictions sum to 1 and so does the one-hot vector), so the update moves probability between classes and does not create it.
Example. For the scores a=(−1,1,2,3) we found y^=(0.012,0.089,0.242,0.657). With the true class being the third, the gradient is
The fourth class has the highest predicted probability but is wrong, so it receives the strongest push down. The true class is pushed up.
The worked network from earlier. There y^=(0.376,0.624) and the true class is 2, so ∇a2L=(0.376,0.624−1)=(0.376,−0.376).
Regression gives the same shape
For regression with a linear output, y^=a, and squared error L=21∑j(y^j−yj)2, we get
∂aj∂L=y^j−yj
It is again "prediction minus target". That is part of why softmax with cross-entropy and linear output with squared error are such natural pairings: the gradient at the output takes the same simple form in both.
EasySoftmaxGradient
The network outputs (0.7, 0.2, 0.1) and the true class is the first. What is the gradient with respect to a_L?
MediumSoftmaxGradient
Why do the entries of the softmax-cross-entropy gradient always sum to zero?
HardSoftmaxChain rule
Why do the y-hat-l factors cancel when combining Step 1 and Step 2?