Part 6 · Credit assignment

Backpropagation asks who influenced the mistake.

The final loss was produced by a chain of earlier calculations. Backpropagation walks that chain in reverse and measures the loss’s sensitivity to every trainable parameter.

Reason before calculating

If a parameter has a positive gradient, should increasing it make the loss rise or fall?

A derivative is local sensitivity

\[\frac{\partial J}{\partial w}\approx\frac{\text{small change in loss}}{\text{small change in }w}\]

If the derivative is large and positive, a small increase in the parameter increases the loss noticeably. If it is negative, increasing the parameter locally decreases the loss. If it is near zero, the loss is locally insensitive to that parameter.

Do not read a gradient as blame in a moral sense. It is a precise counterfactual: “If this number moved slightly, how would the loss move?”

The chain rule joins local effects

Suppose one parameter w changes a score z, which changes a prediction a, which changes the loss:

Parameter
w
Weighted sum
z
Activation
a
Prediction
ŷ
Loss
J
\[\frac{\partial J}{\partial w}=\frac{\partial J}{\partial a}\cdot\frac{\partial a}{\partial z}\cdot\frac{\partial z}{\partial w}\]

Each factor answers one local question. Multiplication combines them into the total effect of w on the loss.

A tiny scalar example

Let z = wx, a = z², and J = 3a. With x = 2 and w = 4:

\[z=8,\qquad a=64,\qquad J=192\]\[\frac{\partial J}{\partial a}=3,\qquad \frac{\partial a}{\partial z}=2z=16,\qquad \frac{\partial z}{\partial w}=x=2\]\[\frac{\partial J}{\partial w}=3\times16\times2=96\]

A tiny increase of 0.001 in w should increase the loss by approximately 96 × 0.001 = 0.096. This approximation becomes more accurate as the change becomes smaller.

Why the journey runs backward

We know the derivative of the loss at the output first. Earlier layers depend on that downstream signal, so gradients naturally arrive in reverse order:

Output error: compare the current probability with the label.
Output parameters: determine how hidden activations contributed to that error.
Hidden activations: send the output error through the output weights.
Hidden parameters: include the activation derivative and original inputs.

Backpropagation computes gradients. Gradient descent uses those gradients to update parameters. They are related, but they are not the same operation.

Previous: LossNext: Backprop math