Backpropagation asks who influenced the mistake.
The final loss was produced by a chain of earlier calculations. Backpropagation walks that chain in reverse and measures the loss’s sensitivity to every trainable parameter.
Reason before calculating
If a parameter has a positive gradient, should increasing it make the loss rise or fall?
A derivative is local sensitivity
If the derivative is large and positive, a small increase in the parameter increases the loss noticeably. If it is negative, increasing the parameter locally decreases the loss. If it is near zero, the loss is locally insensitive to that parameter.
Do not read a gradient as blame in a moral sense. It is a precise counterfactual: “If this number moved slightly, how would the loss move?”
The chain rule joins local effects
Suppose one parameter w changes a score z, which changes a prediction a, which changes the loss:
w
z
a
ŷ
J
Each factor answers one local question. Multiplication combines them into the total effect of w on the loss.
A tiny scalar example
Let z = wx, a = z², and J = 3a. With x = 2 and w = 4:
A tiny increase of 0.001 in w should increase the loss by approximately 96 × 0.001 = 0.096. This approximation becomes more accurate as the change becomes smaller.
Why the journey runs backward
We know the derivative of the loss at the output first. Earlier layers depend on that downstream signal, so gradients naturally arrive in reverse order:
Backpropagation computes gradients. Gradient descent uses those gradients to update parameters. They are related, but they are not the same operation.