Now send one error through the entire network.
We reuse the exact forward-pass values. Every gradient has the same shape as the parameter it will update, which gives us a powerful correctness check.
Predict the first gradient
Our row has y = 1 and ŷ = 0.5794. Is the output-score gradient positive or negative?
ŷ − y = 0.5794 − 1 = −0.4206. Increasing the output score would increase the class-1 probability and reduce the loss.Watch the complete forward and backward journey
Blue values are calculated and cached during the forward pass. Orange values are gradients calculated in reverse during backpropagation. Use the controls to pause at any equation.
Forward 1: supply the example and parameters
The six blue labels on the connections are the current weights. They remain fixed during this forward and backward pass; gradient descent updates them only after all gradients are known.
Forward 2: hidden weighted sums
Each hidden neuron receives its own weighted combination of both inputs.
Forward 3: hidden activations
Tanh turns the two scores into the learned features sent to the output layer.
Forward 4: output score
The output neuron combines both hidden activations into one raw score.
Forward 5: predicted probability
The sigmoid converts the score into the predicted class-1 probability.
Forward 6: loss
Because y = 1, binary cross-entropy keeps only the first logarithmic term.
Backward 1: output-score gradient
The backward journey begins at the output. Orange now represents sensitivity of the loss.
Backward 2: output parameter gradients
The hidden activations scale how strongly each output connection contributed.
Backward 3: gradient reaching hidden activations
The output weights route the loss gradient back to the two hidden activations.
Backward 4: gradient through tanh
The local tanh derivative determines how much of each returning gradient passes through.
Backward 5: hidden parameter gradients
Every trainable parameter now has a gradient with the same shape. Gradient descent can perform the update.
Values saved during the forward pass
This is the complete numerical state before backpropagation begins: the example, all trainable parameters, and every cached intermediate value.
What is cached? Implementations must retain the intermediate activations needed by the chain rule. The parameters are already stored by the model; showing them here makes every backward calculation independently traceable.
Step 1: output error signal
Sigmoid plus binary cross-entropy gives the simplified result:
The negative sign means moving the output score upward would reduce this example’s loss.
Step 2: output weight and bias gradients
Because Z[2] = A[1]W[2] + b[2], each output weight’s influence is scaled by the hidden activation entering it:
Step 3: pass the signal into the hidden layer
First ask how changing each hidden activation would affect the loss:
The output weights route and scale the error. Their signs can even reverse its direction.
Step 4: pass through tanh
For tanh, the derivative can be computed directly from the saved activation:
The symbol ⊙ means element-by-element multiplication. A saturated tanh activation would have a derivative near zero and would weaken the returning signal.
Step 5: hidden weight and bias gradients
Batch formulas for m examples
We place the factor 1/m inside dZ[2]. Therefore, do not divide by m again in the later lines.
Bias gradients sum across examples because one shared bias was broadcast to every row during the forward pass.
Convention warning: another correct implementation may divide dW and db by m instead. Use exactly one averaging convention, not both.