Part 7 · Derive every gradient

Now send one error through the entire network.

We reuse the exact forward-pass values. Every gradient has the same shape as the parameter it will update, which gives us a powerful correctness check.

Predict the first gradient

Our row has y = 1 and ŷ = 0.5794. Is the output-score gradient positive or negative?

Watch the complete forward and backward journey

Blue values are calculated and cached during the forward pass. Orange values are gradients calculated in reverse during backpropagation. Use the controls to pause at any equation.

Forward values and flowBackward gradients and reverse flow

Forward 1: supply the example and parameters

\[x=\begin{bmatrix}1.0&0.5\end{bmatrix},\qquad y=1\]\[W^{[1]}=\begin{bmatrix}0.4&-0.2\\0.1&0.3\end{bmatrix},\quad b^{[1]}=\begin{bmatrix}0&0\end{bmatrix}\]\[W^{[2]}=\begin{bmatrix}0.7\\-0.5\end{bmatrix},\qquad b^{[2]}=0\]

The six blue labels on the connections are the current weights. They remain fixed during this forward and backward pass; gradient descent updates them only after all gradients are known.

Values saved during the forward pass

This is the complete numerical state before backpropagation begins: the example, all trainable parameters, and every cached intermediate value.

\[x=\begin{bmatrix}1.0&0.5\end{bmatrix},\qquad y=1\]\[W^{[1]}=\begin{bmatrix}0.4&-0.2\\0.1&0.3\end{bmatrix},\qquad b^{[1]}=\begin{bmatrix}0&0\end{bmatrix}\]\[Z^{[1]}=\begin{bmatrix}0.45&-0.05\end{bmatrix},\qquad A^{[1]}\approx\begin{bmatrix}0.4219&-0.0500\end{bmatrix}\]\[W^{[2]}=\begin{bmatrix}0.7\\-0.5\end{bmatrix},\qquad b^{[2]}=0\]\[Z^{[2]}\approx0.3203,\qquad A^{[2]}=\widehat y\approx0.5794\]\[\mathcal L(y,\widehat y)\approx0.5458\]

What is cached? Implementations must retain the intermediate activations needed by the chain rule. The parameters are already stored by the model; showing them here makes every backward calculation independently traceable.

Step 1: output error signal

Sigmoid plus binary cross-entropy gives the simplified result:

\[dZ^{[2]}=\frac{\partial \mathcal L}{\partial Z^{[2]}}=A^{[2]}-y=0.5794-1=-0.4206\]

The negative sign means moving the output score upward would reduce this example’s loss.

Step 2: output weight and bias gradients

Because Z[2] = A[1]W[2] + b[2], each output weight’s influence is scaled by the hidden activation entering it:

\[dW^{[2]}=(A^{[1]})^T dZ^{[2]}\]\[=\begin{bmatrix}0.4219\\-0.0500\end{bmatrix}(-0.4206)\approx\begin{bmatrix}-0.1775\\0.0210\end{bmatrix}\]\[db^{[2]}=dZ^{[2]}=-0.4206\]
ParameterW[2]: 2 × 1
GradientdW[2]: 2 × 1
Parameterb[2]: 1 × 1
Gradientdb[2]: 1 × 1

Step 3: pass the signal into the hidden layer

First ask how changing each hidden activation would affect the loss:

\[dA^{[1]}=dZ^{[2]}(W^{[2]})^T\]\[=(-0.4206)\begin{bmatrix}0.7&-0.5\end{bmatrix}\approx\begin{bmatrix}-0.2944&0.2103\end{bmatrix}\]

The output weights route and scale the error. Their signs can even reverse its direction.

Step 4: pass through tanh

For tanh, the derivative can be computed directly from the saved activation:

\[\frac{d}{dz}\tanh(z)=1-\tanh^2(z)=1-a^2\]\[dZ^{[1]}=dA^{[1]}\odot\left(1-(A^{[1]})^2\right)\]\[\approx\begin{bmatrix}-0.2944&0.2103\end{bmatrix}\odot\begin{bmatrix}0.8220&0.9975\end{bmatrix}\approx\begin{bmatrix}-0.2420&0.2098\end{bmatrix}\]

The symbol ⊙ means element-by-element multiplication. A saturated tanh activation would have a derivative near zero and would weaken the returning signal.

Step 5: hidden weight and bias gradients

\[dW^{[1]}=x^TdZ^{[1]}\]\[=\begin{bmatrix}1.0\\0.5\end{bmatrix}\begin{bmatrix}-0.2420&0.2098\end{bmatrix}=\begin{bmatrix}-0.2420&0.2098\\-0.1210&0.1049\end{bmatrix}\]\[db^{[1]}=dZ^{[1]}=\begin{bmatrix}-0.2420&0.2098\end{bmatrix}\]
ParameterW[1]: 2 × 2
GradientdW[1]: 2 × 2
Parameterb[1]: 1 × 2
Gradientdb[1]: 1 × 2

Batch formulas for m examples

We place the factor 1/m inside dZ[2]. Therefore, do not divide by m again in the later lines.

\[dZ^{[2]}=\frac{A^{[2]}-Y}{m}\]\[dW^{[2]}=(A^{[1]})^TdZ^{[2]},\qquad db^{[2]}=\sum_{i=1}^{m}dZ_i^{[2]}\]\[dA^{[1]}=dZ^{[2]}(W^{[2]})^T\]\[dZ^{[1]}=dA^{[1]}\odot\left(1-(A^{[1]})^2\right)\]\[dW^{[1]}=X^TdZ^{[1]},\qquad db^{[1]}=\sum_{i=1}^{m}dZ_i^{[1]}\]

Bias gradients sum across examples because one shared bias was broadcast to every row during the forward pass.

Convention warning: another correct implementation may divide dW and db by m instead. Use exactly one averaging convention, not both.

Previous: Backprop intuitionNext: Updates