Part 5 · Turn prediction quality into one number

The loss is the network’s measurable objective.

A label says right or wrong, but training needs a smooth signal that says how wrong. Binary cross-entropy strongly penalizes confident mistakes and rewards confident correct predictions.

Rank these predictions

The true class is 1. Which should receive the largest loss: p = 0.9, p = 0.6, or p = 0.01?

Binary cross-entropy for one example

\[\mathcal L(y,\widehat y)=-\left[y\log(\widehat y)+(1-y)\log(1-\widehat y)\right]\]

When y = 1

\[\mathcal L=-\log(\widehat y)\]

As the predicted class-1 probability approaches 1, loss approaches 0.

When y = 0

\[\mathcal L=-\log(1-\widehat y)\]

As the predicted class-1 probability approaches 0, loss approaches 0.

The curve rises sharply near a confident mistake. Correct confidence moves the loss toward zero.

Apply it to our forward pass

Our row has y = 1 and the network predicted ŷ = 0.5794:

\[\mathcal L=-\log(0.5794)\approx0.5458\]
True yPredicted ŷLossInterpretation
10.900.105Confident and correct
10.600.511Correct side, uncertain
10.102.303Confident and wrong
00.902.303Confident and wrong

For m examples, training minimizes the mean:

\[J=-\frac{1}{m}\sum_{i=1}^{m}\left[y_i\log(\widehat y_i)+(1-y_i)\log(1-\widehat y_i)\right]\]

The useful cancellation: sigmoid plus BCE

Let a = σ(z). First differentiate the loss with respect to the probability:

\[\frac{\partial \mathcal L}{\partial a}=-\frac{y}{a}+\frac{1-y}{1-a}\]

The sigmoid derivative is:

\[\frac{\partial a}{\partial z}=a(1-a)\]

Apply the chain rule and simplify:

\[\frac{\partial \mathcal L}{\partial z}=\left(-\frac{y}{a}+\frac{1-y}{1-a}\right)a(1-a)\]\[=-y(1-a)+(1-y)a=-y+ya+a-ya=\boxed{a-y}\]

Meaning: at the output score, the error signal is simply predicted probability − true label. A correct prediction close to its label has a small gradient; a wrong confident prediction has a large one.

Continue from the score to an output weight

The output score is a weighted sum of the hidden activations. For output weight wⱼ[2]:

\[z^{[2]}=\sum_j a_j^{[1]}w_j^{[2]}+b^{[2]}\]

Only one term contains wⱼ[2], so its local derivative is the input flowing through that connection:

\[\frac{\partial z^{[2]}}{\partial w_j^{[2]}}=a_j^{[1]}\]

Now connect this local derivative to the loss using the chain rule:

\[\frac{\partial \mathcal L}{\partial w_j^{[2]}}=\frac{\partial \mathcal L}{\partial z^{[2]}}\frac{\partial z^{[2]}}{\partial w_j^{[2]}}\]\[=(a^{[2]}-y)a_j^{[1]}\]

Intuition: a weight’s gradient is the output error multiplied by the signal that used that weight. A connection carrying a large activation has more influence; a zero activation gives that connection zero gradient for this example.

Apply the result to both output weights

For our running example:

\[A^{[1]}=\begin{bmatrix}0.4219&-0.0500\end{bmatrix},\qquad a^{[2]}-y=0.5794-1=-0.4206\]

The first output-weight gradient is:

\[\frac{\partial \mathcal L}{\partial w_1^{[2]}}=(-0.4206)(0.4219)\approx-0.1775\]

The second output-weight gradient is:

\[\frac{\partial \mathcal L}{\partial w_2^{[2]}}=(-0.4206)(-0.0500)\approx0.0210\]

Stacking both derivatives gives the matrix form:

\[\boxed{\frac{\partial \mathcal L}{\partial W^{[2]}}=(A^{[1]})^T(A^{[2]}-y)}\]\[=\begin{bmatrix}0.4219\\-0.0500\end{bmatrix}(-0.4206)\approx\begin{bmatrix}-0.1775\\0.0210\end{bmatrix}\]
Hidden activationA[1]: 1 × 2
TransposeA[1]ᵀ: 2 × 1
Output errorA[2] − y: 1 × 1
Weight gradientdW[2]: 2 × 1

The bias gradient follows the same chain

Because the bias enters the output score with coefficient 1:

\[\frac{\partial z^{[2]}}{\partial b^{[2]}}=1\]\[\boxed{\frac{\partial \mathcal L}{\partial b^{[2]}}=a^{[2]}-y=-0.4206}\]

For a batch of m examples, average the example-level gradients:

\[\boxed{\frac{\partial J}{\partial W^{[2]}}=\frac{1}{m}(A^{[1]})^T(A^{[2]}-Y)}\]\[\boxed{\frac{\partial J}{\partial b^{[2]}}=\frac{1}{m}\sum_{i=1}^{m}\left(A_i^{[2]}-Y_i\right)}\]

Where the derivation stops here: these are the output-layer gradients. To reach W[1], the error must additionally pass through W[2] and the hidden activation derivative. That full journey is derived on the backpropagation mathematics page.

Numerical stability is part of correctness

log(0) is undefined. A simple teaching implementation clips probabilities to a small interval such as [10⁻¹², 1 − 10⁻¹²]. Production libraries usually combine sigmoid and BCE into a stable “loss from logits” calculation that avoids explicitly forming extreme probabilities.

Previous: Forward passNext: Backprop intuition