The loss is the network’s measurable objective.
A label says right or wrong, but training needs a smooth signal that says how wrong. Binary cross-entropy strongly penalizes confident mistakes and rewards confident correct predictions.
Rank these predictions
The true class is 1. Which should receive the largest loss: p = 0.9, p = 0.6, or p = 0.01?
p = 0.01. It is not merely wrong; it is extremely confident in the wrong class.Binary cross-entropy for one example
When y = 1
As the predicted class-1 probability approaches 1, loss approaches 0.
When y = 0
As the predicted class-1 probability approaches 0, loss approaches 0.
The curve rises sharply near a confident mistake. Correct confidence moves the loss toward zero.
Apply it to our forward pass
Our row has y = 1 and the network predicted ŷ = 0.5794:
| True y | Predicted ŷ | Loss | Interpretation |
|---|---|---|---|
| 1 | 0.90 | 0.105 | Confident and correct |
| 1 | 0.60 | 0.511 | Correct side, uncertain |
| 1 | 0.10 | 2.303 | Confident and wrong |
| 0 | 0.90 | 2.303 | Confident and wrong |
For m examples, training minimizes the mean:
The useful cancellation: sigmoid plus BCE
Let a = σ(z). First differentiate the loss with respect to the probability:
The sigmoid derivative is:
Apply the chain rule and simplify:
Meaning: at the output score, the error signal is simply predicted probability − true label. A correct prediction close to its label has a small gradient; a wrong confident prediction has a large one.
Continue from the score to an output weight
The output score is a weighted sum of the hidden activations. For output weight wⱼ[2]:
Only one term contains wⱼ[2], so its local derivative is the input flowing through that connection:
Now connect this local derivative to the loss using the chain rule:
Intuition: a weight’s gradient is the output error multiplied by the signal that used that weight. A connection carrying a large activation has more influence; a zero activation gives that connection zero gradient for this example.
Apply the result to both output weights
For our running example:
The first output-weight gradient is:
The second output-weight gradient is:
Stacking both derivatives gives the matrix form:
The bias gradient follows the same chain
Because the bias enters the output score with coefficient 1:
For a batch of m examples, average the example-level gradients:
Where the derivation stops here: these are the output-layer gradients. To reach W[1], the error must additionally pass through W[2] and the hidden activation derivative. That full journey is derived on the backpropagation mathematics page.
Numerical stability is part of correctness
log(0) is undefined. A simple teaching implementation clips probabilities to a small interval such as [10⁻¹², 1 − 10⁻¹²]. Production libraries usually combine sigmoid and BCE into a stable “loss from logits” calculation that avoids explicitly forming extreme probabilities.