Log loss and training logistic regression
Logistic regression learns coefficients by minimizing a probability-aware loss function called log loss or binary cross entropy.
Why not use MSE?
MSE can be used with probabilities, but it is not the standard loss for logistic regression. Log loss matches the probability model and strongly punishes confident wrong predictions.
| Prediction | True label | Intuition |
|---|---|---|
| \(\hat{p}=0.99\) | \(y=1\) | Very confident and correct, tiny loss. |
| \(\hat{p}=0.51\) | \(y=1\) | Correct side, but not confident. |
| \(\hat{p}=0.01\) | \(y=1\) | Very confident and wrong, huge loss. |
Confidence check
Which should be punished more: predicting 0.49 for a positive example or predicting 0.01?
Predicting 0.01 should be punished much more because the model is confidently wrong.
Binary cross entropy loss
For one training example:
This compact formula contains two cases:
If \(y=1\)
Loss is small when \(\hat{p}\) is close to 1.
If \(y=0\)
Loss is small when \(\hat{p}\) is close to 0.
Numerical safety
In code, probabilities are often clipped slightly away from 0 and 1 before taking logs.
Loss curve intuition
For \(y=1\), loss decreases as predicted probability approaches 1. For \(y=0\), loss decreases as probability approaches 0.
Cost function over all examples
The training cost is average binary cross entropy across all \(n\) examples:
Here:
Negative log likelihood, step by step
Likelihood means: “How likely is the observed answer under the probability predicted by the model?”
For one row, if the true class is \(1\), the model assigns likelihood \(\hat{p}\). If the true class is \(0\), the model assigns likelihood \(1-\hat{p}\).
| True label | Predicted probability \(\hat{p}\) | Likelihood of true label | NLL |
|---|---|---|---|
| \(y=1\) | 0.90 | 0.90 | \(-\log(0.90)=0.105\) |
| \(y=1\) | 0.20 | 0.20 | \(-\log(0.20)=1.609\) |
| \(y=0\) | 0.20 | \(1-0.20=0.80\) | \(-\log(0.80)=0.223\) |
| \(y=0\) | 0.95 | \(1-0.95=0.05\) | \(-\log(0.05)=2.996\) |
For many rows, likelihood multiplies all row-wise probabilities:
Products of many probabilities become tiny, so we take log. Log turns products into sums, which are easier to optimize:
Maximum likelihood wants to maximize this value. Most ML training code minimizes a loss, so we multiply by \(-1\):
Initial NLL for an untrained model
If all coefficients start at zero, then every linear score is zero:
The sigmoid gives the same probability for every row:
For either class, the likelihood of the true label is \(0.5\), so the NLL per example is:
For \(n\) examples:
| Dataset size | Untrained probability | Total NLL | Average NLL |
|---|---|---|---|
| 10 rows | 0.5 for every row | \(10\log(2)=6.93\) | 0.693 |
| 100 rows | 0.5 for every row | \(100\log(2)=69.3\) | 0.693 |
Initial NLL for multiclass classification
For \(K\) classes, softmax converts class scores into class probabilities:
If the untrained model starts with all class scores equal to zero, then each class gets equal probability:
The NLL for the correct class is:
| Number of classes | Untrained probability per class | Initial average NLL |
|---|---|---|
| 2 | \(1/2=0.5\) | \(\log(2)=0.693\) |
| 3 | \(1/3=0.333\) | \(\log(3)=1.099\) |
| 10 | \(1/10=0.1\) | \(\log(10)=2.303\) |
This gives students a useful benchmark: a trained model should usually reduce average NLL below the “random equal probability” starting point.
Gradient descent
Using matrix notation:
The gradient of average log loss is:
The update rule is:
The term \(\hat{p}-y\) is the probability error. If the model predicts too high, the update pushes coefficients down; if too low, it pushes them up.
During successful training, log loss should generally move downward over iterations.
Full gradient derivation
For one training example, start with:
The single-row loss is:
Differentiate the loss with respect to \(\hat{p}_i\):
The sigmoid derivative is:
Using the chain rule:
Since \(z_i=x_i^T\beta\), for coefficient \(\beta_j\):
So the coefficient-level derivative is:
Averaging over all examples gives:
In matrix notation:
Gradient intuition
If \(y=1\) but the model predicts \(\hat{p}=0.2\), what is \(\hat{p}-y\)?
It is \(-0.8\), so the update pushes the model to increase the probability for similar examples.
Gradient descent steps
The training loop is:
| Step | Meaning |
|---|---|
| Initialize \(\beta\) | Start with zeros or small random values. |
| Forward pass | Compute scores and probabilities. |
| Loss | Measure average NLL/log loss. |
| Gradient | Find how each coefficient should change. |
| Update | Move coefficients in the direction that reduces loss. |
Maximum likelihood intuition
Logistic regression can be understood as choosing coefficients that make the observed labels most likely.
Suppose two models predict probabilities for four examples:
| True labels | Model A probabilities for true labels | Model B probabilities for true labels |
|---|---|---|
| \(1,0,1,1\) | 0.90, 0.80, 0.70, 0.85 | 0.55, 0.52, 0.51, 0.50 |
Model A gives higher probability to what actually happened, so it has higher likelihood and lower NLL.
This is why minimizing log loss is the same as maximizing the probability assigned to the observed training labels.