Concept 3

Log loss and training logistic regression

Logistic regression learns coefficients by minimizing a probability-aware loss function called log loss or binary cross entropy.

Why not use MSE?

MSE can be used with probabilities, but it is not the standard loss for logistic regression. Log loss matches the probability model and strongly punishes confident wrong predictions.

PredictionTrue labelIntuition
\(\hat{p}=0.99\)\(y=1\)Very confident and correct, tiny loss.
\(\hat{p}=0.51\)\(y=1\)Correct side, but not confident.
\(\hat{p}=0.01\)\(y=1\)Very confident and wrong, huge loss.

Confidence check

Which should be punished more: predicting 0.49 for a positive example or predicting 0.01?

Predicting 0.01 should be punished much more because the model is confidently wrong.

Binary cross entropy loss

For one training example:

\[ L(y,\hat{p})=-\left[y\log(\hat{p})+(1-y)\log(1-\hat{p})\right] \]

This compact formula contains two cases:

If \(y=1\)

\[ L=-\log(\hat{p}) \]

Loss is small when \(\hat{p}\) is close to 1.

If \(y=0\)

\[ L=-\log(1-\hat{p}) \]

Loss is small when \(\hat{p}\) is close to 0.

Numerical safety

In code, probabilities are often clipped slightly away from 0 and 1 before taking logs.

Loss curve intuition

predicted probability p loss loss when y=1 loss when y=0

For \(y=1\), loss decreases as predicted probability approaches 1. For \(y=0\), loss decreases as probability approaches 0.

Cost function over all examples

The training cost is average binary cross entropy across all \(n\) examples:

\[ J(\beta)= -\frac{1}{n}\sum_{i=1}^{n} \left[ y_i\log(\hat{p}_i)+(1-y_i)\log(1-\hat{p}_i) \right] \]

Here:

\[ \hat{p}_i=\sigma(x_i^T\beta) \]
Logistic regression has no closed-form OLS-style solution. We usually optimize this cost function iteratively.

Negative log likelihood, step by step

Likelihood means: “How likely is the observed answer under the probability predicted by the model?”

For one row, if the true class is \(1\), the model assigns likelihood \(\hat{p}\). If the true class is \(0\), the model assigns likelihood \(1-\hat{p}\).

\[ P(y_i\mid x_i)=\hat{p}_i^{y_i}(1-\hat{p}_i)^{1-y_i} \]
True labelPredicted probability \(\hat{p}\)Likelihood of true labelNLL
\(y=1\)0.900.90\(-\log(0.90)=0.105\)
\(y=1\)0.200.20\(-\log(0.20)=1.609\)
\(y=0\)0.20\(1-0.20=0.80\)\(-\log(0.80)=0.223\)
\(y=0\)0.95\(1-0.95=0.05\)\(-\log(0.05)=2.996\)

For many rows, likelihood multiplies all row-wise probabilities:

\[ \mathcal{L}(\beta)=\prod_{i=1}^{n}\hat{p}_i^{y_i}(1-\hat{p}_i)^{1-y_i} \]

Products of many probabilities become tiny, so we take log. Log turns products into sums, which are easier to optimize:

\[ \log\mathcal{L}(\beta)=\sum_{i=1}^{n}\left[y_i\log(\hat{p}_i)+(1-y_i)\log(1-\hat{p}_i)\right] \]

Maximum likelihood wants to maximize this value. Most ML training code minimizes a loss, so we multiply by \(-1\):

\[ \text{NLL}(\beta)=-\log\mathcal{L}(\beta) \]
Negative log likelihood is simply log loss written from the maximum-likelihood point of view.

Initial NLL for an untrained model

If all coefficients start at zero, then every linear score is zero:

\[ z_i=x_i^T\beta=0 \]

The sigmoid gives the same probability for every row:

\[ \hat{p}_i=\sigma(0)=\frac{1}{2}=0.5 \]

For either class, the likelihood of the true label is \(0.5\), so the NLL per example is:

\[ -\log(0.5)=\log(2)\approx0.693 \]

For \(n\) examples:

\[ \text{Total NLL}=n\log(2) \qquad \text{Average NLL}=\log(2)\approx0.693 \]
Dataset sizeUntrained probabilityTotal NLLAverage NLL
10 rows0.5 for every row\(10\log(2)=6.93\)0.693
100 rows0.5 for every row\(100\log(2)=69.3\)0.693
If we initialize or fit only the intercept to match the class rate, the initial average NLL becomes the entropy of the class distribution. But with all logits initialized at zero, the clean starting value is \(\log(2)\).

Initial NLL for multiclass classification

For \(K\) classes, softmax converts class scores into class probabilities:

\[ P(y=c\mid x)=\frac{e^{z_c}}{\sum_{k=1}^{K}e^{z_k}} \]

If the untrained model starts with all class scores equal to zero, then each class gets equal probability:

\[ P(y=c\mid x)=\frac{e^0}{e^0+e^0+\cdots+e^0}=\frac{1}{K} \]

The NLL for the correct class is:

\[ -\log\left(\frac{1}{K}\right)=\log(K) \]
Number of classesUntrained probability per classInitial average NLL
2\(1/2=0.5\)\(\log(2)=0.693\)
3\(1/3=0.333\)\(\log(3)=1.099\)
10\(1/10=0.1\)\(\log(10)=2.303\)

This gives students a useful benchmark: a trained model should usually reduce average NLL below the “random equal probability” starting point.

Gradient descent

Using matrix notation:

\[ \hat{p}=\sigma(X\beta) \]

The gradient of average log loss is:

\[ \nabla_\beta J=\frac{1}{n}X^T(\hat{p}-y) \]

The update rule is:

\[ \beta := \beta-\alpha\nabla_\beta J \]

The term \(\hat{p}-y\) is the probability error. If the model predicts too high, the update pushes coefficients down; if too low, it pushes them up.

gradient descent iterations cost cost should decrease

During successful training, log loss should generally move downward over iterations.

Full gradient derivation

For one training example, start with:

\[ z_i=x_i^T\beta \qquad \hat{p}_i=\sigma(z_i) \]

The single-row loss is:

\[ L_i=-\left[y_i\log(\hat{p}_i)+(1-y_i)\log(1-\hat{p}_i)\right] \]

Differentiate the loss with respect to \(\hat{p}_i\):

\[ \frac{\partial L_i}{\partial \hat{p}_i} =-\frac{y_i}{\hat{p}_i}+\frac{1-y_i}{1-\hat{p}_i} \]

The sigmoid derivative is:

\[ \frac{\partial \hat{p}_i}{\partial z_i} =\hat{p}_i(1-\hat{p}_i) \]

Using the chain rule:

\[ \frac{\partial L_i}{\partial z_i} = \frac{\partial L_i}{\partial \hat{p}_i} \cdot \frac{\partial \hat{p}_i}{\partial z_i} \]
\[ \frac{\partial L_i}{\partial z_i} = \left(-\frac{y_i}{\hat{p}_i}+\frac{1-y_i}{1-\hat{p}_i}\right) \hat{p}_i(1-\hat{p}_i) \]
\[ \frac{\partial L_i}{\partial z_i} = -y_i(1-\hat{p}_i)+(1-y_i)\hat{p}_i = \hat{p}_i-y_i \]

Since \(z_i=x_i^T\beta\), for coefficient \(\beta_j\):

\[ \frac{\partial z_i}{\partial \beta_j}=x_{ij} \]

So the coefficient-level derivative is:

\[ \frac{\partial L_i}{\partial \beta_j} = (\hat{p}_i-y_i)x_{ij} \]

Averaging over all examples gives:

\[ \frac{\partial J}{\partial \beta_j} = \frac{1}{n}\sum_{i=1}^{n}(\hat{p}_i-y_i)x_{ij} \]

In matrix notation:

\[ \nabla_\beta J=\frac{1}{n}X^T(\hat{p}-y) \]

Gradient intuition

If \(y=1\) but the model predicts \(\hat{p}=0.2\), what is \(\hat{p}-y\)?

It is \(-0.8\), so the update pushes the model to increase the probability for similar examples.

Gradient descent steps

The training loop is:

Initialize \(\beta\) Compute \(z=X\beta\) Compute \(\hat{p}=\sigma(z)\) Compute gradient Update \(\beta\)
\[ \beta^{(t+1)} = \beta^{(t)} - \alpha\frac{1}{n}X^T(\hat{p}-y) \]
StepMeaning
Initialize \(\beta\)Start with zeros or small random values.
Forward passCompute scores and probabilities.
LossMeasure average NLL/log loss.
GradientFind how each coefficient should change.
UpdateMove coefficients in the direction that reduces loss.

Maximum likelihood intuition

Logistic regression can be understood as choosing coefficients that make the observed labels most likely.

\[ \text{Choose }\beta\text{ so the labels we actually observed get high probability.} \]

Suppose two models predict probabilities for four examples:

True labelsModel A probabilities for true labelsModel B probabilities for true labels
\(1,0,1,1\)0.90, 0.80, 0.70, 0.850.55, 0.52, 0.51, 0.50

Model A gives higher probability to what actually happened, so it has higher likelihood and lower NLL.

\[ \text{higher likelihood} \Longleftrightarrow \text{lower negative log likelihood} \]

This is why minimizing log loss is the same as maximizing the probability assigned to the observed training labels.

Previous: Sigmoid Next: Interpretation