Interpreting coefficients and controlling complexity
Logistic regression coefficients are not direct probability changes. They operate on log-odds, which can be translated into odds ratios.
Odds and log-odds
If \(p\) is the probability of class 1, odds are:
Examples:
| Probability \(p\) | Odds \(p/(1-p)\) | Plain meaning |
|---|---|---|
| 0.50 | 1 | Class 1 and class 0 are equally likely. |
| 0.75 | 3 | Class 1 is 3 times as likely as class 0. |
| 0.20 | 0.25 | Class 1 is less likely than class 0. |
Logistic regression models the log of the odds:
Why does logistic regression model log-odds?
The model starts with a linear score:
Then sigmoid converts that score into probability:
Now derive the odds from the sigmoid:
Taking log on both sides gives:
Since \(z\) is the linear score:
Probability, odds, and log-odds examples
Probabilities are bounded between 0 and 1. Odds range from 0 to infinity. Log-odds range from \(-\infty\) to \(+\infty\), which makes them suitable for a linear score.
| Probability \(p\) | Odds \(p/(1-p)\) | Log-odds \(\log(p/(1-p))\) | Meaning |
|---|---|---|---|
| 0.10 | 0.111 | -2.197 | Strongly leans toward class 0. |
| 0.25 | 0.333 | -1.099 | Class 0 is more likely. |
| 0.50 | 1.000 | 0.000 | Boundary point. |
| 0.75 | 3.000 | 1.099 | Class 1 is more likely. |
| 0.90 | 9.000 | 2.197 | Strongly leans toward class 1. |
Scale check
Why not model probability directly as a linear function?
A line can go below 0 or above 1. Log-odds can safely take any real value, so it matches the linear score.
Coefficient interpretation
A one-unit increase in \(x_j\), holding other features fixed, changes log-odds by \(\beta_j\).
| Coefficient | Odds multiplier | Interpretation |
|---|---|---|
| \(\beta_j=0\) | \(e^0=1\) | No change in odds. |
| \(\beta_j=0.405\) | \(e^{0.405}\approx1.5\) | Odds increase by about 50%. |
| \(\beta_j=-0.693\) | \(e^{-0.693}\approx0.5\) | Odds are cut roughly in half. |
Odds interpretation
If \(e^{\beta_j}=2\), what happens to the odds when \(x_j\) increases by one unit?
The odds of class 1 double, holding other features fixed.
Why coefficients are not direct probability changes
The same coefficient can cause different probability changes depending on where the current probability starts.
Suppose a feature increases log-odds by \(0.693\). That means the odds multiply by \(e^{0.693}\approx2\).
| Starting probability | Starting odds | After odds double | New probability | Probability change |
|---|---|---|---|---|
| 0.10 | 0.111 | 0.222 | 0.182 | +0.082 |
| 0.50 | 1.000 | 2.000 | 0.667 | +0.167 |
| 0.80 | 4.000 | 8.000 | 0.889 | +0.089 |
The odds multiplier is fixed, but the probability change is not fixed.
Feature scaling
Scaling is not required for the sigmoid formula itself, but it matters in practice when using gradient descent or regularization.
If one feature ranges from 0 to 1 and another ranges from 0 to 100000, coefficient updates and penalties can become uneven.
Scaling makes optimization and regularization behave more evenly across features.
Regularization
Logistic regression can overfit when features are many, noisy, or strongly correlated. Regularization adds a penalty to discourage overly large coefficients.
| Method | Penalty | Effect |
|---|---|---|
| L2 / Ridge | \(\lambda\sum_{j=1}^{p}\beta_j^2\) | Shrinks coefficients smoothly toward zero. |
| L1 / Lasso | \(\lambda\sum_{j=1}^{p}|\beta_j|\) | Can set some coefficients exactly to zero. |
| Elastic Net | Mix of L1 and L2 | Combines shrinkage and feature selection behavior. |
Regularization controls overfitting by limiting how large coefficients can become.
Linearly separable data
Sometimes the classes can be perfectly separated by a line. This sounds ideal, but it creates a subtle training problem for logistic regression.
If a boundary already separates all points correctly, making the coefficients larger can make predicted probabilities even more confident:
As \(\beta\) grows, correctly classified positive examples can move toward \(\hat{p}\approx1\), and correctly classified negative examples can move toward \(\hat{p}\approx0\).
That can keep improving likelihood without meaningfully changing the boundary location.
When data is perfectly separable, larger coefficients can mean more confidence, not necessarily a better real-world model.
Regularization reason
If the data is already perfectly separated, why might coefficient size still keep growing?
Larger coefficients make probabilities closer to 0 or 1, increasing likelihood, even if the boundary already classifies all training points correctly.
Multiclass extension
For more than two classes, logistic regression can be extended in two common ways.
| Approach | Idea | Example |
|---|---|---|
| One-vs-rest | Train one binary classifier per class. | Cat vs rest, dog vs rest, horse vs rest. |
| Softmax regression | Predict probabilities across all classes that sum to 1. | \(P(cat)+P(dog)+P(horse)=1\) |