Concept 4

Interpreting coefficients and controlling complexity

Logistic regression coefficients are not direct probability changes. They operate on log-odds, which can be translated into odds ratios.

Odds and log-odds

If \(p\) is the probability of class 1, odds are:

\[ \text{odds}=\frac{p}{1-p} \]

Examples:

Probability \(p\)Odds \(p/(1-p)\)Plain meaning
0.501Class 1 and class 0 are equally likely.
0.753Class 1 is 3 times as likely as class 0.
0.200.25Class 1 is less likely than class 0.

Logistic regression models the log of the odds:

\[ \log\left(\frac{p}{1-p}\right)=\beta_0+\beta_1x_1+\cdots+\beta_px_p \]

Why does logistic regression model log-odds?

The model starts with a linear score:

\[ z=\beta_0+\beta_1x_1+\cdots+\beta_px_p \]

Then sigmoid converts that score into probability:

\[ p=\frac{1}{1+e^{-z}} \]

Now derive the odds from the sigmoid:

\[ 1-p = 1-\frac{1}{1+e^{-z}} = \frac{e^{-z}}{1+e^{-z}} \]
\[ \frac{p}{1-p} = \frac{\frac{1}{1+e^{-z}}}{\frac{e^{-z}}{1+e^{-z}}} = e^z \]

Taking log on both sides gives:

\[ \log\left(\frac{p}{1-p}\right)=z \]

Since \(z\) is the linear score:

\[ \log\left(\frac{p}{1-p}\right) = \beta_0+\beta_1x_1+\cdots+\beta_px_p \]
This is the heart of logistic regression: probability is nonlinear, but log-odds is modeled as a linear function of features.

Probability, odds, and log-odds examples

Probabilities are bounded between 0 and 1. Odds range from 0 to infinity. Log-odds range from \(-\infty\) to \(+\infty\), which makes them suitable for a linear score.

Probability \(p\)Odds \(p/(1-p)\)Log-odds \(\log(p/(1-p))\)Meaning
0.100.111-2.197Strongly leans toward class 0.
0.250.333-1.099Class 0 is more likely.
0.501.0000.000Boundary point.
0.753.0001.099Class 1 is more likely.
0.909.0002.197Strongly leans toward class 1.

Scale check

Why not model probability directly as a linear function?

A line can go below 0 or above 1. Log-odds can safely take any real value, so it matches the linear score.

Coefficient interpretation

A one-unit increase in \(x_j\), holding other features fixed, changes log-odds by \(\beta_j\).

\[ \text{odds multiplier}=e^{\beta_j} \]
CoefficientOdds multiplierInterpretation
\(\beta_j=0\)\(e^0=1\)No change in odds.
\(\beta_j=0.405\)\(e^{0.405}\approx1.5\)Odds increase by about 50%.
\(\beta_j=-0.693\)\(e^{-0.693}\approx0.5\)Odds are cut roughly in half.
A coefficient is not a fixed probability change. Probability changes depend on the starting probability because the sigmoid curve is nonlinear.

Odds interpretation

If \(e^{\beta_j}=2\), what happens to the odds when \(x_j\) increases by one unit?

The odds of class 1 double, holding other features fixed.

Why coefficients are not direct probability changes

The same coefficient can cause different probability changes depending on where the current probability starts.

Suppose a feature increases log-odds by \(0.693\). That means the odds multiply by \(e^{0.693}\approx2\).

Starting probabilityStarting oddsAfter odds doubleNew probabilityProbability change
0.100.1110.2220.182+0.082
0.501.0002.0000.667+0.167
0.804.0008.0000.889+0.089

The odds multiplier is fixed, but the probability change is not fixed.

Feature scaling

Scaling is not required for the sigmoid formula itself, but it matters in practice when using gradient descent or regularization.

If one feature ranges from 0 to 1 and another ranges from 0 to 100000, coefficient updates and penalties can become uneven.

Standard practice: scale numerical features before regularized logistic regression.
age: 18-80 income: 0-200000 scaled features

Scaling makes optimization and regularization behave more evenly across features.

Regularization

Logistic regression can overfit when features are many, noisy, or strongly correlated. Regularization adds a penalty to discourage overly large coefficients.

\[ J_{regularized}=J_{logloss}+\lambda\cdot\text{penalty} \]
MethodPenaltyEffect
L2 / Ridge\(\lambda\sum_{j=1}^{p}\beta_j^2\)Shrinks coefficients smoothly toward zero.
L1 / Lasso\(\lambda\sum_{j=1}^{p}|\beta_j|\)Can set some coefficients exactly to zero.
Elastic NetMix of L1 and L2Combines shrinkage and feature selection behavior.
without penalty with penalty large coefficients shrunk coefficients

Regularization controls overfitting by limiting how large coefficients can become.

Linearly separable data

Sometimes the classes can be perfectly separated by a line. This sounds ideal, but it creates a subtle training problem for logistic regression.

If a boundary already separates all points correctly, making the coefficients larger can make predicted probabilities even more confident:

\[ z=x^T\beta \qquad \hat{p}=\sigma(z) \]

As \(\beta\) grows, correctly classified positive examples can move toward \(\hat{p}\approx1\), and correctly classified negative examples can move toward \(\hat{p}\approx0\).

That can keep improving likelihood without meaningfully changing the boundary location.

Without regularization, perfectly separable data can make logistic-regression coefficients grow extremely large. Regularization keeps the solution finite and stable.
many nearby boundaries separate perfectly

When data is perfectly separable, larger coefficients can mean more confidence, not necessarily a better real-world model.

Regularization reason

If the data is already perfectly separated, why might coefficient size still keep growing?

Larger coefficients make probabilities closer to 0 or 1, increasing likelihood, even if the boundary already classifies all training points correctly.

Multiclass extension

For more than two classes, logistic regression can be extended in two common ways.

ApproachIdeaExample
One-vs-restTrain one binary classifier per class.Cat vs rest, dog vs rest, horse vs rest.
Softmax regressionPredict probabilities across all classes that sum to 1.\(P(cat)+P(dog)+P(horse)=1\)
For a first logistic regression session, keep multiclass brief. The core binary ideas are the foundation.
Previous: Loss Next: Metrics