Classification from scratch

Logistic regression: from linear score to probability to decision.

Logistic regression is the bridge from regression thinking to classification. It starts with a linear score, converts it into a probability, learns using log loss, and makes decisions using a threshold.

Teaching story

Classification problem Linear score Sigmoid probability Log loss training Threshold and metrics

Opening question

If linear regression predicts 1.3 for “spam” and -0.2 for “not spam,” what should those numbers mean?

This motivates the need for probabilities between 0 and 1.

Session roadmap

Core formulas in one place

\[ z = \beta_0+\beta_1x_1+\cdots+\beta_px_p \qquad \hat{p}=\sigma(z)=\frac{1}{1+e^{-z}} \]
\[ L(y,\hat{p})=-\left[y\log(\hat{p})+(1-y)\log(1-\hat{p})\right] \]
\[ \nabla_\beta J=\frac{1}{n}X^T(\hat{p}-y) \qquad \beta := \beta-\alpha\nabla_\beta J \]
The session should feel familiar after linear regression: linear combination, cost function, gradient descent. The new pieces are sigmoid probability, log loss, and classification metrics.
Previous Topic Start Logistic Regression