Concept 1

From regression to classification

Linear regression predicts a continuous number. Logistic regression predicts the probability of a class, usually for a yes/no outcome.

Classification problem setup

In binary classification, the target has two possible values:

\[ y \in \{0,1\} \]

The goal is not just to predict a label. A better goal is to predict the probability of the positive class:

\[ \hat{p}=P(y=1\mid x) \]

Spam detection

\(y=1\): spam. \(y=0\): not spam.

Customer churn

\(y=1\): customer leaves. \(y=0\): customer stays.

Loan default

\(y=1\): default. \(y=0\): repaid.

Target definition

In fraud detection, which class should be coded as \(1\): fraud or non-fraud?

Usually fraud is coded as \(1\), because it is the event we care about detecting.

Why linear regression is not enough

A linear model can output any real number:

\[ \hat{y}=\beta_0+\beta_1x \]

For classification, that creates two problems:

ProblemWhy it hurts
Predictions can be below 0 or above 1Those values cannot be probabilities.
Error is not probability-awareConfident wrong classification should be punished more strongly.
1 0 linear line goes beyond probability range

A straight-line output is not naturally limited to the range \([0,1]\).

The logistic regression idea

Logistic regression keeps the useful linear part, but converts it into a probability.

features \(x\) linear score \(z\) sigmoid probability \(\hat{p}\) class decision
\[ z=\beta_0+\beta_1x_1+\cdots+\beta_px_p \qquad \hat{p}=\sigma(z) \]
The word “regression” remains in the name because the model estimates a continuous probability and uses a linear combination of features. But the task is classification.
Overview Next: Sigmoid