Concept 2

Sigmoid probability and decision boundary

The sigmoid function turns any real-valued linear score into a probability between 0 and 1.

Linear score: logit

First compute a linear score:

\[ z = \beta_0+\beta_1x_1+\beta_2x_2+\cdots+\beta_px_p \]

This score can be negative, positive, or zero. By itself, it is not a probability.

Score \(z\)Meaning before sigmoid
Large positiveEvidence leans toward class 1.
Near zeroModel is near the decision boundary.
Large negativeEvidence leans toward class 0.
features score z any real value sigmoid probability 0 to 1

The sigmoid is the probability-conversion step.

Sigmoid function

\[ \sigma(z)=\frac{1}{1+e^{-z}} \]
z=0, p=0.5 1 0 z probability

Negative scores become probabilities near 0. Positive scores become probabilities near 1. Zero becomes 0.5.

Probability check

If \(z=0\), what is \(\sigma(z)\)?

\(\sigma(0)=0.5\), because the model is exactly at the boundary between the two classes.

Prediction threshold

After computing the probability, we convert it into a class using a threshold.

\[ \hat{y}= \begin{cases} 1, & \text{if } \hat{p}\ge t \\ 0, & \text{if } \hat{p}\lt t \end{cases} \]

The default threshold is often \(t=0.5\), but it can be changed based on the problem. The equality case \(\hat{p}=t\) is a convention; here we assign it to class \(1\).

Predicted probabilityThresholdPredicted class
0.720.501
0.720.800
0.180.500

Threshold thinking

For cancer screening, would you usually prefer a lower threshold or a higher threshold?

A lower threshold may catch more true cases, increasing recall, but it can also create more false positives.

Mini numerical walkthrough

Let us compute one prediction by hand. Suppose we are predicting whether a customer will churn.

\[ z=-3+0.04(\text{support calls})+1.2(\text{monthly contract}) \]

For a customer with 8 support calls and a monthly contract value of 1:

\[ z=-3+0.04(8)+1.2(1)=-1.48 \]

Convert score into probability:

\[ \hat{p}=\sigma(-1.48)=\frac{1}{1+e^{1.48}}\approx0.185 \]

With threshold \(t=0.5\), the predicted class is:

\[ 0.185 \lt 0.5 \Rightarrow \hat{y}=0 \]
QuantityValueMeaning
\(z\)-1.48Evidence leans toward class 0.
\(\hat{p}\)0.185Estimated churn probability is 18.5%.
\(t\)0.5Decision cutoff.
\(\hat{y}\)0Predicted not churn.

Threshold check

If the business wants to proactively call risky customers, might it use a threshold lower than 0.5?

Yes. A lower threshold catches more possible churners, but also creates more false alarms.

Decision boundary

With threshold \(0.5\), the decision boundary happens where:

\[ \hat{p}=0.5 \Rightarrow z=0 \]

For two features:

\[ \beta_0+\beta_1x_1+\beta_2x_2=0 \]

This is a line in the two-feature input space. Logistic regression produces a linear decision boundary in the original feature space.

Important intuition: the sigmoid makes the predicted probability nonlinear, but the \(0.5\) decision boundary is still where the linear score \(z\) equals zero. So with original features, logistic regression separates classes using a line, plane, or hyperplane.

If we add transformed features, logistic regression can create nonlinear boundaries in the original input space:

\[ z=\beta_0+\beta_1x_1+\beta_2x_2+\beta_3x_1^2+\beta_4x_2^2+\beta_5x_1x_2 \]

The boundary is still linear in the transformed features, but it may look curved when plotted against the original \(x_1,x_2\).

z=0 boundary x1 x2

Points on one side are predicted class 0; points on the other side are predicted class 1.

Linear vs nonlinear boundary with feature engineering

Without transformed features, logistic regression draws a linear boundary. With transformed features, the model is still linear in the new feature space, but the boundary can look curved in the original plot.

original x1, x2 only: line struggles

A straight boundary cannot naturally separate a center cluster from an outside ring.

with x1^2 and x2^2: curved boundary

Feature engineering can make a logistic model draw a curved boundary in the original feature space.

\[ z=\beta_0+\beta_1x_1+\beta_2x_2+\beta_3x_1^2+\beta_4x_2^2 \]
The model is still logistic regression. The difference is that we changed the input features.
Previous: Setup Next: Loss