Sigmoid probability and decision boundary
The sigmoid function turns any real-valued linear score into a probability between 0 and 1.
Linear score: logit
First compute a linear score:
This score can be negative, positive, or zero. By itself, it is not a probability.
| Score \(z\) | Meaning before sigmoid |
|---|---|
| Large positive | Evidence leans toward class 1. |
| Near zero | Model is near the decision boundary. |
| Large negative | Evidence leans toward class 0. |
The sigmoid is the probability-conversion step.
Sigmoid function
Negative scores become probabilities near 0. Positive scores become probabilities near 1. Zero becomes 0.5.
Probability check
If \(z=0\), what is \(\sigma(z)\)?
\(\sigma(0)=0.5\), because the model is exactly at the boundary between the two classes.
Prediction threshold
After computing the probability, we convert it into a class using a threshold.
The default threshold is often \(t=0.5\), but it can be changed based on the problem. The equality case \(\hat{p}=t\) is a convention; here we assign it to class \(1\).
| Predicted probability | Threshold | Predicted class |
|---|---|---|
| 0.72 | 0.50 | 1 |
| 0.72 | 0.80 | 0 |
| 0.18 | 0.50 | 0 |
Threshold thinking
For cancer screening, would you usually prefer a lower threshold or a higher threshold?
A lower threshold may catch more true cases, increasing recall, but it can also create more false positives.
Mini numerical walkthrough
Let us compute one prediction by hand. Suppose we are predicting whether a customer will churn.
For a customer with 8 support calls and a monthly contract value of 1:
Convert score into probability:
With threshold \(t=0.5\), the predicted class is:
| Quantity | Value | Meaning |
|---|---|---|
| \(z\) | -1.48 | Evidence leans toward class 0. |
| \(\hat{p}\) | 0.185 | Estimated churn probability is 18.5%. |
| \(t\) | 0.5 | Decision cutoff. |
| \(\hat{y}\) | 0 | Predicted not churn. |
Threshold check
If the business wants to proactively call risky customers, might it use a threshold lower than 0.5?
Yes. A lower threshold catches more possible churners, but also creates more false alarms.
Decision boundary
With threshold \(0.5\), the decision boundary happens where:
For two features:
This is a line in the two-feature input space. Logistic regression produces a linear decision boundary in the original feature space.
If we add transformed features, logistic regression can create nonlinear boundaries in the original input space:
The boundary is still linear in the transformed features, but it may look curved when plotted against the original \(x_1,x_2\).
Points on one side are predicted class 0; points on the other side are predicted class 1.
Linear vs nonlinear boundary with feature engineering
Without transformed features, logistic regression draws a linear boundary. With transformed features, the model is still linear in the new feature space, but the boundary can look curved in the original plot.
A straight boundary cannot naturally separate a center cluster from an outside ring.
Feature engineering can make a logistic model draw a curved boundary in the original feature space.