Evaluation metrics, thresholds, and imbalanced classes
Logistic regression outputs probabilities. Evaluation depends on how those probabilities are converted into decisions and what mistakes matter most.
Confusion matrix
A confusion matrix compares predicted classes with actual classes.
| Predicted 0 | Predicted 1 | |
|---|---|---|
| Actual 0 | True Negative (TN) | False Positive (FP) |
| Actual 1 | False Negative (FN) | True Positive (TP) |
Mistake cost
In cancer detection, which error is usually more dangerous: false positive or false negative?
False negative is often more dangerous because a real case is missed.
Core metrics
Accuracy
Overall fraction of correct predictions.
Precision
Of predicted positives, how many were actually positive?
Recall
Of actual positives, how many did we catch?
Threshold tuning
Changing the threshold changes the confusion matrix.
| Threshold | Effect | Risk |
|---|---|---|
| Lower threshold | Predict more positives, recall usually increases. | More false positives. |
| Higher threshold | Predict fewer positives, precision may increase. | More false negatives. |
Move the threshold left or right depending on which mistakes are costlier.
Threshold as a business decision
The best threshold is not always \(0.5\). It depends on which mistake is more expensive.
| Problem | Costlier mistake | Threshold direction | Metric to watch |
|---|---|---|---|
| Cancer screening | False negative: missing a real case | Lower threshold | Recall, sensitivity |
| Fraud detection | Usually false negative, but false positives also create investigation cost | Often lower, then tune carefully | Recall, precision, PR-AUC |
| Spam filtering | False positive: important email goes to spam | Higher threshold | Precision |
| Loan approval risk | False negative: approving a likely defaulter | Higher risk cutoff or stricter policy | Precision for risky class, recall for defaults |
| Customer churn outreach | False negative: missing a likely churner | Lower threshold if outreach is cheap | Recall and campaign ROI |
Decision cost
If calling a customer is cheap but missing a churner is expensive, should the threshold be high or low?
Lower. The business may accept more false positives to catch more likely churners.
Class imbalance
Accuracy can be misleading when one class is much more common than the other.
A model that predicts “not fraud” for everyone gets 99% accuracy, but catches no fraud.
| Metric | Useful when |
|---|---|
| Precision | False positives are costly. |
| Recall | False negatives are costly. |
| F1-score | Need a balance between precision and recall. |
| PR-AUC | Positive class is rare. |
| ROC-AUC | Need threshold-independent ranking quality, especially with moderate imbalance. |
Metric choice
For fraud detection, why might 99% accuracy be a bad sign rather than a good sign?
The model may simply be predicting the majority class and missing rare fraud cases.
ROC curve and PR curve
Both ROC and PR curves are created by sweeping the classification threshold from high to low and recalculating metrics at each threshold.
ROC curve shows recall/TPR against false positive rate across thresholds.
PR curve is often more informative when the positive class is rare.
ROC curve in detail
ROC stands for Receiver Operating Characteristic. It asks: as we lower the threshold and catch more positives, how many false alarms do we create?
| Threshold | TPR / Recall | FPR | Interpretation |
|---|---|---|---|
| Very high | Low | Low | Model predicts very few positives. |
| Medium | Higher | Some false positives | Model catches more positives but creates more alarms. |
| Very low | Very high | High | Model predicts many positives. |
ROC-AUC summarizes the curve into one number. A higher ROC-AUC means the model tends to rank true positives above true negatives.
ROC interpretation
If recall increases because we lowered the threshold, what usually happens to the false positive rate?
It usually increases, because more negative examples also cross the lower threshold.
PR curve in detail
PR stands for Precision-Recall. It focuses on the positive class.
PR curves are especially useful when the positive class is rare. In rare-event problems, ROC can sometimes look good even when precision is poor.
| Threshold | Precision | Recall | Interpretation |
|---|---|---|---|
| High | Often high | Low | Only the most confident positives are predicted. |
| Medium | Moderate | Moderate to high | More positives are caught, with more false positives. |
| Low | Can fall | High | Most positives are caught, but many predicted positives may be wrong. |
As threshold decreases, recall usually increases because more positives are captured. Precision may decrease because more false positives are included.
Choosing between ROC and PR
| Situation | Prefer | Why |
|---|---|---|
| Classes are fairly balanced | ROC-AUC or PR-AUC | Both can be informative. |
| Positive class is rare | PR-AUC | It focuses on how useful positive predictions are. |
| Need ranking quality across all thresholds | ROC-AUC | It measures separation between positives and negatives. |
| Business cares about catching positives | Recall and PR curve | Missing positives is costly. |
| Business cares about false alarms | Precision and PR curve | Predicted positives must be trustworthy. |
Curve choice
In fraud detection with 1% fraud cases, why might PR curve be more useful than ROC curve?
Because PR curve focuses on the rare positive class and directly shows the tradeoff between catching fraud and keeping alerts trustworthy.
Assumptions and limitations
| Point | Meaning |
|---|---|
| Linear decision boundary | In original features, logistic regression separates classes using a line, plane, or hyperplane. |
| Feature engineering matters | Nonlinear boundaries need transformed features or another model. |
| Probability calibration | Predicted probabilities may need calibration for high-stakes decisions. |
| Outliers can influence coefficients | Strong outliers or mislabeled points can affect the fitted boundary. |
| Regularized models need scaling | Scaling makes the penalty treat features more fairly. |
Final mental model
End check
What is the biggest conceptual difference between linear regression and logistic regression?
Linear regression predicts a continuous value. Logistic regression predicts a probability for classification.