Concept 5

Evaluation metrics, thresholds, and imbalanced classes

Logistic regression outputs probabilities. Evaluation depends on how those probabilities are converted into decisions and what mistakes matter most.

Confusion matrix

A confusion matrix compares predicted classes with actual classes.

Predicted 0Predicted 1
Actual 0True Negative (TN)False Positive (FP)
Actual 1False Negative (FN)True Positive (TP)

Mistake cost

In cancer detection, which error is usually more dangerous: false positive or false negative?

False negative is often more dangerous because a real case is missed.

Core metrics

Accuracy

\[ \frac{TP+TN}{TP+TN+FP+FN} \]

Overall fraction of correct predictions.

Precision

\[ \frac{TP}{TP+FP} \]

Of predicted positives, how many were actually positive?

Recall

\[ \frac{TP}{TP+FN} \]

Of actual positives, how many did we catch?

\[ F1 = 2\cdot\frac{\text{Precision}\cdot\text{Recall}}{\text{Precision}+\text{Recall}} \]
Precision asks whether positive predictions are trustworthy. Recall asks whether actual positives are being missed.

Threshold tuning

Changing the threshold changes the confusion matrix.

ThresholdEffectRisk
Lower thresholdPredict more positives, recall usually increases.More false positives.
Higher thresholdPredict fewer positives, precision may increase.More false negatives.
threshold predicted probability

Move the threshold left or right depending on which mistakes are costlier.

Threshold as a business decision

The best threshold is not always \(0.5\). It depends on which mistake is more expensive.

ProblemCostlier mistakeThreshold directionMetric to watch
Cancer screeningFalse negative: missing a real caseLower thresholdRecall, sensitivity
Fraud detectionUsually false negative, but false positives also create investigation costOften lower, then tune carefullyRecall, precision, PR-AUC
Spam filteringFalse positive: important email goes to spamHigher thresholdPrecision
Loan approval riskFalse negative: approving a likely defaulterHigher risk cutoff or stricter policyPrecision for risky class, recall for defaults
Customer churn outreachFalse negative: missing a likely churnerLower threshold if outreach is cheapRecall and campaign ROI

Decision cost

If calling a customer is cheap but missing a churner is expensive, should the threshold be high or low?

Lower. The business may accept more false positives to catch more likely churners.

Class imbalance

Accuracy can be misleading when one class is much more common than the other.

\[ \text{Fraud rate}=1\% \]

A model that predicts “not fraud” for everyone gets 99% accuracy, but catches no fraud.

MetricUseful when
PrecisionFalse positives are costly.
RecallFalse negatives are costly.
F1-scoreNeed a balance between precision and recall.
PR-AUCPositive class is rare.
ROC-AUCNeed threshold-independent ranking quality, especially with moderate imbalance.

Metric choice

For fraud detection, why might 99% accuracy be a bad sign rather than a good sign?

The model may simply be predicting the majority class and missing rare fraud cases.

ROC curve and PR curve

Both ROC and PR curves are created by sweeping the classification threshold from high to low and recalculating metrics at each threshold.

false positive rate true positive rate ROC curve

ROC curve shows recall/TPR against false positive rate across thresholds.

recall precision PR curve

PR curve is often more informative when the positive class is rare.

ROC curve in detail

ROC stands for Receiver Operating Characteristic. It asks: as we lower the threshold and catch more positives, how many false alarms do we create?

\[ \text{True Positive Rate}=\text{Recall}=\frac{TP}{TP+FN} \]
\[ \text{False Positive Rate}=\frac{FP}{FP+TN} \]
ThresholdTPR / RecallFPRInterpretation
Very highLowLowModel predicts very few positives.
MediumHigherSome false positivesModel catches more positives but creates more alarms.
Very lowVery highHighModel predicts many positives.

ROC-AUC summarizes the curve into one number. A higher ROC-AUC means the model tends to rank true positives above true negatives.

ROC-AUC is about ranking quality across thresholds. It does not directly tell you which threshold to use in production.

ROC interpretation

If recall increases because we lowered the threshold, what usually happens to the false positive rate?

It usually increases, because more negative examples also cross the lower threshold.

PR curve in detail

PR stands for Precision-Recall. It focuses on the positive class.

\[ \text{Precision}=\frac{TP}{TP+FP} \qquad \text{Recall}=\frac{TP}{TP+FN} \]

PR curves are especially useful when the positive class is rare. In rare-event problems, ROC can sometimes look good even when precision is poor.

ThresholdPrecisionRecallInterpretation
HighOften highLowOnly the most confident positives are predicted.
MediumModerateModerate to highMore positives are caught, with more false positives.
LowCan fallHighMost positives are caught, but many predicted positives may be wrong.
precision often decreases recall increases lowering threshold metric value

As threshold decreases, recall usually increases because more positives are captured. Precision may decrease because more false positives are included.

For fraud, disease detection, rare failure detection, and other rare-positive problems, PR-AUC is often more informative than ROC-AUC.

Choosing between ROC and PR

SituationPreferWhy
Classes are fairly balancedROC-AUC or PR-AUCBoth can be informative.
Positive class is rarePR-AUCIt focuses on how useful positive predictions are.
Need ranking quality across all thresholdsROC-AUCIt measures separation between positives and negatives.
Business cares about catching positivesRecall and PR curveMissing positives is costly.
Business cares about false alarmsPrecision and PR curvePredicted positives must be trustworthy.

Curve choice

In fraud detection with 1% fraud cases, why might PR curve be more useful than ROC curve?

Because PR curve focuses on the rare positive class and directly shows the tradeoff between catching fraud and keeping alerts trustworthy.

Assumptions and limitations

PointMeaning
Linear decision boundaryIn original features, logistic regression separates classes using a line, plane, or hyperplane.
Feature engineering mattersNonlinear boundaries need transformed features or another model.
Probability calibrationPredicted probabilities may need calibration for high-stakes decisions.
Outliers can influence coefficientsStrong outliers or mislabeled points can affect the fitted boundary.
Regularized models need scalingScaling makes the penalty treat features more fairly.
Logistic regression is a strong baseline because it is simple, interpretable, fast, and often surprisingly competitive.

Final mental model

\[ \text{linear score} \rightarrow \text{sigmoid probability} \rightarrow \text{thresholded class} \rightarrow \text{classification metrics} \]

End check

What is the biggest conceptual difference between linear regression and logistic regression?

Linear regression predicts a continuous value. Logistic regression predicts a probability for classification.

Previous: Interpretation Next: Handling Imbalance