Concept 6

Handling imbalanced data in logistic regression

Imbalanced data does not mean the model is useless. It means accuracy is no longer enough, and training/evaluation must respect the minority class.

The problem

In many real classification problems, the positive class is rare.

\[ \text{Positive class rate}=\frac{\text{number of positive examples}}{\text{total examples}} \]

Suppose a churn dataset has 950 non-churn customers and 50 churn customers.

ModelPrediction strategyAccuracyRecall for churn
Naive modelPredicts everyone as non-churn\(95\%\)\(0\%\)
Useful modelFinds many churnersMay be lowerMuch higher
High accuracy can hide a model that completely ignores the minority class.
majority class minority class 950 50

The model sees many more majority examples than minority examples.

Quick check

If only 5% of customers churn, why can 95% accuracy still be a poor result?

Because the model may simply predict “not churn” for everyone and miss every actual churner.

Method 1: use metrics that focus on the minority class

For imbalanced data, start by changing how the model is evaluated.

Precision

\[ \frac{TP}{TP+FP} \]

When the model says positive, how often is it right?

Recall

\[ \frac{TP}{TP+FN} \]

Out of actual positives, how many did it catch?

F1-score

\[ 2\cdot\frac{PR}{P+R} \]

A single-number balance between precision and recall.

MetricBest whenExample
RecallMissing positives is costlyChurn, cancer screening, fraud
PrecisionFalse alarms are costlySpam filtering, manual investigation queues
PR-AUCPositive class is rareFraud, rare disease, rare churn event
ROC-AUCRanking quality across thresholds mattersModerately imbalanced classification
For rare positives, PR-AUC is often more honest than ROC-AUC because precision is directly affected by false positives.

Method 2: tune the prediction threshold

Logistic regression produces a probability. The class label comes after choosing a threshold.

\[ \hat{y}= \begin{cases} 1, & \hat{p}\ge t \\ 0, & \hat{p}\lt t \end{cases} \]

The default threshold \(t=0.5\) is not always the best choice.

Threshold changeWhat happens?Tradeoff
Lower thresholdMore positives predictedHigher recall, more false positives
Higher thresholdFewer positives predictedHigher precision, more false negatives

For churn, if calling a customer is cheap, we may lower the threshold to catch more likely churners.

t = 0.35 t = 0.50 predicted probability

Moving the threshold left catches more positives but also increases false positives.

Decision point

For churn prediction, when would you prefer recall over precision?

When outreach is cheap and missing a churner is more costly than contacting someone who may not churn.

Method 3: class-weighted logistic regression

Class weights change the loss function. Mistakes on the minority class receive a larger penalty.

\[ J(\beta)= -\frac{1}{n}\sum_{i=1}^{n}w_{y_i} \left[y_i\log(\hat{p}_i)+(1-y_i)\log(1-\hat{p}_i)\right] \]

For two classes, a common balanced weight is:

\[ w_c=\frac{n}{K\cdot n_c} \]

where \(n\) is total rows, \(K\) is number of classes, and \(n_c\) is the number of rows in class \(c\).

ClassCountBalanced weight
Non-churn950\(\frac{1000}{2\cdot950}=0.526\)
Churn50\(\frac{1000}{2\cdot50}=10\)

A churn mistake is weighted about \(19\) times more than a non-churn mistake in this example.

\[ \nabla_\beta J=\frac{1}{n}X^T\left(w_y\odot(\hat{p}-y)\right) \]

In normal training, the gradient is dominated by the majority class because there are many more majority rows. With class weights, the minority rows produce larger gradient contributions, so the model cannot reduce loss by simply ignoring them.

Training setupWhat the model learns
Without weightsMajority-class mistakes dominate total loss.
With weightsMinority-class mistakes become expensive, so the boundary shifts toward catching more positives.
majority mistake minority mistake 0.5x 10x

Class weighting changes the training objective without changing the rows in the dataset.

In scikit-learn, this is usually `LogisticRegression(class_weight="balanced")`.

Intuition check

If one churn mistake counts like many non-churn mistakes, what should happen to recall?

Recall often increases because the model is now pushed to catch more minority-class examples. Precision may decrease because more positives may be predicted.

Method 4: oversampling the minority class

Oversampling increases the number of minority examples in the training set.

\[ \text{Before: } 950\text{ majority},\ 50\text{ minority} \qquad \text{After: } 950\text{ majority},\ 950\text{ minority} \]

Simple oversampling duplicates minority rows. The model now sees the minority class more often during training.

BenefitRisk
Often improves minority recallCan overfit repeated minority rows
Easy to explain and implementTraining data becomes larger
Oversampling must happen only on the training set, never before train-test split.
before after oversampling

Minority examples are repeated until the class balance improves.

Method 5: undersampling the majority class

Undersampling reduces majority rows so the classes become closer in size.

\[ \text{Before: } 950\text{ majority},\ 50\text{ minority} \qquad \text{After: } 50\text{ majority},\ 50\text{ minority} \]

This can work well when the majority class has many examples and losing some rows is acceptable.

BenefitRisk
Fast and simpleThrows away useful majority information
Can improve minority sensitivityMay increase variance and reduce stability
before after undersampling

Majority examples are removed to create a more balanced training set.

Method 6: SMOTE

SMOTE creates synthetic minority examples instead of simply duplicating existing ones.

\[ x_{\text{new}} = x_i + \lambda(x_{\text{neighbor}}-x_i), \qquad 0\le \lambda\le 1 \]

The new point is placed between a minority example and one of its minority neighbors.

\[ A=(10,80),\quad B=(20,100),\quad \lambda=0.4 \]
\[ B-A=(10,20),\quad 0.4(B-A)=(4,8) \]
\[ x_{new}=A+0.4(B-A)=(10,80)+(4,8)=(14,88) \]

If the two features are tenure and monthly charges, SMOTE has created a synthetic churn customer with tenure \(14\) and monthly charges \(88\).

BenefitRisk
Creates more diverse minority examplesCan create unrealistic examples
Often better than pure duplicationNeeds care with categorical features
Can improve recallCan blur class boundaries if classes overlap
SMOTE should be applied inside the training workflow, not on the full dataset before splitting.
For categorical-heavy data, plain SMOTE can create strange synthetic values. In that case, use categorical-aware methods or start with class weights and threshold tuning.
minority point neighbor synthetic

SMOTE interpolates between minority points to create synthetic rows.

Method 7: combine over and under sampling

Sometimes the best practical balance is to oversample the minority class and undersample the majority class.

\[ 950:50 \quad\longrightarrow\quad 300:300 \]

This avoids making the dataset too large while also avoiding throwing away too much majority data.

ApproachWhat it doesWhen useful
Only oversamplingDuplicates or synthesizes minority rowsSmall datasets, recall focus
Only undersamplingRemoves majority rowsVery large majority class
Combined samplingMoves both classes toward a target ratioSevere imbalance with enough data

Method 8: stratified split and stratified cross-validation

Stratification preserves the class ratio in train/test split and in each cross-validation fold.

\[ \frac{n_{positive}^{train}}{n^{train}} \approx \frac{n_{positive}^{test}}{n^{test}} \approx \frac{n_{positive}^{full}}{n^{full}} \]

Example: suppose there are \(1000\) rows with \(900\) class 0 and \(100\) class 1.

SplitClass 0Class 1Positive rate
Full data90010010%
Train, 80%7208010%
Test, 20%1802010%

Without stratification, a small test set may accidentally contain too few minority examples, making evaluation noisy.

TechniqueUse
`train_test_split(..., stratify=y)`Stable train/test class ratios
`StratifiedKFold`Stable cross-validation class ratios

With 5-fold stratified cross-validation on the same \(900:100\) dataset, each fold has about \(180\) class 0 rows and \(20\) class 1 rows. Each fold becomes the test fold once, while the other four folds become training data.

fold 1 fold 2 fold 3 same class ratio in each fold

Each split keeps a similar minority-class proportion.

FoldValidation foldTraining foldsValidation class ratio
1Fold 1Folds 2, 3, 4, 5180:20
2Fold 2Folds 1, 3, 4, 5180:20
3Fold 3Folds 1, 2, 4, 5180:20
4Fold 4Folds 1, 2, 3, 5180:20
5Fold 5Folds 1, 2, 3, 4180:20

Why stratification matters

If one validation fold accidentally has only two positive examples, can recall be trusted?

Not much. One correct or incorrect prediction can change recall dramatically. Stratified folds make evaluation more stable.

Method 9: collect more minority class data

The most powerful solution is often not algorithmic. More real minority examples improve what the model can learn.

ProblemUseful extra data
ChurnMore customers who actually churned, cancellation reasons, support history
FraudConfirmed fraud cases, investigation labels, transaction context
Medical screeningMore confirmed positive cases and follow-up outcomes
Sampling methods can help, but they cannot invent missing real-world patterns that are absent from the data.

Recommended modeling workflow

Stratified split Baseline model Check PR-AUC and recall Tune threshold Try class weights Try sampling Compare on validation data Pick business threshold
StepReason
Baseline logistic regressionCreates a reference point.
Threshold tuningOften gives a large practical gain without retraining.
Class weightsChanges training so minority mistakes matter more.
Oversampling / SMOTEGives the model more minority examples to learn from.
UndersamplingUseful when the majority class is very large.
Final decisionChoose based on business cost, not accuracy alone.

Final discussion

If two models have the same ROC-AUC, but one has much better recall at acceptable precision, which one should a churn team choose?

The second model is usually more useful because it performs better at the actual operating threshold.

Summary

MethodChanges training data?Changes loss?Main purpose
Better metricsNoNoEvaluate the minority class honestly
Threshold tuningNoNoChoose the right precision-recall tradeoff
Class weightsNoYesPenalize minority mistakes more
OversamplingYesNoShow minority examples more often
UndersamplingYesNoReduce majority dominance
SMOTEYesNoCreate synthetic minority examples
StratificationNoNoKeep fair class ratios during evaluation
Previous: Metrics Back to Overview