Handling imbalanced data in logistic regression
Imbalanced data does not mean the model is useless. It means accuracy is no longer enough, and training/evaluation must respect the minority class.
The problem
In many real classification problems, the positive class is rare.
Suppose a churn dataset has 950 non-churn customers and 50 churn customers.
| Model | Prediction strategy | Accuracy | Recall for churn |
|---|---|---|---|
| Naive model | Predicts everyone as non-churn | \(95\%\) | \(0\%\) |
| Useful model | Finds many churners | May be lower | Much higher |
The model sees many more majority examples than minority examples.
Quick check
If only 5% of customers churn, why can 95% accuracy still be a poor result?
Because the model may simply predict “not churn” for everyone and miss every actual churner.
Method 1: use metrics that focus on the minority class
For imbalanced data, start by changing how the model is evaluated.
Precision
When the model says positive, how often is it right?
Recall
Out of actual positives, how many did it catch?
F1-score
A single-number balance between precision and recall.
| Metric | Best when | Example |
|---|---|---|
| Recall | Missing positives is costly | Churn, cancer screening, fraud |
| Precision | False alarms are costly | Spam filtering, manual investigation queues |
| PR-AUC | Positive class is rare | Fraud, rare disease, rare churn event |
| ROC-AUC | Ranking quality across thresholds matters | Moderately imbalanced classification |
Method 2: tune the prediction threshold
Logistic regression produces a probability. The class label comes after choosing a threshold.
The default threshold \(t=0.5\) is not always the best choice.
| Threshold change | What happens? | Tradeoff |
|---|---|---|
| Lower threshold | More positives predicted | Higher recall, more false positives |
| Higher threshold | Fewer positives predicted | Higher precision, more false negatives |
For churn, if calling a customer is cheap, we may lower the threshold to catch more likely churners.
Moving the threshold left catches more positives but also increases false positives.
Decision point
For churn prediction, when would you prefer recall over precision?
When outreach is cheap and missing a churner is more costly than contacting someone who may not churn.
Method 3: class-weighted logistic regression
Class weights change the loss function. Mistakes on the minority class receive a larger penalty.
For two classes, a common balanced weight is:
where \(n\) is total rows, \(K\) is number of classes, and \(n_c\) is the number of rows in class \(c\).
| Class | Count | Balanced weight |
|---|---|---|
| Non-churn | 950 | \(\frac{1000}{2\cdot950}=0.526\) |
| Churn | 50 | \(\frac{1000}{2\cdot50}=10\) |
A churn mistake is weighted about \(19\) times more than a non-churn mistake in this example.
In normal training, the gradient is dominated by the majority class because there are many more majority rows. With class weights, the minority rows produce larger gradient contributions, so the model cannot reduce loss by simply ignoring them.
| Training setup | What the model learns |
|---|---|
| Without weights | Majority-class mistakes dominate total loss. |
| With weights | Minority-class mistakes become expensive, so the boundary shifts toward catching more positives. |
Class weighting changes the training objective without changing the rows in the dataset.
Intuition check
If one churn mistake counts like many non-churn mistakes, what should happen to recall?
Recall often increases because the model is now pushed to catch more minority-class examples. Precision may decrease because more positives may be predicted.
Method 4: oversampling the minority class
Oversampling increases the number of minority examples in the training set.
Simple oversampling duplicates minority rows. The model now sees the minority class more often during training.
| Benefit | Risk |
|---|---|
| Often improves minority recall | Can overfit repeated minority rows |
| Easy to explain and implement | Training data becomes larger |
Minority examples are repeated until the class balance improves.
Method 5: undersampling the majority class
Undersampling reduces majority rows so the classes become closer in size.
This can work well when the majority class has many examples and losing some rows is acceptable.
| Benefit | Risk |
|---|---|
| Fast and simple | Throws away useful majority information |
| Can improve minority sensitivity | May increase variance and reduce stability |
Majority examples are removed to create a more balanced training set.
Method 6: SMOTE
SMOTE creates synthetic minority examples instead of simply duplicating existing ones.
The new point is placed between a minority example and one of its minority neighbors.
If the two features are tenure and monthly charges, SMOTE has created a synthetic churn customer with tenure \(14\) and monthly charges \(88\).
| Benefit | Risk |
|---|---|
| Creates more diverse minority examples | Can create unrealistic examples |
| Often better than pure duplication | Needs care with categorical features |
| Can improve recall | Can blur class boundaries if classes overlap |
SMOTE interpolates between minority points to create synthetic rows.
Method 7: combine over and under sampling
Sometimes the best practical balance is to oversample the minority class and undersample the majority class.
This avoids making the dataset too large while also avoiding throwing away too much majority data.
| Approach | What it does | When useful |
|---|---|---|
| Only oversampling | Duplicates or synthesizes minority rows | Small datasets, recall focus |
| Only undersampling | Removes majority rows | Very large majority class |
| Combined sampling | Moves both classes toward a target ratio | Severe imbalance with enough data |
Method 8: stratified split and stratified cross-validation
Stratification preserves the class ratio in train/test split and in each cross-validation fold.
Example: suppose there are \(1000\) rows with \(900\) class 0 and \(100\) class 1.
| Split | Class 0 | Class 1 | Positive rate |
|---|---|---|---|
| Full data | 900 | 100 | 10% |
| Train, 80% | 720 | 80 | 10% |
| Test, 20% | 180 | 20 | 10% |
Without stratification, a small test set may accidentally contain too few minority examples, making evaluation noisy.
| Technique | Use |
|---|---|
| `train_test_split(..., stratify=y)` | Stable train/test class ratios |
| `StratifiedKFold` | Stable cross-validation class ratios |
With 5-fold stratified cross-validation on the same \(900:100\) dataset, each fold has about \(180\) class 0 rows and \(20\) class 1 rows. Each fold becomes the test fold once, while the other four folds become training data.
Each split keeps a similar minority-class proportion.
| Fold | Validation fold | Training folds | Validation class ratio |
|---|---|---|---|
| 1 | Fold 1 | Folds 2, 3, 4, 5 | 180:20 |
| 2 | Fold 2 | Folds 1, 3, 4, 5 | 180:20 |
| 3 | Fold 3 | Folds 1, 2, 4, 5 | 180:20 |
| 4 | Fold 4 | Folds 1, 2, 3, 5 | 180:20 |
| 5 | Fold 5 | Folds 1, 2, 3, 4 | 180:20 |
Why stratification matters
If one validation fold accidentally has only two positive examples, can recall be trusted?
Not much. One correct or incorrect prediction can change recall dramatically. Stratified folds make evaluation more stable.
Method 9: collect more minority class data
The most powerful solution is often not algorithmic. More real minority examples improve what the model can learn.
| Problem | Useful extra data |
|---|---|
| Churn | More customers who actually churned, cancellation reasons, support history |
| Fraud | Confirmed fraud cases, investigation labels, transaction context |
| Medical screening | More confirmed positive cases and follow-up outcomes |
Recommended modeling workflow
| Step | Reason |
|---|---|
| Baseline logistic regression | Creates a reference point. |
| Threshold tuning | Often gives a large practical gain without retraining. |
| Class weights | Changes training so minority mistakes matter more. |
| Oversampling / SMOTE | Gives the model more minority examples to learn from. |
| Undersampling | Useful when the majority class is very large. |
| Final decision | Choose based on business cost, not accuracy alone. |
Final discussion
If two models have the same ROC-AUC, but one has much better recall at acceptable precision, which one should a churn team choose?
The second model is usually more useful because it performs better at the actual operating threshold.
Summary
| Method | Changes training data? | Changes loss? | Main purpose |
|---|---|---|---|
| Better metrics | No | No | Evaluate the minority class honestly |
| Threshold tuning | No | No | Choose the right precision-recall tradeoff |
| Class weights | No | Yes | Penalize minority mistakes more |
| Oversampling | Yes | No | Show minority examples more often |
| Undersampling | Yes | No | Reduce majority dominance |
| SMOTE | Yes | No | Create synthetic minority examples |
| Stratification | No | No | Keep fair class ratios during evaluation |