KNN imputation fills missing values using similar rows
Instead of filling missing values with a global mean, KNN imputation uses nearby rows to make a context-aware replacement.
Why imputation?
Real datasets often contain missing values.
| Simple imputation | KNN imputation |
|---|---|
| Fill with global mean/median/mode | Fill using similar rows |
| Fast and simple | More context-aware |
| Ignores relationships between features | Uses relationships between features |
Procedure
- Choose features used to calculate similarity.
- Scale numeric features.
- For a row with a missing value, find its \(K\) nearest complete or usable rows.
- For numeric missing values, fill with mean or weighted mean.
- For categorical missing values, fill with mode or weighted mode.
Small working example
Suppose `BMI` is missing for one person.
| Neighbor | Age | Exercise score | BMI |
|---|---|---|---|
| 1 | 31 | 7 | 24 |
| 2 | 33 | 6 | 26 |
| 3 | 30 | 8 | 25 |
The missing value is estimated from similar rows.
Weighted KNN imputation
Closer rows can be given more weight.
This is useful when one neighbor is much more similar than the others.
Advantages of KNN imputation
Context-aware
Uses similar rows instead of a global average.
Nonlinear
Can capture local patterns without fitting an equation.
Easy intuition
Students can reason about it visually.
Limitations
| Limitation | Why it matters |
|---|---|
| Needs scaling | Distance is scale-sensitive. |
| Can be slow | Needs neighbor search for missing rows. |
| Bad with irrelevant features | Irrelevant features distort similarity. |
| Missingness pattern matters | If many features are missing, distances become unreliable. |
Simulation idea for class
Imputation question
If two customers have similar age, income, and spending score, would their purchase amount be more useful than the global average?
Often yes. That is the core intuition behind KNN imputation.
A simple classroom simulation:
- Create a small dataset with `age`, `income`, `spending_score`, and `purchase_amount`.
- Hide some `purchase_amount` values.
- Scale the features.
- For each missing row, find nearest rows using available features.
- Fill missing values with neighbor average.
- Compare mean imputation vs KNN imputation.
Validation-safe imputation workflow
Workflow question
If we impute missing values using the full dataset before splitting, what information can leak?
The test/validation distribution can influence training preprocessing, making evaluation look better than it really is.
KNN imputation should be fit only on training data inside the modeling pipeline. Then the learned imputation process is applied to validation or test rows.
Final checklist
| Before using KNN | Why |
|---|---|
| Scale numeric features | Distance is scale-sensitive. |
| Choose distance metric | Similarity definition changes neighbors. |
| Tune \(K\) | Controls bias-variance tradeoff. |
| Use validation or CV | Avoid choosing \(K\) on test data. |
| Use a pipeline | Prevent leakage from scaling or imputation. |
| Check prediction speed | KNN can be slow for large datasets. |