Concept 4

KNN imputation fills missing values using similar rows

Instead of filling missing values with a global mean, KNN imputation uses nearby rows to make a context-aware replacement.

Why imputation?

Real datasets often contain missing values.

Simple imputationKNN imputation
Fill with global mean/median/modeFill using similar rows
Fast and simpleMore context-aware
Ignores relationships between featuresUses relationships between features

Procedure

  1. Choose features used to calculate similarity.
  2. Scale numeric features.
  3. For a row with a missing value, find its \(K\) nearest complete or usable rows.
  4. For numeric missing values, fill with mean or weighted mean.
  5. For categorical missing values, fill with mode or weighted mode.
\[ x_{missing} \leftarrow \frac{1}{K}\sum_{i\in N_K(x)}x_i \]
In practical libraries, KNN imputation computes distance using the features available for both rows. For example, scikit-learn's `KNNImputer` uses a missing-value-aware distance rather than requiring every row to be fully complete.

Small working example

Suppose `BMI` is missing for one person.

NeighborAgeExercise scoreBMI
131724
233626
330825
\[ \text{imputed BMI}=\frac{24+26+25}{3}=25 \]
BMI missing use nearest rows to fill

The missing value is estimated from similar rows.

Weighted KNN imputation

Closer rows can be given more weight.

\[ w_i=\frac{1}{d_i+\epsilon} \]
\[ x_{missing} \leftarrow \frac{\sum_{i\in N_K(x)}w_i x_i}{\sum_{i\in N_K(x)}w_i} \]

This is useful when one neighbor is much more similar than the others.

Advantages of KNN imputation

Context-aware

Uses similar rows instead of a global average.

Nonlinear

Can capture local patterns without fitting an equation.

Easy intuition

Students can reason about it visually.

Limitations

LimitationWhy it matters
Needs scalingDistance is scale-sensitive.
Can be slowNeeds neighbor search for missing rows.
Bad with irrelevant featuresIrrelevant features distort similarity.
Missingness pattern mattersIf many features are missing, distances become unreliable.

Simulation idea for class

Imputation question

If two customers have similar age, income, and spending score, would their purchase amount be more useful than the global average?

Often yes. That is the core intuition behind KNN imputation.

A simple classroom simulation:

  1. Create a small dataset with `age`, `income`, `spending_score`, and `purchase_amount`.
  2. Hide some `purchase_amount` values.
  3. Scale the features.
  4. For each missing row, find nearest rows using available features.
  5. Fill missing values with neighbor average.
  6. Compare mean imputation vs KNN imputation.

Validation-safe imputation workflow

Workflow question

If we impute missing values using the full dataset before splitting, what information can leak?

The test/validation distribution can influence training preprocessing, making evaluation look better than it really is.

KNN imputation should be fit only on training data inside the modeling pipeline. Then the learned imputation process is applied to validation or test rows.

Split Fit imputer on train Transform validation/test Train model Evaluate once
For teaching, this is a good place to remind students that preprocessing can cause data leakage just like model training can.

Final checklist

Before using KNNWhy
Scale numeric featuresDistance is scale-sensitive.
Choose distance metricSimilarity definition changes neighbors.
Tune \(K\)Controls bias-variance tradeoff.
Use validation or CVAvoid choosing \(K\) on test data.
Use a pipelinePrevent leakage from scaling or imputation.
Check prediction speedKNN can be slow for large datasets.
Previous: Choosing K Back to Overview