Distance metrics decide who counts as a neighbor
KNN is only as good as the similarity measure. Change the distance metric, and the nearest neighbors can change.
Choose distance based on the data
Similarity question
Should two documents be compared by word-count magnitude, or by whether they point in a similar topic direction?
For text and embeddings, direction is often more useful, so cosine similarity is commonly preferred.
| Data type | Useful distance or similarity | Reason |
|---|---|---|
| Dense numeric features | Euclidean or Manhattan | Coordinates have direct numeric meaning. |
| Features with outliers | Manhattan | Absolute differences can be less dominated by large squared gaps. |
| Text vectors or embeddings | Cosine similarity/distance | Direction often matters more than magnitude. |
| Categorical values | Hamming distance or careful encoding | Categories do not have natural numeric distance. |
| Mixed data | Preprocessing plus a suitable metric | Numeric, categorical, and text features may need different handling. |
Euclidean distance
Euclidean distance is straight-line distance.
Manhattan distance
Manhattan distance adds absolute coordinate differences.
It is useful when movement happens along grid-like paths or when absolute changes are easier to interpret.
Euclidean goes straight. Manhattan moves across axes.
Minkowski distance
Minkowski distance generalizes Euclidean and Manhattan distance.
| \(q\) | Distance |
|---|---|
| \(q=1\) | Manhattan distance |
| \(q=2\) | Euclidean distance |
Cosine similarity
Cosine similarity measures angle, not magnitude. It is common for text, documents, embeddings, and high-dimensional sparse data.
Cosine distance is often written as:
Two vectors pointing in the same direction have high similarity even if their lengths differ.
Working example
Cosine asks whether vectors point in a similar direction.
Feature scaling is not optional
Scaling question
If income has values in lakhs and age has values below 100, which feature will dominate Euclidean distance?
Income, unless we scale the features.
KNN uses distances, so large-scale features can dominate.
| Feature | Scale | Risk |
|---|---|---|
| Age | 18 to 65 | Moderate contribution |
| Annual income | 200000 to 5000000 | Dominates distance |
Scaling changes the geometry of the feature space, so the nearest neighbors can change.
Irrelevant features can break neighborhoods
Feature question
If we add a random customer ID number as a feature, should it help us find similar customers?
No. It can distort distance even though it has no useful meaning.
KNN assumes that distance in feature space represents real similarity. Irrelevant or noisy features add extra distance without adding signal.
If the noise gap is large, it can dominate the useful gap and make the wrong points appear nearest.
| Feature | Useful for similarity? | Action |
|---|---|---|
| Age, income, usage frequency | Often yes | Scale and validate. |
| Random ID, row number, accidental codes | No | Remove before KNN. |
| Highly duplicated or redundant columns | Maybe harmful | Check validation performance. |
Use pipelines to avoid leakage
Leakage question
Should the scaler learn mean and standard deviation from the full dataset before train-test split?
No. The validation/test data should stay unseen while preprocessing is learned.
For KNN, preprocessing is part of the model. A good workflow is:
In practice, use a pipeline so scaling, encoding, and KNN are tuned together without leaking information from validation or test data.
Other important concepts
| Concept | Why it matters |
|---|---|
| Irrelevant features | They add noise to distances. |
| High dimensions | Distances become less meaningful. This is part of the curse of dimensionality. |
| Categorical variables | Need encoding or special distance measures. |
| Outliers | Can affect local neighborhoods, especially for small \(K\). |