Concept 2

Distance metrics decide who counts as a neighbor

KNN is only as good as the similarity measure. Change the distance metric, and the nearest neighbors can change.

Choose distance based on the data

Similarity question

Should two documents be compared by word-count magnitude, or by whether they point in a similar topic direction?

For text and embeddings, direction is often more useful, so cosine similarity is commonly preferred.

Data typeUseful distance or similarityReason
Dense numeric featuresEuclidean or ManhattanCoordinates have direct numeric meaning.
Features with outliersManhattanAbsolute differences can be less dominated by large squared gaps.
Text vectors or embeddingsCosine similarity/distanceDirection often matters more than magnitude.
Categorical valuesHamming distance or careful encodingCategories do not have natural numeric distance.
Mixed dataPreprocessing plus a suitable metricNumeric, categorical, and text features may need different handling.
The distance metric is part of the model design. Changing it can change the neighbors, the prediction, and the business interpretation.

Euclidean distance

Euclidean distance is straight-line distance.

\[ d(x,z)=\sqrt{\sum_{j=1}^{p}(x_j-z_j)^2} \]
\[ x=(2,3),\quad z=(5,7) \]
\[ d(x,z)=\sqrt{(2-5)^2+(3-7)^2}=\sqrt{9+16}=5 \]

Manhattan distance

Manhattan distance adds absolute coordinate differences.

\[ d(x,z)=\sum_{j=1}^{p}|x_j-z_j| \]
\[ d(x,z)=|2-5|+|3-7|=3+4=7 \]

It is useful when movement happens along grid-like paths or when absolute changes are easier to interpret.

Euclidean Manhattan

Euclidean goes straight. Manhattan moves across axes.

Minkowski distance

Minkowski distance generalizes Euclidean and Manhattan distance.

\[ d(x,z)=\left(\sum_{j=1}^{p}|x_j-z_j|^q\right)^{1/q} \]
\(q\)Distance
\(q=1\)Manhattan distance
\(q=2\)Euclidean distance

Cosine similarity

Cosine similarity measures angle, not magnitude. It is common for text, documents, embeddings, and high-dimensional sparse data.

\[ \cos(\theta)=\frac{x\cdot z}{\|x\|\|z\|} \]

Cosine distance is often written as:

\[ d_{\cos}(x,z)=1-\cos(\theta) \]

Two vectors pointing in the same direction have high similarity even if their lengths differ.

Working example

\[ x=(1,2,2),\qquad z=(2,0,1) \]
\[ x\cdot z=(1)(2)+(2)(0)+(2)(1)=4 \]
\[ \|x\|=\sqrt{1^2+2^2+2^2}=3,\qquad \|z\|=\sqrt{2^2+0^2+1^2}=\sqrt{5} \]
\[ \cos(\theta)=\frac{4}{3\sqrt{5}}\approx 0.596 \qquad d_{\cos}=1-0.596=0.404 \]
angle doc A doc B

Cosine asks whether vectors point in a similar direction.

Feature scaling is not optional

Scaling question

If income has values in lakhs and age has values below 100, which feature will dominate Euclidean distance?

Income, unless we scale the features.

KNN uses distances, so large-scale features can dominate.

FeatureScaleRisk
Age18 to 65Moderate contribution
Annual income200000 to 5000000Dominates distance
\[ x_j^{scaled}=\frac{x_j-\mu_j}{\sigma_j} \]
Before scaling: income dominates After scaling: both features matter x-axis spread controls distance neighborhood becomes more balanced

Scaling changes the geometry of the feature space, so the nearest neighbors can change.

Without scaling, KNN may become “nearest income neighbors” rather than truly similar people.

Irrelevant features can break neighborhoods

Feature question

If we add a random customer ID number as a feature, should it help us find similar customers?

No. It can distort distance even though it has no useful meaning.

KNN assumes that distance in feature space represents real similarity. Irrelevant or noisy features add extra distance without adding signal.

\[ d(x,z)=\sqrt{(\text{useful gap})^2+(\text{noise gap})^2} \]

If the noise gap is large, it can dominate the useful gap and make the wrong points appear nearest.

FeatureUseful for similarity?Action
Age, income, usage frequencyOften yesScale and validate.
Random ID, row number, accidental codesNoRemove before KNN.
Highly duplicated or redundant columnsMaybe harmfulCheck validation performance.

Use pipelines to avoid leakage

Leakage question

Should the scaler learn mean and standard deviation from the full dataset before train-test split?

No. The validation/test data should stay unseen while preprocessing is learned.

For KNN, preprocessing is part of the model. A good workflow is:

Split data Fit scaler on train Transform validation/test Fit KNN on train Evaluate

In practice, use a pipeline so scaling, encoding, and KNN are tuned together without leaking information from validation or test data.

Other important concepts

ConceptWhy it matters
Irrelevant featuresThey add noise to distances.
High dimensionsDistances become less meaningful. This is part of the curse of dimensionality.
Categorical variablesNeed encoding or special distance measures.
OutliersCan affect local neighborhoods, especially for small \(K\).
Previous: Basics Next: Choosing K