Instance-based learning

K Nearest Neighbors: learn by comparing with similar examples.

KNN is one of the most intuitive machine learning algorithms: to predict a new point, look at the most similar past examples and let them vote or average.

Opening intuition

Quick question

If your three closest friends all recommend the same restaurant, how much does that influence your choice?

That is the everyday intuition of KNN: nearby examples influence the decision.

Imagine a new customer enters our dataset. Instead of learning coefficients or a tree, KNN asks:

\[ \text{Who are the }K\text{ most similar customers we have already seen?} \]

If most similar customers churned, classify this customer as likely churn. If their average bill was high, predict a high bill.

KNN is called a lazy learning algorithm because most of the work happens during prediction time, not training time.
new point feature 1 feature 2

The prediction depends on nearby known examples.

Session roadmap

Core idea in one line

\[ \hat{y}(x_{new}) = \text{aggregate}\left(y_i: x_i \in N_K(x_{new})\right) \]
TaskAggregationOutput
ClassificationMajority voteClass label
RegressionMean or weighted meanNumeric value
ImputationMean/mode of neighbor valuesMissing value replacement
Previous Topic Start KNN