KNN for classification and regression
KNN makes predictions from local evidence. First understand the idea, then the steps, then how classification and regression use the same neighbor logic.
What is KNN?
K Nearest Neighbors stores the training data. For a new point, it calculates distance to known points, finds the closest \(K\), and combines their target values.
Algorithm procedure
- Choose \(K\).
- Choose a distance or similarity measure.
- Scale features if needed.
- For the new point, compute distance to all training points.
- Sort by distance and take the closest \(K\).
- Predict using majority vote or average.
KNN for classification
For classification, KNN uses majority vote among neighbors.
| Neighbor | Class |
|---|---|
| 1 | Churn |
| 2 | No churn |
| 3 | Churn |
| 4 | Churn |
| 5 | No churn |
For \(K=5\), churn wins \(3\) vs \(2\), so the predicted class is churn.
KNN can also produce a simple probability estimate from vote fraction:
Classification predicts the majority class among the \(K\) nearest neighbors.
Probability estimates and tie-breaking
Vote question
If 4 neighbors vote as 2 churn and 2 no churn, is the model truly confident?
No. Equal votes indicate uncertainty, and the final class depends on the tie-breaking rule.
KNN classification can estimate class probability using the fraction of neighbors from each class.
For example, if 3 out of 5 neighbors are churn customers, then the local probability estimate is:
These probabilities are not learned from a smooth equation. They are local vote fractions, so they can change sharply when the neighborhood changes.
KNN for regression
For regression, KNN averages the target values of the nearest neighbors.
If the nearest house prices are:
Regression predicts a local average rather than a class vote.
KNN decision boundary
Boundary question
If KNN predicts from nearby examples, does it need to draw one straight line between classes?
No. KNN can create curved and local boundaries because each region is decided by nearby points.
Linear and logistic regression usually learn one global shape from the data. KNN does not learn a single equation for the boundary. The boundary appears from repeated local voting across the feature space.
| Model | Boundary intuition |
|---|---|
| Logistic regression | Usually one linear boundary unless features are transformed. |
| KNN | Many local voting regions; boundary can be nonlinear. |
KNN boundaries come from local neighborhoods, so they can bend around data patterns.
What is stored as the trained KNN model?
Model storage question
If KNN has no learned slope or coefficient, what must it remember to predict later?
It must remember the training examples and the rule for measuring similarity.
In linear regression and logistic regression, training mainly learns parameters such as coefficients and intercept.
KNN is different. A fitted KNN model mainly stores the training examples, their labels, and the rules needed to compare future points.
| Stored item | Why it is needed |
|---|---|
| \(X_{train}\) | New points are compared with past feature values. |
| \(y_{train}\) | Neighbor labels or target values are used for voting/averaging. |
| \(K\) | Controls how many neighbors are used. |
| Distance metric | Defines what “near” means. |
| Scaler/encoder | Future data must be transformed exactly like training data. |
When do we use KNN?
Where would KNN fit?
Would KNN make sense for recommending similar users, similar products, or similar documents?
Yes, because similarity is the central idea in those problems.
| Use KNN when | Be careful when |
|---|---|
| Similar examples should have similar labels. | There are many irrelevant features. |
| The dataset is not too large. | Prediction latency must be extremely low. |
| The decision boundary may be nonlinear. | Features have very different scales. |
| You need a simple baseline. | Data is high-dimensional and sparse. |
How do we evaluate KNN?
Metric question
If missing a churn customer is costly, should accuracy be the only metric?
No. Recall, precision, F1, and confusion matrix become important when mistakes have different costs.
| Problem type | Useful metrics | What they tell us |
|---|---|---|
| Classification | Accuracy, precision, recall, F1, confusion matrix | How often classes are predicted correctly and which errors happen. |
| Imbalanced classification | Recall, precision, F1, ROC-AUC, PR-AUC | Whether minority-class cases are being found. |
| Regression | MAE, MSE, RMSE, \(R^2\) | How far numeric predictions are from actual values. |
KNN compared with earlier models
| Property | Linear/logistic regression | KNN |
|---|---|---|
| Training | Learns coefficients | Stores examples |
| Prediction | Fast equation calculation | Neighbor search can be slower |
| Boundary shape | Usually global and simple | Local and flexible |
| Feature scaling | Helpful for optimization/regularization | Essential for meaningful distances |
| Interpretability | Coefficients can be inspected | Explain using nearest examples |
Training cost vs prediction cost
KNN is lazy: training is almost just storing data, but prediction requires distance calculations.
| Stage | What happens | Cost intuition |
|---|---|---|
| Training | Store training examples | Cheap |
| Prediction | Compare new point with many training examples | Can be expensive |
where \(n\) is number of training rows and \(p\) is number of features.