Regression evaluation metrics
Metrics translate model mistakes into numbers. Students should learn both the formula and the sentence they would say to a business user.
Error metrics
| Metric | Formula | Best classroom explanation |
|---|---|---|
| MAE | \(\frac{1}{n}\sum_{i=1}^{n}|y_i - \hat{y}_i|\) | On average, how far are predictions from actual values? |
| MSE | \(\frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2\) | Average squared error, useful for optimization but harder to interpret. |
| RMSE | \(\sqrt{\mathrm{MSE}}\) | Typical prediction error in the same unit as y. |
Metric choice question
If a house-price model has RMSE = 5 lakh, is that good or bad?
Expected answer: it depends on the price range. 5 lakh error is different for 20 lakh homes versus 5 crore homes.
R2 score
R2 measures how much variation in the target is explained by the model compared with a baseline that always predicts the mean.
- SS_res: unexplained error left by the model.
- SS_tot: total variation around the mean.
- R2 = 1 means perfect prediction.
- R2 = 0 means no better than predicting the average.
- R2 can be negative when the model is worse than the average baseline.
Baseline challenge
If our model gets R2 = 0, what simple prediction is it roughly matching?
Expected answer: always predicting the average target value.
R2 asks: how much better is the regression line than the average line?
Adjusted R2
R2 usually increases when more features are added, even if the feature is not useful. Adjusted R2 penalizes unnecessary features.
Here, n is number of rows and p is number of predictors.
Feature debate
Should we add every column we have into a regression model?
Let students argue yes/no. Then introduce noise, overfitting, interpretability, and adjusted R2.