Regression evaluation metrics
Metrics translate model mistakes into numbers. Students should learn both the formula and the sentence they would say to a business user.
Error metrics
| Metric | Formula | Plain-English meaning |
|---|---|---|
| MAE | \(\frac{1}{n}\sum_{i=1}^{n}|y_i - \hat{y}_i|\) | On average, how far are predictions from actual values? |
| MSE | \(\frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2\) | Average squared error, useful for optimization but harder to interpret. |
| RMSE | \(\sqrt{\mathrm{MSE}}\) | Typical prediction error in the same unit as y. |
Metric choice
If a house-price model has RMSE = 5 lakh, is that good or bad?
Hint: it depends on the price range. 5 lakh error is different for 20 lakh homes versus 5 crore homes.
R2 score
R2 measures how much variation in the target is explained by the model compared with a baseline that always predicts the mean.
- SS_res: unexplained error left by the model.
- SS_tot: total variation around the mean.
- R2 = 1 means perfect prediction.
- R2 = 0 means no better than predicting the average.
- R2 can be negative when the model is worse than the average baseline.
Baseline challenge
If our model gets R2 = 0, what simple prediction is it roughly matching?
Takeaway: R2 = 0 roughly matches always predicting the average target value.
R2 asks: how much better is the regression line than the average line?
What does SS mean in R2?
In R2, SS means Sum of Squares. The formula compares two kinds of squared error: error before using the model and error after using the model.
Total Sum of Squares
This is the total variation in the target around its average. It is the baseline error from a very simple model that always predicts \(\bar{y}\).
Residual Sum of Squares
This is the leftover error after using the regression model. Lower residual error means the model explains more of the variation.
So R2 asks: how much better is the regression model compared to simply predicting the average every time?
Baseline: predict the average
\(SS_{tot}\): how much error exists if we only use the mean.
Regression: predict with a line
\(SS_{res}\): how much error remains after the model.
Adjusted R2
R2 usually increases when more features are added, even if the feature is not useful. Adjusted R2 penalizes unnecessary features.
Here, n is number of rows and p is number of predictors.
Feature choice
Should we add every column we have into a regression model?
Takeaway: extra columns can add noise, overfitting, and interpretation problems; adjusted R2 penalizes unnecessary predictors.