Concept 4

Regression evaluation metrics

Metrics translate model mistakes into numbers. Students should learn both the formula and the sentence they would say to a business user.

Error metrics

MetricFormulaPlain-English meaning
MAE\(\frac{1}{n}\sum_{i=1}^{n}|y_i - \hat{y}_i|\)On average, how far are predictions from actual values?
MSE\(\frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2\)Average squared error, useful for optimization but harder to interpret.
RMSE\(\sqrt{\mathrm{MSE}}\)Typical prediction error in the same unit as y.

Metric choice

If a house-price model has RMSE = 5 lakh, is that good or bad?

Hint: it depends on the price range. 5 lakh error is different for 20 lakh homes versus 5 crore homes.

R2 score

R2 measures how much variation in the target is explained by the model compared with a baseline that always predicts the mean.

\[R^2 = 1 - \frac{SS_{res}}{SS_{tot}}\]
  • SS_res: unexplained error left by the model.
  • SS_tot: total variation around the mean.
  • R2 = 1 means perfect prediction.
  • R2 = 0 means no better than predicting the average.
  • R2 can be negative when the model is worse than the average baseline.

Baseline challenge

If our model gets R2 = 0, what simple prediction is it roughly matching?

Takeaway: R2 = 0 roughly matches always predicting the average target value.

mean baseline model

R2 asks: how much better is the regression line than the average line?

What does SS mean in R2?

In R2, SS means Sum of Squares. The formula compares two kinds of squared error: error before using the model and error after using the model.

Total Sum of Squares

\[SS_{tot} = \sum_{i=1}^{n}(y_i - \bar{y})^2\]

This is the total variation in the target around its average. It is the baseline error from a very simple model that always predicts \(\bar{y}\).

Residual Sum of Squares

\[SS_{res} = \sum_{i=1}^{n}(y_i - \hat{y}_i)^2\]

This is the leftover error after using the regression model. Lower residual error means the model explains more of the variation.

\[R^2 = 1 - \frac{SS_{res}}{SS_{tot}}\]

So R2 asks: how much better is the regression model compared to simply predicting the average every time?

Baseline: predict the average

average line

\(SS_{tot}\): how much error exists if we only use the mean.

Regression: predict with a line

model line

\(SS_{res}\): how much error remains after the model.

Example: if \(SS_{tot}=1000\) and \(SS_{res}=200\), then \(R^2 = 1 - \frac{200}{1000} = 0.8\). The model explains 80% of the target variation.

Adjusted R2

R2 usually increases when more features are added, even if the feature is not useful. Adjusted R2 penalizes unnecessary features.

\[\mathrm{Adjusted}\ R^2 = 1 - \frac{(1 - R^2)(n - 1)}{n - p - 1}\]

Here, n is number of rows and p is number of predictors.

For beginners, teach adjusted R2 after R2 is comfortable. Do not let it distract from MAE, RMSE, and test-set evaluation.

Feature choice

Should we add every column we have into a regression model?

Takeaway: extra columns can add noise, overfitting, and interpretation problems; adjusted R2 penalizes unnecessary predictors.

Previous: Session 1 Next: Multiple regression