Errors and cost functions
A model learns by comparing predictions with actual values. The math of linear regression is mostly the math of measuring mistakes.
Residual
A residual is the vertical distance between the actual point and the predicted line. Positive means the model predicted too low. Negative means the model predicted too high.
Why not sum raw errors?
A model with two big mistakes could look perfect if positive and negative errors cancel. That is why we use absolute values or squares.
Try this
Model A errors are +10 and -10. Model B errors are +1 and -1. If we sum raw errors, both look like 0. Are both equally good?
Takeaway: raw errors can cancel, so MAE and MSE give a more honest error summary.
A best-fit line tries to make these vertical distances small overall.
Main cost formulas
| Name | Formula | What to say in class |
|---|---|---|
| RSS | \(\sum_{i=1}^{n}(y_i - \hat{y}_i)^2\) | Total squared error. Bigger datasets naturally produce bigger RSS. |
| MSE | \(\frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2\) | Average squared error. Common training objective. |
| RMSE | \(\sqrt{\mathrm{MSE}}\) | Error in the same unit as the target. Easier to report. |
| MAE | \(\frac{1}{n}\sum_{i=1}^{n}|y_i - \hat{y}_i|\) | Average absolute error. Easy for business users. |
Choose a metric
For predicting delivery time, which metric would you explain to a business manager: MAE, MSE, or RMSE?
Hint: MAE/RMSE are easier because they are in minutes; MSE is useful internally but less intuitive.
Outlier demo: why RMSE reacts strongly
MAE and RMSE both measure prediction error, but they react differently when one prediction is very bad. RMSE squares errors before averaging, so one large error becomes much more visible.
Normal errors
One large outlier
Checkpoint
Which metric changed more dramatically after adding one large outlier: MAE or RMSE?
Key idea: RMSE is useful when large mistakes are especially costly, but it is also more sensitive to outliers.
Mini example
| Actual | Predicted | Error | Squared error | Absolute error |
|---|---|---|---|---|
| 100 | 90 | 10 | 100 | 10 |
| 200 | 220 | -20 | 400 | 20 |
| 300 | 280 | 20 | 400 | 20 |
RSS = 900, MSE = 300, RMSE = 17.32, MAE = 16.67.
Quick calculation
Change the last prediction from 280 to 250. What happens to MAE and RMSE?
This helps students see that RMSE reacts more strongly to large mistakes.