Concept 2

Errors and cost functions

A model learns by comparing predictions with actual values. The math of linear regression is mostly the math of measuring mistakes.

Residual

\[e_i = y_i - \hat{y}_i\]

A residual is the vertical distance between the actual point and the predicted line. Positive means the model predicted too low. Negative means the model predicted too high.

Why not sum raw errors?

\[(+10) + (-10) = 0\]

A model with two big mistakes could look perfect if positive and negative errors cancel. That is why we use absolute values or squares.

Board question

Model A errors are +10 and -10. Model B errors are +1 and -1. If we sum raw errors, both look like 0. Are both equally good?

Let students answer first, then use this to motivate MAE and MSE.

residual

A best-fit line tries to make these vertical distances small overall.

Main cost formulas

NameFormulaWhat to say in class
RSS\(\sum_{i=1}^{n}(y_i - \hat{y}_i)^2\)Total squared error. Bigger datasets naturally produce bigger RSS.
MSE\(\frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2\)Average squared error. Common training objective.
RMSE\(\sqrt{\mathrm{MSE}}\)Error in the same unit as the target. Easier to report.
MAE\(\frac{1}{n}\sum_{i=1}^{n}|y_i - \hat{y}_i|\)Average absolute error. Easy for business users.
MSE punishes large errors more strongly because errors are squared. This is useful, but it also makes MSE sensitive to outliers.

Poll

For predicting delivery time, which metric would you explain to a business manager: MAE, MSE, or RMSE?

Expected direction: MAE/RMSE are easier because they are in minutes; MSE is useful internally but less intuitive.

Mini example

ActualPredictedErrorSquared errorAbsolute error
100901010010
200220-2040020
3002802040020

RSS = 900, MSE = 300, RMSE = 17.32, MAE = 16.67.

2-minute calculation

Change the last prediction from 280 to 250. What happens to MAE and RMSE?

This helps students see that RMSE reacts more strongly to large mistakes.

Previous: Line Next: Learning