Concept 6

Assumptions and practical risks

Assumptions explain when linear regression is trustworthy and when its output needs caution.

Core assumptions

AssumptionMeaningHow to explain to students
LinearityThe average relationship between x and y is roughly linear.A straight line should be a reasonable simplification.
Independent errorsOne residual should not depend on another residual.Time series data often violates this because nearby observations are related.
Constant varianceResidual spread should be similar across prediction levels.The model should not be accurate for small values and wildly noisy for large values.
Normal residualsResiduals should be roughly bell-shaped for inference.Mainly important for confidence intervals and hypothesis tests, not just prediction.
No severe multicollinearityInput features should not duplicate each other too strongly.If TV and radio budgets were almost identical, coefficients become hard to trust.

Diagnostic question

If the residual plot shows a clear curve, which assumption is probably violated?

Expected answer: linearity. The model may need new features, transformations, or a different algorithm.

Multicollinearity

Multicollinearity happens when two or more input features are strongly related. It may not destroy prediction accuracy, but it can make coefficient interpretation unstable.

\[\mathrm{corr}(x_1, x_2) \approx 1\]\[\text{coefficient interpretation becomes unreliable}\]
  • Use a correlation heatmap as the first check.
  • Use VIF as an advanced diagnostic.
  • Remove or combine redundant features if interpretation matters.

Student prediction

If two features give almost the same information, should both coefficients be trusted separately?

Expected answer: not always. Prediction may still work, but individual coefficient interpretation becomes unstable.

TVRadioNews
TV1.000.120.06
Radio0.121.000.88
News0.060.881.00

High feature-feature correlation is the warning sign. Feature-target correlation is different and often useful.

Feature scaling

Standard OLS linear regression does not require scaling for prediction because the solver can handle different numeric ranges. Scaling becomes important when:

Scaling check

Which feature has a bigger number: house area in square feet or number of bedrooms? Does bigger number automatically mean bigger importance?

Expected answer: no. Units and scale affect magnitude; coefficient interpretation must consider feature scale.

Regularization teaser

Regularization adds a penalty for large coefficients. It helps reduce overfitting and stabilize models with many features.

\[\text{Ridge: } \min\left(\mathrm{MSE} + \lambda\sum_{j=1}^{p}\beta_j^2\right)\]\[\text{Lasso: } \min\left(\mathrm{MSE} + \lambda\sum_{j=1}^{p}|\beta_j|\right)\]
Good stopping point for linear regression: tell students regularization is the next tool after they understand ordinary linear regression.
Previous: Multiple regression Back to overview