Assumptions and practical risks
Assumptions explain when linear regression is trustworthy and when its output needs caution.
Core assumptions
| Assumption | Meaning | How to explain to students |
|---|---|---|
| Linearity | The average relationship between x and y is roughly linear. | A straight line should be a reasonable simplification. |
| Independent errors | One residual should not depend on another residual. | Time series data often violates this because nearby observations are related. |
| Constant variance | Residual spread should be similar across prediction levels. | The model should not be accurate for small values and wildly noisy for large values. |
| Normal residuals | Residuals should be roughly bell-shaped for inference. | Mainly important for confidence intervals and hypothesis tests, not just prediction. |
| No severe multicollinearity | Input features should not duplicate each other too strongly. | If TV and radio budgets were almost identical, coefficients become hard to trust. |
Diagnostic question
If the residual plot shows a clear curve, which assumption is probably violated?
Expected answer: linearity. The model may need new features, transformations, or a different algorithm.
Multicollinearity
Multicollinearity happens when two or more input features are strongly related. It may not destroy prediction accuracy, but it can make coefficient interpretation unstable.
- Use a correlation heatmap as the first check.
- Use VIF as an advanced diagnostic.
- Remove or combine redundant features if interpretation matters.
Student prediction
If two features give almost the same information, should both coefficients be trusted separately?
Expected answer: not always. Prediction may still work, but individual coefficient interpretation becomes unstable.
| TV | Radio | News | |
|---|---|---|---|
| TV | 1.00 | 0.12 | 0.06 |
| Radio | 0.12 | 1.00 | 0.88 |
| News | 0.06 | 0.88 | 1.00 |
High feature-feature correlation is the warning sign. Feature-target correlation is different and often useful.
Feature scaling
Standard OLS linear regression does not require scaling for prediction because the solver can handle different numeric ranges. Scaling becomes important when:
- you train using gradient descent or SGD,
- you use regularization such as Ridge, Lasso, or ElasticNet,
- you compare coefficient magnitudes across features,
- you use distance-based models later, such as KNN.
Scaling check
Which feature has a bigger number: house area in square feet or number of bedrooms? Does bigger number automatically mean bigger importance?
Expected answer: no. Units and scale affect magnitude; coefficient interpretation must consider feature scale.
Regularization teaser
Regularization adds a penalty for large coefficients. It helps reduce overfitting and stabilize models with many features.