Concept 6

Assumptions and practical risks

Assumptions explain when linear regression is trustworthy and when its output needs caution.

Core assumptions

AssumptionMeaningPlain-English meaning
LinearityThe average relationship between x and y is roughly linear.A straight line should be a reasonable simplification.
Independent errorsOne residual should not depend on another residual.Time series data often violates this because nearby observations are related.
Constant varianceResidual spread should be similar across prediction levels.The model should not be accurate for small values and wildly noisy for large values.
Normal residualsResiduals should be roughly bell-shaped for inference.Mainly important for confidence intervals and hypothesis tests, not just prediction.
No severe multicollinearityInput features should not duplicate each other too strongly.If TV and radio budgets were almost identical, coefficients become hard to trust.

Diagnostic check

If the residual plot shows a clear curve, which assumption is probably violated?

Takeaway: this points to a linearity problem. The model may need new features, transformations, or a different algorithm.

Assumptions visual gallery

These quick visual checks help diagnose when linear regression is behaving well and when the model may need feature engineering, transformations, or another algorithm.

Multicollinearity

Multicollinearity happens when two or more input features are strongly related. It may not destroy prediction accuracy, but it can make coefficient interpretation unstable.

\[\mathrm{corr}(x_1, x_2) \approx 1\]\[\text{coefficient interpretation becomes unreliable}\]
  • Use a correlation heatmap as the first check.
  • Use VIF as an advanced diagnostic.
  • Remove or combine redundant features if interpretation matters.

Prediction check

If two features give almost the same information, should both coefficients be trusted separately?

Takeaway: not always. Prediction may still work, but individual coefficient interpretation becomes unstable.

TVRadioNews
TV1.000.120.06
Radio0.121.000.88
News0.060.881.00

High feature-feature correlation is the warning sign. Feature-target correlation is different and often useful.

Feature scaling

Standard OLS linear regression does not require scaling for prediction because the solver can handle different numeric ranges. Scaling becomes important when:

Scale check

Which feature has a bigger number: house area in square feet or number of bedrooms? Does bigger number automatically mean bigger importance?

Takeaway: no. Units and scale affect magnitude; coefficient interpretation must consider feature scale.

Regularization teaser

Regularization adds a penalty for large coefficients. It helps reduce overfitting and stabilize models with many features.

\[\text{Ridge: } \min\left(\mathrm{MSE} + \lambda\sum_{j=1}^{p}\beta_j^2\right)\]\[\text{Lasso: } \min\left(\mathrm{MSE} + \lambda\sum_{j=1}^{p}|\beta_j|\right)\]
Next step: regularization is the natural tool after ordinary linear regression.
Previous: Multiple regression Back to overview