Polynomial regression: flexibility, generalization, and control.
This session begins with a simple question: what if the relationship is clearly not a straight line? From there, students learn polynomial features, model complexity, bias-variance tradeoff, validation, cross-validation, and regularization.
Teaching story
Opening question
Can every useful prediction problem be solved with a straight line?
Use examples like house size vs price, experience vs salary, medicine dosage vs response, and ad spend vs sales. Many relationships increase, slow down, saturate, or bend.
Session roadmap
Polynomial regression
What it is, why it is needed, how it handles non-linear data, and why it is still a linear model in the coefficients.
2Simulation and generalization
Create curved data, compare degree 1, 2, 5, and 12 models, and separate training fit from future performance.
3Bias and variance
Understand underfitting, overfitting, the tradeoff, and why model complexity can both help and hurt.
4Validation and cross-validation
Train, validation, test splits; choosing polynomial degrees to try; K-fold cross-validation; and model selection using validation error.
5Regularization
Ridge, Lasso, Elastic Net, lambda, coefficient shrinkage, and simulations for controlling overfitting.
Reference anchors for the session
These are the big ideas the session will keep returning to:
Curve fitting is broader than regression
A fitted curve can summarize a relationship, interpolate between known observations, or support prediction. Regression focuses on learning from noisy data rather than forcing an exact pass through every point.
Polynomial regression is multiple regression
After creating \(x,x^2,x^3,\ldots\), the model becomes ordinary linear regression on transformed columns.
Complexity needs control
Higher degree can reduce bias, but it can also increase variance. Validation and regularization decide how much flexibility is useful.
Suggested 3 hour pacing
| Time | Focus | Class mode |
|---|---|---|
| 0:00-0:35 | Motivation, polynomial equation, non-linear data | Visual explanation and questions |
| 0:35-1:05 | Data simulation and model degrees | Compare fits and discuss generalization |
| 1:05-1:45 | Bias, variance, underfitting, overfitting | Concept visuals and quick checks |
| 1:45-2:20 | Validation, model selection, K-fold CV | Workflow and algorithm walkthrough |
| 2:20-3:00 | Regularization, lambda, Ridge/Lasso simulation | Coefficient shrinkage story and wrap-up |
Useful reading links
| Link | Use in class |
|---|---|
| Features and Polynomial Regression notes | Feature scaling, square-root features, and feature engineering examples. |
| Curve fitting | Context for fitting, interpolation, smoothing, least-squares approximation, and extrapolation risk. |
| Polynomial regression | Definition, matrix form, and why the method is linear in estimated coefficients. |
| Bias-variance tradeoff | Interactive intuition for repeated model fits, systematic error, variance, and the U-shaped error curve. |