Controlling overfitting with regularization
Polynomial features can make a model powerful, but also unstable. Regularization adds a cost for unnecessary coefficient size, encouraging simpler curves.
Why regularization?
High-degree polynomial models often use large positive and negative coefficients to create sharp bends. Those bends may fit training noise.
Regularization keeps the model from becoming too dependent on extreme coefficient values.
Complexity check
If two models have similar validation error, but one has much smaller coefficients, which model feels safer?
The simpler model is often preferred because it is usually more stable.
Ridge regularization
Ridge regression adds an \(L2\) penalty, which penalizes squared coefficient size.
Large coefficients become expensive, so Ridge shrinks them toward zero. It usually does not make coefficients exactly zero.
Lasso regularization
Lasso regression adds an \(L1\) penalty, which penalizes absolute coefficient size.
Lasso can shrink some coefficients exactly to zero, so it can act like feature selection.
| Method | Penalty | Main effect |
|---|---|---|
| Ridge | \(\lambda\sum\beta_j^2\) | Shrinks coefficients smoothly. |
| Lasso | \(\lambda\sum|\beta_j|\) | Can set some coefficients to zero. |
| Elastic Net | Mix of \(L1\) and \(L2\) | Combines shrinkage and feature selection behavior. |
Shrinkage parameter: lambda
The parameter \(\lambda\) controls how strongly we punish large coefficients.
Lambda is not “more is always better.” It is another model choice that should be selected using validation or cross-validation.
Coefficient shrinkage simulation
Imagine a degree-8 polynomial model. Without regularization, high-degree terms may receive large coefficients to chase small noise patterns.
Regularization discourages extreme coefficient values, which often produces smoother polynomial curves.
Regularized polynomial workflow
Final session check
What are the two main knobs we used to control polynomial regression?
Polynomial degree controls flexibility. Lambda controls coefficient shrinkage.
Feature engineering beyond powers
Polynomial regression is one feature engineering strategy, but the larger idea is to create features whose shape matches the problem.
| Feature idea | Example | When it may help |
|---|---|---|
| Power feature | \(size^2\), \(experience^3\) | Relationship bends or changes direction. |
| Square-root feature | \(\sqrt{size}\) | Effect grows quickly first, then slows down. |
| Interaction feature | \(x_1x_2\) | One feature changes the effect of another feature. |
Feature shape
If house price rises with size but the increase slows for very large houses, which feature shape might be more natural: \(size^2\) or \(\sqrt{size}\)?
A square-root style feature may better represent diminishing returns.
Takeaway
Polynomial regression teaches a general ML lesson:
Curves are useful, but uncontrolled curves can memorize. The complete workflow is to create flexible features, validate carefully, and regularize when needed.