Concept 2

Exploring data, simulation, and generalization

Polynomial regression is easiest to understand when students can see several models trying to fit the same curved dataset.

Build a simulated problem

Suppose the true relationship is curved, but the observations contain noise.

\[y = f(x) + \epsilon\]
\[f(x)=3+2x-0.7x^2\qquad \epsilon\sim \mathcal{N}(0,\sigma^2)\]

The model never sees the clean function \(f(x)\). It only sees noisy samples. That is why a model can accidentally fit noise instead of signal.

Signal vs noise

When a point is far away from the smooth pattern, should the model bend itself to pass through that point?

This question separates fitting the training data from learning the underlying relationship.

Compare model degrees

Degree 1: too rigid Degree 2: good shape Degree 5: flexible Degree 12: memorizes Gray curve: true relationship. Colored curve: learned model. Black points: noisy training data.

The goal is not to pass through every point. The goal is to recover a pattern that works on future data.

Generalization

Generalization means the model performs well on new data, not only on the data it saw during training.

\[\text{good ML model} \neq \text{lowest training error}\]
\[\text{good ML model} = \text{low future error}\]

A high-degree polynomial can reduce training error almost to zero, but its predictions between and beyond training points may become unreasonable.

Training error

Error measured on the data used to learn the coefficients.

Training error usually decreases as model complexity increases.

Test or validation error

Error measured on data not used for fitting.

This is a better estimate of how the model will behave in the real world.

Generalization check

If a degree-15 model has almost zero training error, should we trust it more than a degree-3 model?

Answer depends on validation performance. Without validation, low training error can be misleading.

Extrapolation risk

Extrapolation means using the fitted curve outside the range of observed training data. Polynomial curves can behave very strangely outside the data range, especially at higher degrees.

training data range outside range outside range

Inside the observed range, the curve may look reasonable. Outside it, high-degree terms can dominate and produce unrealistic predictions.

Prediction boundary

If the training data has house sizes from 500 to 3000 sq ft, should we confidently predict prices for a 20,000 sq ft mansion using the same polynomial curve?

Be careful. That is extrapolation, and polynomial models can be unreliable outside the observed range.

Polynomial feature matrix

For one original feature \(x\), a degree-3 polynomial model creates this design matrix:

\[ X_{poly}= \begin{bmatrix} 1 & x_1 & x_1^2 & x_1^3 \\ 1 & x_2 & x_2^2 & x_2^3 \\ 1 & x_3 & x_3^2 & x_3^3 \\ \vdots & \vdots & \vdots & \vdots \\ 1 & x_n & x_n^2 & x_n^3 \end{bmatrix} \]

After this transformation, the model is fit using the same linear regression machinery:

\[\hat{y}=X_{poly}\beta\]

This matrix is a polynomial design matrix. For one input feature, the columns follow the pattern \(1,x,x^2,\ldots,x^d\). The same idea extends to multiple features by adding powers and interaction terms.

\[x_1x_2,\quad x_1^2,\quad x_2^2,\quad x_1^2x_2,\quad x_1x_2^2\]

Interaction terms let the effect of one feature depend on another feature.

Polynomial features can become numerically large. In practice, scaling features before regularized polynomial regression is usually important.
Previous: Intro Next: Bias-Variance