Concept 4

Validation in model training

Validation is the habit that protects us from being fooled by training error. It gives a practical way to choose polynomial degree, lambda, and model settings.

Why validation?

Training error tells us how well the model fits known data. Validation error estimates how well the model may perform on unseen data.

\[\text{choose model using validation error, not training error}\]

Model choice

If degree 10 has lower training error than degree 3, but degree 3 has lower validation error, which model should we prefer?

Prefer degree 3, because the goal is generalization.

Train, validation, and test split

train validation test fit coefficients select degree/lambda final unbiased check

The test set should be touched only after model choices are finalized.

SplitUsed forDo we learn from it?
Training setFit model coefficients.Yes
Validation setChoose degree, lambda, features, and hyperparameters.Indirectly
Test setEstimate final performance after decisions are complete.No

Which polynomial degrees should we experiment with?

Polynomial degree is a hyperparameter. The model learns coefficients such as \(\beta_0,\beta_1,\beta_2\), but we decide which degrees are worth trying.

\[ d \in \{1,2,3,4,5\} \]

A good practical approach is to start simple, increase complexity gradually, and let validation performance decide when extra flexibility stops helping.

Start with the story of the data

If the relationship bends once, try degree 2. If it may bend more than once, try degree 3 or 4. Avoid jumping directly to degree 20.

Respect data size and noise

Small or noisy datasets should use a smaller search range. High-degree polynomials can memorize random fluctuations.

Use validation to decide

The final choice should come from validation or cross-validation error, not from training error.

ProblemReasonable degrees to tryWhy this range?
Study hours vs marks\(\{1,2,3\}\)Scores may rise quickly first and then saturate. A low-degree curve is usually enough.
Experience vs salary\(\{1,2,3,4\}\)Salary growth may accelerate early and flatten later, but very high degrees are hard to explain.
House size vs price\(\{1,2,3\}\)There may be diminishing returns for very large houses. Extrapolation risk is high.
Ad spend vs sales\(\{1,2,3,4,5\}\)Sales can rise, flatten, or show diminishing returns. Validation should choose the useful complexity.
Very small dataset\(\{1,2\}\)Few points cannot reliably support a highly flexible curve.
A useful default for many classroom examples is \(d \in \{1,2,3,4,5\}\). If degree 5 is best and validation error is still decreasing, expand carefully. If validation error starts increasing, stop.

Search range decision

For 20 noisy data points, would you try degrees 1 to 5 or degrees 1 to 25 first?

Start with 1 to 5. A degree-25 polynomial has too much freedom for such a small noisy dataset.

Solid example: choosing degree by validation error

Suppose we try several polynomial degrees for ad spend vs sales. Training error keeps improving as degree increases, but validation error tells a different story.

DegreeTraining RMSEValidation RMSEDecision
18.28.6Too simple. Likely underfitting.
25.96.2Much better. Captures curvature.
35.45.8Best validation score.
45.16.1Training improves, validation worsens.
54.77.4Overfitting risk is increasing.
\[ \text{selected degree}=3 \]

The lesson: degree 5 looks best on training data, but degree 3 is better for future data because it has the lowest validation error.

degree 3 1 2 4 5 lowest validation error training RMSE validation RMSE polynomial degree error

Degree 3 is selected because the validation RMSE reaches its lowest point there, even though training RMSE keeps improving at higher degrees.

Validation curve

A validation curve compares model complexity with training and validation error.

polynomial degree error training error decreases validation error is U-shaped best degree

Training error may keep falling, but validation error usually reveals the overfitting point.

Cross-validation

Cross-validation makes validation more reliable by rotating which part of the data is used for validation.

\[\text{CV score}=\frac{1}{K}\sum_{k=1}^{K}\text{ValidationError}_k\]

Instead of trusting one validation split, K-fold cross-validation averages performance across \(K\) different validation folds.

Fold 1 val Fold 2 val Fold 3 val Fold 4 val Fold 5 val

Each fold gets a turn as validation data. The final CV score averages the validation errors.

Implementation idea for K-fold cross-validation

The practical algorithm is simple:

Split data into K folds Hold out one fold Train on remaining folds Measure validation error Average all errors
\[ \text{for degree in } \{1,2,3,\ldots\}: \quad \text{score degree using K-fold CV} \]
\[ \text{selected degree}=\arg\min_d \text{CVError}(d) \]
Use cross-validation for model selection. Keep a final test set separate if you need a final trustworthy performance number.

Data leakage check

Should polynomial feature creation and scaling be fit before or inside each cross-validation fold?

They should be fit inside each fold using only the training part. Pipelines help avoid leakage.

What cross-validation helps us choose

Polynomial regression has choices that are not learned directly as coefficients. These choices are called hyperparameters.

HyperparameterControlsIf too lowIf too high
Polynomial degreeCurve flexibilityUnderfitting, high biasOverfitting, high variance
Ridge/Lasso lambdaRegularization strengthWeak control of overfittingToo much shrinkage, underfitting
\[ \text{choose hyperparameters by validation performance, then refit the final model} \]

After choosing the degree and lambda, a common workflow is to refit the selected model on the available training data, then use the untouched test set once for the final performance estimate.

Previous: Bias-Variance Next: Regularization