Validation in model training
Validation is the habit that protects us from being fooled by training error. It gives a practical way to choose polynomial degree, lambda, and model settings.
Why validation?
Training error tells us how well the model fits known data. Validation error estimates how well the model may perform on unseen data.
Model choice
If degree 10 has lower training error than degree 3, but degree 3 has lower validation error, which model should we prefer?
Prefer degree 3, because the goal is generalization.
Train, validation, and test split
The test set should be touched only after model choices are finalized.
| Split | Used for | Do we learn from it? |
|---|---|---|
| Training set | Fit model coefficients. | Yes |
| Validation set | Choose degree, lambda, features, and hyperparameters. | Indirectly |
| Test set | Estimate final performance after decisions are complete. | No |
Which polynomial degrees should we experiment with?
Polynomial degree is a hyperparameter. The model learns coefficients such as \(\beta_0,\beta_1,\beta_2\), but we decide which degrees are worth trying.
A good practical approach is to start simple, increase complexity gradually, and let validation performance decide when extra flexibility stops helping.
Start with the story of the data
If the relationship bends once, try degree 2. If it may bend more than once, try degree 3 or 4. Avoid jumping directly to degree 20.
Respect data size and noise
Small or noisy datasets should use a smaller search range. High-degree polynomials can memorize random fluctuations.
Use validation to decide
The final choice should come from validation or cross-validation error, not from training error.
| Problem | Reasonable degrees to try | Why this range? |
|---|---|---|
| Study hours vs marks | \(\{1,2,3\}\) | Scores may rise quickly first and then saturate. A low-degree curve is usually enough. |
| Experience vs salary | \(\{1,2,3,4\}\) | Salary growth may accelerate early and flatten later, but very high degrees are hard to explain. |
| House size vs price | \(\{1,2,3\}\) | There may be diminishing returns for very large houses. Extrapolation risk is high. |
| Ad spend vs sales | \(\{1,2,3,4,5\}\) | Sales can rise, flatten, or show diminishing returns. Validation should choose the useful complexity. |
| Very small dataset | \(\{1,2\}\) | Few points cannot reliably support a highly flexible curve. |
Search range decision
For 20 noisy data points, would you try degrees 1 to 5 or degrees 1 to 25 first?
Start with 1 to 5. A degree-25 polynomial has too much freedom for such a small noisy dataset.
Solid example: choosing degree by validation error
Suppose we try several polynomial degrees for ad spend vs sales. Training error keeps improving as degree increases, but validation error tells a different story.
| Degree | Training RMSE | Validation RMSE | Decision |
|---|---|---|---|
| 1 | 8.2 | 8.6 | Too simple. Likely underfitting. |
| 2 | 5.9 | 6.2 | Much better. Captures curvature. |
| 3 | 5.4 | 5.8 | Best validation score. |
| 4 | 5.1 | 6.1 | Training improves, validation worsens. |
| 5 | 4.7 | 7.4 | Overfitting risk is increasing. |
The lesson: degree 5 looks best on training data, but degree 3 is better for future data because it has the lowest validation error.
Degree 3 is selected because the validation RMSE reaches its lowest point there, even though training RMSE keeps improving at higher degrees.
Validation curve
A validation curve compares model complexity with training and validation error.
Training error may keep falling, but validation error usually reveals the overfitting point.
Cross-validation
Cross-validation makes validation more reliable by rotating which part of the data is used for validation.
Instead of trusting one validation split, K-fold cross-validation averages performance across \(K\) different validation folds.
Each fold gets a turn as validation data. The final CV score averages the validation errors.
Implementation idea for K-fold cross-validation
The practical algorithm is simple:
Data leakage check
Should polynomial feature creation and scaling be fit before or inside each cross-validation fold?
They should be fit inside each fold using only the training part. Pipelines help avoid leakage.
What cross-validation helps us choose
Polynomial regression has choices that are not learned directly as coefficients. These choices are called hyperparameters.
| Hyperparameter | Controls | If too low | If too high |
|---|---|---|---|
| Polynomial degree | Curve flexibility | Underfitting, high bias | Overfitting, high variance |
| Ridge/Lasso lambda | Regularization strength | Weak control of overfitting | Too much shrinkage, underfitting |
After choosing the degree and lambda, a common workflow is to refit the selected model on the available training data, then use the untouched test set once for the final performance estimate.