Concept 3

Bias, variance, underfitting, and overfitting

Polynomial regression is a perfect place to teach the central ML tension: simple models miss patterns, while very flexible models chase noise.

What is bias?

Bias is error caused by a model being too simple or making assumptions that are too restrictive.

\[\text{high bias} \Rightarrow \text{model cannot capture the true pattern}\]

A degree-1 line trying to fit a clear curve usually has high bias.

A useful mental model: imagine training the same type of model on many different datasets sampled from the same real-world process. Bias measures how far the average prediction is from the true pattern.

Bias is not about social or ethical bias here. In this context, bias means systematic prediction error from an overly simple model.

What is variance?

Variance is error caused by a model being too sensitive to the particular training sample.

\[\text{high variance} \Rightarrow \text{model changes a lot when training data changes}\]

A very high-degree polynomial can bend around small random noise. If we collect a slightly different dataset, the fitted curve may look very different.

Using the same repeated-datasets mental model, variance measures how spread out the different model predictions are at the same input value.

Sensitivity check

If removing one training point changes the curve dramatically, what does that suggest?

It suggests high variance. The model is too sensitive to the exact sample.

Repeated samples intuition

Bias and variance are easier to understand when we imagine several possible training datasets from the same source. Each dataset gives a slightly different fitted model.

High bias similar wrong models Good balance similar useful models High variance different unstable models

Bias asks whether the average model is wrong. Variance asks whether the fitted model changes too much across samples.

Underfitting and overfitting

Underfitting high bias, low variance Good balance lower total error Overfitting low bias, high variance

Underfitting misses the pattern. Overfitting follows noise. A useful model captures the signal and ignores random fluctuations.

Bias-variance decomposition

For squared error, prediction error at a fixed input value \(x\) can be understood as three parts:

\[ \mathbb{E}\left[(y-\hat{f}(x))^2\right] = \text{Bias}^2 + \text{Variance} + \text{Irreducible noise} \]

The expectation symbol \(\mathbb{E}\) simply means average. So the left side asks:

\[ \text{If we trained this model many times on many slightly different datasets, how wrong would it be on average?} \]

The formula is not saying one single prediction error has three visible parts. It is saying that if we repeat the whole training process many times and predict at the same \(x\), the model's average squared error comes from three sources.

\[ \text{Average prediction error} = \text{error from simplicity} + \text{error from instability} + \text{unavoidable randomness} \]
PartMeaningPolynomial regression example
Bias squaredError from wrong or too-simple model assumptions.Degree 1 line for a curved relationship.
VarianceError from sensitivity to training data.Degree 15 curve changes wildly with small data changes.
Irreducible noiseRandomness no model can remove.Measurement error, missing factors, natural randomness.

Bias

The model is consistently wrong because its shape is too simple. A straight line trying to fit a curve will keep missing the curve even if we collect a new sample.

Variance

The model changes too much depending on the exact training data. A very wiggly polynomial may look completely different if a few points change.

Irreducible noise

Some error remains because the world is partly unpredictable or because important information is missing from the dataset.

Formula translation

Which part of the formula increases when a model starts memorizing small random ups and downs in the training data?

Variance increases because the model has become too sensitive to the sample.

The goal is not to make bias zero or variance zero. The goal is to minimize total expected error.

Why is there a tradeoff?

As degree increases, the model becomes more flexible. Flexibility reduces bias because the model can express more shapes. But flexibility can increase variance because the model can also fit random noise.

model complexity / polynomial degree error best balance bias^2 decreases variance increases total error underfit overfit

As polynomial degree increases, bias usually goes down, variance usually goes up, and total expected error is minimized at a middle level of complexity.

Degree selection

If degree 1 underfits and degree 15 overfits, how should we choose the degree?

Use validation data or cross-validation instead of choosing by training error.

Previous: Simulation Next: Validation