Part 8

The mathematics of a forest is a story about variance and correlation.

Averaging does not erase every error. It reduces the part of the error that varies from tree to tree, especially when tree predictions are not too correlated.

Start with two trees

Variance question

If two tree predictions fluctuate independently around the same average, what should happen when we average them?

Let each tree's prediction have variance \(\sigma^2\). If trees were independent, averaging \(B\) of them would divide variance by \(B\):

\[\operatorname{Var}\left(\frac{1}{B}\sum_{b=1}^B f_b\right)=\frac{\sigma^2}{B}\]

Intuition: random high and low deviations partially cancel. More trees make the average steadier.

Real trees are correlated

Correlation question

What if every tree still uses the same dominant feature and makes almost the same error?

Let \(\rho\) be the pairwise correlation between tree predictions. The variance of the average becomes:

\[\begin{aligned}\operatorname{Var}(\bar f)&=\frac{1}{B^2}\left[B\sigma^2+B(B-1)\rho\sigma^2\right]\\&=\sigma^2\left(\rho+\frac{1-\rho}{B}\right)\end{aligned}\]
SymbolMeaning
\(B\)Number of trees.
\(\sigma^2\)Variance of an individual tree's prediction.
\(\rho\)How similarly trees move and make errors.

Key insight: as \(B\) becomes very large, the \(1/B\) term vanishes, but the correlated part \(\rho\sigma^2\) remains. More trees cannot cancel a shared mistake.

Explore the variance factor

Relative ensemble variance \(\sigma^2\) times: 0.208

Try increasing \(B\), then increasing \(\rho\). The first change keeps helping, while the second creates a floor that more trees cannot remove.

How Random Forest targets the equation

Bootstrap rows

Changes the training examples.

Feature subsets

Changes the candidate questions at nodes.

Aggregation

Turns many noisy predictions into one stable estimate.

Previous: OOBNext: Tuning