Random Forests

One decision tree is a useful opinion. A forest makes that opinion more dependable.

A progressive path from ensemble intuition to bootstrap sampling, bagging, random feature selection, classification, regression, OOB evaluation, variance mathematics, tuning, and practical model choice.

The central question

Opening question

If a small change in the training data can change one tree, how could several trees make a more stable prediction?

Random Forest combines two sources of randomness: each tree sees a bootstrap sample of the rows, and each node considers only a random subset of features. The trees become different enough that their mistakes can cancel when their predictions are aggregated.

\[\hat f(x)=\frac{1}{B}\sum_{b=1}^B f_b(x)\]

Intuition: one tree may be wrong for a particular row. If the other trees are not wrong in exactly the same way, averaging or voting makes the final answer steadier.

Tree 1Tree 2Tree 3 YesNoYesYesNoYes Forest: Yes

The forest aggregates many different tree opinions.

Roadmap

Three definitions to keep separate

IdeaWhat is randomized?What is combined?
EnsembleNot necessarily randomizedPredictions from several models
BaggingBootstrap samples of rowsPredictions from base learners
Random ForestBootstrap rows plus feature subsets at each splitPredictions from decision trees

Memory hook: ensemble is the umbrella; bagging is a way to create an ensemble; Random Forest is bagged decision trees with extra feature randomness.

References

Next: Why ensemble