Before building a forest, understand why one tree can be fragile.
Decision trees are flexible and intuitive, but a small change in the observed sample can change the first split and therefore the whole tree.
Prediction instability
Compare the models
Two trees see almost the same data. Should their decision boundaries be identical?
A deep tree can create very specific rectangular regions. A few changed or duplicated rows can change which feature wins a split. This is high variance: the model is sensitive to the particular sample used for training.
Both trees may fit their own samples well but disagree near a boundary. An ensemble changes the question from “Which one tree is correct?” to “What prediction is common across many plausible trees?”
Intuition: bagging is mainly a variance-reduction strategy. It does not magically remove bad features, label noise, or systematic bias.
Ensemble learning
Voting question
If five independent doctors give four “Yes” opinions and one “No” opinion, what should a simple majority vote predict?
An ensemble combines multiple base estimators. For classification, the final class may be the most common vote. For regression, the final value is usually an average.
Hard vote
Each classifier gives one class; choose the mode.
Soft vote
Each classifier gives probabilities; average the probabilities.
Average
Each regressor gives a number; average the numbers.
When does combining help?
Correlation question
Would averaging help if every tree made exactly the same mistake on every row?
The ensemble benefits when individual models are reasonably strong but do not make perfectly correlated errors. Random Forest creates diversity deliberately through rows and features.
| Ingredient | Effect |
|---|---|
| Strong trees | Each tree should contain useful signal. |
| Different trees | Their errors should not line up perfectly. |
| Aggregation | Common signal is preserved; some random error cancels. |