Bagging changes the rows. Random Forest also changes which features compete at each split.
If one very strong feature is always available, every tree may make the same first split. Feature subsampling encourages different trees to discover different useful structures.
Bagging alone can still create similar trees
Diversity question
If CreditScore is the best split in nearly every bootstrap sample, what might happen to tree correlation?
Bootstrap samples differ, but a dominant feature may still win repeatedly. The trees then make similar decisions and similar errors. Random Forest limits the candidates seen at a node.
A fresh subset is considered at every split. The best split is chosen only among that subset.
The two randomness knobs
| Randomness | Where? | What it buys us |
|---|---|---|
| Bootstrap rows | Different training sample for each tree | Different fitted trees |
| Feature subset | At each node while a tree grows | Less dominance by one feature and lower correlation |
Intuition: a feature can be important without being available at every split. That temporary absence forces the forest to learn backup routes.
Strength versus correlation
Trade-off question
What happens if \(m\) is extremely small: do trees become more different, more accurate individually, or both?
A smaller \(m\) usually lowers correlation, but it can also make individual trees weaker because their best feature is often unavailable. A larger \(m\) strengthens trees but can make them more alike. The useful setting balances both.
How many features should compete?
Choice question
Suppose a dataset has \(p=20\) features. Would you let all 20 compete at every split, or would you deliberately hide some?
Let \(m\) be the number of candidate features sampled at one node. In scikit-learn, this is controlled by max_features. For example, with \(p=20\):
| Setting | Features considered | Typical consequence |
|---|---|---|
1.0 or None | 20 | Strong individual trees, but often more correlated |
"sqrt" | \(\lfloor\sqrt{20}\rfloor=4\) | Common classification starting point |
0.3 | \(\lfloor0.3\times20\rfloor=6\) | A middle level of feature randomness |
3 | 3 | Very diverse candidates; individual trees may be weaker |
Intuition: feature sampling is a controlled trade-off. Smaller \(m\) makes trees disagree more, which can lower ensemble variance; larger \(m\) gives each tree better choices, which can lower its bias.
"sqrt"; regression often starts with all features or a large fraction. These are starting points, not rules.A practical way to choose \(m\)
Validation question
Which value should win: the one that makes individual trees strongest, or the one that gives the best validation performance after averaging?
- Choose a validation design that matches the data: stratified folds for classification, grouped folds for related rows, or time-based splits for time order.
- Try a small set such as
"sqrt","log2",0.3,0.7, and1.0. - Keep the other important settings fixed while comparing, and use the metric that matches the decision.
- Repeat with a few random seeds or cross-validation folds. Prefer a stable setting whose score is strong, not a fragile winner by a tiny margin.
- After choosing \(m\), tune depth and leaf size if needed, then evaluate once on the untouched test set.
Here, \(\mathcal{M}\) is the small set of candidate values. The forest is judged as an ensemble, because the goal is a reliable combined prediction rather than the best-looking single tree.