Part 4

Bagging changes the rows. Random Forest also changes which features compete at each split.

If one very strong feature is always available, every tree may make the same first split. Feature subsampling encourages different trees to discover different useful structures.

Bagging alone can still create similar trees

Diversity question

If CreditScore is the best split in nearly every bootstrap sample, what might happen to tree correlation?

Bootstrap samples differ, but a dominant feature may still win repeatedly. The trees then make similar decisions and similar errors. Random Forest limits the candidates seen at a node.

At one node with six available featuresCreditIncomeAgeTenureBalanceRegionTree 1 sees:CreditAgeRegionTree 2 sees:IncomeTenureBalance

A fresh subset is considered at every split. The best split is chosen only among that subset.

The two randomness knobs

RandomnessWhere?What it buys us
Bootstrap rowsDifferent training sample for each treeDifferent fitted trees
Feature subsetAt each node while a tree growsLess dominance by one feature and lower correlation
\[m=\texttt{max\_features}\quad\text{features considered at one node}\]

Intuition: a feature can be important without being available at every split. That temporary absence forces the forest to learn backup routes.

Strength versus correlation

Trade-off question

What happens if \(m\) is extremely small: do trees become more different, more accurate individually, or both?

A smaller \(m\) usually lowers correlation, but it can also make individual trees weaker because their best feature is often unavailable. A larger \(m\) strengthens trees but can make them more alike. The useful setting balances both.

Rowsbootstrap diversity
Featuressplit diversity
Forestaggregated prediction

How many features should compete?

Choice question

Suppose a dataset has \(p=20\) features. Would you let all 20 compete at every split, or would you deliberately hide some?

Let \(m\) be the number of candidate features sampled at one node. In scikit-learn, this is controlled by max_features. For example, with \(p=20\):

SettingFeatures consideredTypical consequence
1.0 or None20Strong individual trees, but often more correlated
"sqrt"\(\lfloor\sqrt{20}\rfloor=4\)Common classification starting point
0.3\(\lfloor0.3\times20\rfloor=6\)A middle level of feature randomness
33Very diverse candidates; individual trees may be weaker

Intuition: feature sampling is a controlled trade-off. Smaller \(m\) makes trees disagree more, which can lower ensemble variance; larger \(m\) gives each tree better choices, which can lower its bias.

There is no universally correct value. Classification often starts with "sqrt"; regression often starts with all features or a large fraction. These are starting points, not rules.

A practical way to choose \(m\)

Validation question

Which value should win: the one that makes individual trees strongest, or the one that gives the best validation performance after averaging?

  1. Choose a validation design that matches the data: stratified folds for classification, grouped folds for related rows, or time-based splits for time order.
  2. Try a small set such as "sqrt", "log2", 0.3, 0.7, and 1.0.
  3. Keep the other important settings fixed while comparing, and use the metric that matches the decision.
  4. Repeat with a few random seeds or cross-validation folds. Prefer a stable setting whose score is strong, not a fragile winner by a tiny margin.
  5. After choosing \(m\), tune depth and leaf size if needed, then evaluate once on the untouched test set.
\[\widehat{m}=\arg\max_{m\in\mathcal{M}}\;\operatorname{CVScore}(m)\]

Here, \(\mathcal{M}\) is the small set of candidate values. The forest is judged as an ensemble, because the goal is a reliable combined prediction rather than the best-looking single tree.

Previous: BaggingNext: Classification