Part 5

For classification, a forest turns many tree opinions into a class and a probability.

The individual trees use the same target but see different rows and feature candidates. Their predictions are combined at the end.

One loan application, several trees

Prediction question

If a forest contains 100 trees and 72 predict Approved, what is the hard-vote prediction?

Suppose a new application follows a different path through each tree. The forest may produce this vote summary:

PredictionNumber of treesFraction
Approved720.72
Declined280.28
\[\hat y=\operatorname{mode}(72A,28D)=A\]

Intuition: the class is the forest's most common opinion, not the class predicted by a single “best” tree.

Approved 72%Declined 28%probability / vote sharesample of four tree votes

Hard vote versus averaged probability

Probability question

Why might 72% of tree votes be more informative than only the label Approved?

For a binary problem, a hard vote keeps only whether each tree is above its own classification threshold. A probability aggregate retains more information:

\[\hat p(A\mid x)=\frac{1}{B}\sum_{b=1}^B\hat p_b(A\mid x)\]

In scikit-learn, a tree's class probability is the fraction of training samples from that class in the reached leaf, and the forest averages those probabilities. A separate business threshold can then turn probability into an action.

Classification workflow

  1. Split data before fitting the forest; stratify when class proportions matter.
  2. Fit the forest on training rows only.
  3. Inspect validation metrics beyond accuracy when costs are unequal.
  4. Choose the operating threshold using validation data.
  5. Evaluate once on the untouched test set.
A forest probability is not automatically perfectly calibrated. Check calibration if probabilities drive decisions, prices, or interventions.
Previous: FeaturesNext: Regression