Part 2

Bootstrap sampling creates different training worlds for different trees.

A bootstrap sample draws rows with replacement until it has the same size as the original dataset. Some rows repeat; some are left out.

Sampling with replacement

Prediction question

For rows A, B, C, and D, what could a four-draw bootstrap sample look like?

After drawing a row, we put it back before the next draw. One possible sample is:

Original rows: A B C D
Bootstrap sample: C A C D

C appears twice; B is absent.

Intuition: each tree gets a slightly different view of the same training set. This is the source of row-level diversity in bagging.

Original datasetABCDBootstrap drawsCACDB is absentC repeats

Why about 36.8% are out of bag

Probability question

In \(n\) draws, what is the chance that one particular row is never selected?

On one draw, the chance of missing a particular row is \(1-1/n\). Missing it on all \(n\) independent draws gives:

\[P(\text{row is OOB})=\left(1-\frac{1}{n}\right)^n\xrightarrow[n\to\infty]{}e^{-1}\approx0.368\]

Intuition: roughly 36.8% of the original rows are not used by any one tree. They can later act as that tree's unseen mini-test set.

OOB probability: 0.366

Expected unique fraction: 0.634

Small-sample warning

For a small dataset, 36.8% is an expectation, not a guarantee. A particular bootstrap sample may leave out more or fewer rows. Across many trees, however, each row is usually OOB for a useful number of trees.

Previous: Why ensembleNext: Bagging