Information gain measures how much uncertainty a question removes.
Using the complete 14-day cricket dataset from the previous page, we now calculate one split, explain why child impurity is weighted, compare every feature, and grow the first levels of the tree.
The value of a question
Split question
If the parent entropy is 0.940 bits and the expected child entropy is 0.694 bits, what did the question accomplish?
| Symbol | Meaning |
|---|---|
| \(Y\) | The target, here whether cricket is played. |
| \(A\) | A candidate feature or question, such as Weather. |
| \(H(Y)\) | Uncertainty before knowing the answer. |
| \(H(Y\mid A)\) | Expected uncertainty after knowing the answer. |
Intuition: information gain is uncertainty before minus uncertainty after. A larger drop means a more useful question.
Why child entropy must be weighted
Weighting question
Should a child containing 1 row influence the split score as much as a child containing 13 rows?
The fraction \(|D_v|/|D|\) is the chance that a randomly selected training row follows branch \(v\).
Intuition: the post-split score is the impurity a random row is expected to encounter, so large children must count more.
Split the cricket data by Weather
Before calculating
Which Weather branch is already pure, and which branches still need another question?
| Weather | Rows | Play Yes | Play No | Entropy | Weight |
|---|---|---|---|---|---|
| Sunny | 5 | 2 | 3 | 0.971 | 5/14 |
| Overcast | 4 | 4 | 0 | 0.000 | 4/14 |
| Rain | 5 | 3 | 2 | 0.971 | 5/14 |
Intuition: Overcast removes all uncertainty, while Sunny and Rain remain mixed. The weighted average summarizes what a random day experiences.
Weather information gain
Intuition: knowing Weather saves about 0.247 bits of label uncertainty on average.
Pure branches stop; mixed branches repeat the split search locally.
Compare all candidate features
Root question
Should the root use the feature with the largest number of categories, or the feature with the largest measured gain?
| Feature | Information gain |
|---|---|
| Weather | 0.247 |
| Humidity | 0.152 |
| Wind | 0.048 |
| Temperature | 0.029 |
`Weather` becomes the root because it gives the largest immediate reduction in uncertainty.
Repeat locally
Sunny branch
Among Sunny days only, which remaining feature best separates 2 Yes from 3 No?
Inside the Sunny subset, `Humidity` perfectly separates the labels: High gives No and Normal gives Yes. Inside Rain, `Wind` separates Weak from Strong.
Intuition: the algorithm is recursive because each mixed branch becomes a smaller version of the original learning problem.
Greedy does not mean globally optimal
The tree chooses the best immediate split. It does not enumerate every possible future tree, which would become combinatorially expensive.
Intuition: choose the most useful next question, not the provably best complete sequence of all future questions.
ID3, C4.5, and CART
| Family | Main idea | Important distinction |
|---|---|---|
| ID3 | Entropy and information gain | Classic treatment often uses multiway categorical splits. |
| C4.5 | Extension of ID3 | Handles continuous features, missingness strategies, and gain ratio. |
| CART | Classification and regression trees | Uses binary trees; scikit-learn implements an optimized CART-style algorithm. |