Part 2

Impurity tells a tree how mixed the labels are.

Entropy describes expected uncertainty; Gini describes how often a label drawn from the node would disagree with another draw. Both are low for pure nodes and high for mixed nodes.

Begin with purity

Prediction question

Which leaf gives a safer prediction: 10 approvals and 0 declines, or 5 approvals and 5 declines?

10 / 0

Pure. The next label is completely predictable from this node.

9 / 1

Mostly pure. There is a small amount of uncertainty.

5 / 5

Maximally mixed for two classes. Either label is equally plausible.

Surprise before entropy

Information question

Which event is more informative: something expected with probability 0.9, or something rare with probability 0.1?

\[ I(\text{event})=-\log_2 P(\text{event}) \]
ProbabilitySurpriseMeaning
1\(-\log_2(1)=0\) bitsCertain; learning it adds no information.
0.5\(-\log_2(0.5)=1\) bitOne fair yes/no question resolves it.
0.125\(-\log_2(0.125)=3\) bitsRare and therefore more surprising.

Intuition: information is large when reality gives us something we did not expect.

Entropy is average surprise

\[ H(Y)=-\sum_{k=1}^{K}p_k\log_2(p_k) \]
SymbolMeaning
\(Y\)The target label at the current node.
\(K\)Number of target classes.
\(p_k\)Fraction of node samples belonging to class \(k\).
\(H(Y)\)Expected number of bits needed to identify the label.

Intuition: entropy averages the surprise of every possible class, weighted by how often that class occurs.

Meet the complete cricket dataset

Choose the first question

Before calculating anything, scan all 14 days. Would you first ask about Weather, Temperature, Humidity, or Wind?

Each row describes the conditions on one day. The four input features are Weather, Temperature, Humidity, and Wind. The target Play cricket? records whether cricket was played.

DayWeatherTemperatureHumidityWindPlay cricket?
D1SunnyHotHighWeakNo
D2SunnyHotHighStrongNo
D3OvercastHotHighWeakYes
D4RainMildHighWeakYes
D5RainCoolNormalWeakYes
D6RainCoolNormalStrongNo
D7OvercastCoolNormalStrongYes
D8SunnyMildHighWeakNo
D9SunnyCoolNormalWeakYes
D10RainMildNormalWeakYes
D11SunnyMildNormalStrongYes
D12OvercastMildHighStrongYes
D13OvercastHotNormalWeakYes
D14RainMildHighStrongNo

How to read it: Weather has three possible values, while the other inputs each have two or three values. All inputs are categorical in this teaching example. Counting the target column gives 9 Yes days and 5 No days.

The goal is to learn a sequence of questions that predicts Play cricket?. We will first measure the uncertainty in all 14 labels, then test which feature reduces it the most.

Cricket target: calculate every term

From the complete table above, the dataset contains 9 \(Play=Yes\) days and 5 \(Play=No\) days.

\[ p_{Yes}=\frac{9}{14},\qquad p_{No}=\frac{5}{14} \]
\[ \begin{aligned} H(Play)&=-\frac{9}{14}\log_2\frac{9}{14}-\frac{5}{14}\log_2\frac{5}{14}\\ &\approx 0.940\text{ bits} \end{aligned} \]

Intuition: the node is fairly mixed, so its entropy is close to the binary maximum of 1 bit.

9/14 → H ≈ 0.94 00.51 P(Yes)Entropy

Binary entropy is zero at pure endpoints and largest at a 50/50 mixture.

Explore entropy

Before moving the slider

What should happen to entropy as one class moves from 50% toward 100%?

50%50%

Class 0 proportion: 0.50

Entropy: 1.000 bits

Gini impurity

Alternative question

Can we measure mixture without logarithms?

\[ Gini(Y)=1-\sum_{k=1}^{K}p_k^2 \]

Intuition: imagine randomly assigning a label according to the node's class proportions. Gini is the probability that this random label disagrees with the true class.

\[ Gini(9\ Yes,5\ No)=1-\left(\frac{9}{14}\right)^2-\left(\frac{5}{14}\right)^2\approx0.459 \]

Entropy versus Gini

PropertyEntropyGini
Formula\(-\sum p_k\log_2p_k\)\(1-\sum p_k^2\)
Binary maximum1 at 50/500.5 at 50/50
Pure node00
InterpretationExpected uncertainty in bitsExpected disagreement under proportion-based random labelling
Typical resultOften chooses similar splits; neither is universally superior.
Do not compare their raw magnitudes directly. Their scales differ. Compare how much each criterion decreases after a candidate split.

Edge cases

Previous: IntuitionNext: Information Gain