Impurity tells a tree how mixed the labels are.
Entropy describes expected uncertainty; Gini describes how often a label drawn from the node would disagree with another draw. Both are low for pure nodes and high for mixed nodes.
Begin with purity
Prediction question
Which leaf gives a safer prediction: 10 approvals and 0 declines, or 5 approvals and 5 declines?
10 / 0
Pure. The next label is completely predictable from this node.
9 / 1
Mostly pure. There is a small amount of uncertainty.
5 / 5
Maximally mixed for two classes. Either label is equally plausible.
Surprise before entropy
Information question
Which event is more informative: something expected with probability 0.9, or something rare with probability 0.1?
| Probability | Surprise | Meaning |
|---|---|---|
| 1 | \(-\log_2(1)=0\) bits | Certain; learning it adds no information. |
| 0.5 | \(-\log_2(0.5)=1\) bit | One fair yes/no question resolves it. |
| 0.125 | \(-\log_2(0.125)=3\) bits | Rare and therefore more surprising. |
Intuition: information is large when reality gives us something we did not expect.
Entropy is average surprise
| Symbol | Meaning |
|---|---|
| \(Y\) | The target label at the current node. |
| \(K\) | Number of target classes. |
| \(p_k\) | Fraction of node samples belonging to class \(k\). |
| \(H(Y)\) | Expected number of bits needed to identify the label. |
Intuition: entropy averages the surprise of every possible class, weighted by how often that class occurs.
Meet the complete cricket dataset
Choose the first question
Before calculating anything, scan all 14 days. Would you first ask about Weather, Temperature, Humidity, or Wind?
Each row describes the conditions on one day. The four input features are Weather, Temperature, Humidity, and Wind. The target Play cricket? records whether cricket was played.
| Day | Weather | Temperature | Humidity | Wind | Play cricket? |
|---|---|---|---|---|---|
| D1 | Sunny | Hot | High | Weak | No |
| D2 | Sunny | Hot | High | Strong | No |
| D3 | Overcast | Hot | High | Weak | Yes |
| D4 | Rain | Mild | High | Weak | Yes |
| D5 | Rain | Cool | Normal | Weak | Yes |
| D6 | Rain | Cool | Normal | Strong | No |
| D7 | Overcast | Cool | Normal | Strong | Yes |
| D8 | Sunny | Mild | High | Weak | No |
| D9 | Sunny | Cool | Normal | Weak | Yes |
| D10 | Rain | Mild | Normal | Weak | Yes |
| D11 | Sunny | Mild | Normal | Strong | Yes |
| D12 | Overcast | Mild | High | Strong | Yes |
| D13 | Overcast | Hot | Normal | Weak | Yes |
| D14 | Rain | Mild | High | Strong | No |
How to read it: Weather has three possible values, while the other inputs each have two or three values. All inputs are categorical in this teaching example. Counting the target column gives 9 Yes days and 5 No days.
Cricket target: calculate every term
From the complete table above, the dataset contains 9 \(Play=Yes\) days and 5 \(Play=No\) days.
Intuition: the node is fairly mixed, so its entropy is close to the binary maximum of 1 bit.
Binary entropy is zero at pure endpoints and largest at a 50/50 mixture.
Explore entropy
Before moving the slider
What should happen to entropy as one class moves from 50% toward 100%?
Class 0 proportion: 0.50
Entropy: 1.000 bits
Gini impurity
Alternative question
Can we measure mixture without logarithms?
Intuition: imagine randomly assigning a label according to the node's class proportions. Gini is the probability that this random label disagrees with the true class.
Entropy versus Gini
| Property | Entropy | Gini |
|---|---|---|
| Formula | \(-\sum p_k\log_2p_k\) | \(1-\sum p_k^2\) |
| Binary maximum | 1 at 50/50 | 0.5 at 50/50 |
| Pure node | 0 | 0 |
| Interpretation | Expected uncertainty in bits | Expected disagreement under proportion-based random labelling |
| Typical result | Often chooses similar splits; neither is universally superior. | |
Edge cases
- We define \(0\log_2(0)=0\), so an absent class contributes nothing.
- Entropy handles any number of classes; binary entropy is only the easiest visualization.
- Sample or class weights replace simple counts with weighted counts.
- Impurity describes the current training node. It does not directly report future accuracy.