Tree depth is model capacity: too little misses structure; too much memorizes noise.
We connect geometric complexity to train/validation behavior, then control the tree before growth or prune weak branches afterward.
Depth changes what the model can express
Capacity question
Should a tree continue splitting until every training leaf is pure?
A deep tree can create a small region solely to classify one noisy training point.
Bias and variance language
| Tree | Training behavior | Generalization risk |
|---|---|---|
| Too shallow | Misses genuine interactions; training error remains high. | Underfitting: high bias. |
| Moderate | Captures repeatable structure without tiny leaves. | Useful bias-variance balance. |
| Too deep | Training error can approach zero. | Overfitting: high variance and sensitivity to sample noise. |
Intuition: depth buys flexibility. Flexibility helps until the model starts explaining accidents of the training sample.
Training and validation curves
Curve question
Why can validation error rise while training error continues to fall?
Every extra split is chosen to help the training subset. A weak split may capture noise that does not repeat in unseen data.
Pre-pruning controls
Control question
Which parameter directly prevents a leaf supported by only one training row?
| Control | Question it answers | Typical effect |
|---|---|---|
| `max_depth` | How many questions may one path contain? | Limits global complexity. |
| `min_samples_split` | How many rows must a node have before splitting? | Stops very small internal nodes. |
| `min_samples_leaf` | How many rows must every child retain? | Prevents tiny leaves and smooths predictions. |
| `max_leaf_nodes` | How many terminal regions are allowed? | Caps model size directly. |
| `min_impurity_decrease` | How much weighted improvement is required? | Rejects weak questions. |
`min_samples_leaf` is often particularly intuitive: every prediction must be supported by at least that many training examples.
Why weighted impurity decrease includes node size
| Term | Meaning |
|---|---|
| \(N\) | Total weighted training samples. |
| \(N_t\) | Samples reaching the current node. |
| \(I_t,I_L,I_R\) | Parent, left-child, and right-child impurity. |
Intuition: an impressive local improvement deep in a tiny branch should count less than the same improvement affecting half the dataset.
Post-pruning with cost complexity
Pruning question
How can we allow a large tree to grow and then charge it for unnecessary leaves?
| Symbol | Meaning |
|---|---|
| \(R(T)\) | Total leaf error or, in scikit-learn, total sample-weighted leaf impurity. |
| \(|\widetilde T|\) | Number of leaves. |
| \(\alpha\) | Complexity price paid for each leaf. |
Intuition: a branch survives only when its improvement is worth the extra leaves it creates.
A small pruning calculation
Suppose \(\alpha=0.02\):
| Tree | Leaf impurity \(R(T)\) | Leaves | Cost complexity |
|---|---|---|---|
| Large tree | 0.12 | 8 | \(0.12+0.02(8)=0.28\) |
| Smaller tree | 0.18 | 4 | \(0.18+0.02(4)=0.26\) |
The larger tree fits better before the penalty, but the smaller tree wins after complexity is priced.
Weakest-link pruning
Intuition: \(\alpha_{eff}\) measures improvement per extra leaf for a branch. The branch providing the least improvement per added leaf is removed first.
Cross-validation chooses `ccp_alpha` from the pruning path. The untouched test set is used only after that choice.
Practical regularization order
- Start with a shallow, interpretable baseline.
- Tune `max_depth` and `min_samples_leaf` with validation folds.
- Inspect train-versus-validation gaps and leaf sample counts.
- Evaluate a cost-complexity pruning path when useful.
- Choose the simplest tree whose validation performance is competitive.