Part 6

Tree depth is model capacity: too little misses structure; too much memorizes noise.

We connect geometric complexity to train/validation behavior, then control the tree before growth or prune weak branches afterward.

Depth changes what the model can express

Capacity question

Should a tree continue splitting until every training leaf is pure?

Depth 1: underfit Depth 3: useful structure Deep tree: memorized pocket

A deep tree can create a small region solely to classify one noisy training point.

Bias and variance language

TreeTraining behaviorGeneralization risk
Too shallowMisses genuine interactions; training error remains high.Underfitting: high bias.
ModerateCaptures repeatable structure without tiny leaves.Useful bias-variance balance.
Too deepTraining error can approach zero.Overfitting: high variance and sensitivity to sample noise.

Intuition: depth buys flexibility. Flexibility helps until the model starts explaining accidents of the training sample.

Training and validation curves

Curve question

Why can validation error rise while training error continues to fall?

Every extra split is chosen to help the training subset. A weak split may capture noise that does not repeat in unseen data.

Choose complexity using validation, not training fit alone.
training errorvalidation errorselected depth Tree depthError

Pre-pruning controls

Control question

Which parameter directly prevents a leaf supported by only one training row?

ControlQuestion it answersTypical effect
`max_depth`How many questions may one path contain?Limits global complexity.
`min_samples_split`How many rows must a node have before splitting?Stops very small internal nodes.
`min_samples_leaf`How many rows must every child retain?Prevents tiny leaves and smooths predictions.
`max_leaf_nodes`How many terminal regions are allowed?Caps model size directly.
`min_impurity_decrease`How much weighted improvement is required?Rejects weak questions.

`min_samples_leaf` is often particularly intuitive: every prediction must be supported by at least that many training examples.

Why weighted impurity decrease includes node size

\[ \Delta I=\frac{N_t}{N}\left(I_t-\frac{N_L}{N_t}I_L-\frac{N_R}{N_t}I_R\right) \]
TermMeaning
\(N\)Total weighted training samples.
\(N_t\)Samples reaching the current node.
\(I_t,I_L,I_R\)Parent, left-child, and right-child impurity.

Intuition: an impressive local improvement deep in a tiny branch should count less than the same improvement affecting half the dataset.

Post-pruning with cost complexity

Pruning question

How can we allow a large tree to grow and then charge it for unnecessary leaves?

\[ R_\alpha(T)=R(T)+\alpha|\widetilde T| \]
SymbolMeaning
\(R(T)\)Total leaf error or, in scikit-learn, total sample-weighted leaf impurity.
\(|\widetilde T|\)Number of leaves.
\(\alpha\)Complexity price paid for each leaf.

Intuition: a branch survives only when its improvement is worth the extra leaves it creates.

A small pruning calculation

Suppose \(\alpha=0.02\):

TreeLeaf impurity \(R(T)\)LeavesCost complexity
Large tree0.128\(0.12+0.02(8)=0.28\)
Smaller tree0.184\(0.18+0.02(4)=0.26\)

The larger tree fits better before the penalty, but the smaller tree wins after complexity is priced.

Weakest-link pruning

\[ \alpha_{eff}(t)=\frac{R(t)-R(T_t)}{|\widetilde T_t|-1} \]

Intuition: \(\alpha_{eff}\) measures improvement per extra leaf for a branch. The branch providing the least improvement per added leaf is removed first.

Cross-validation chooses `ccp_alpha` from the pruning path. The untouched test set is used only after that choice.

Practical regularization order

  1. Start with a shallow, interpretable baseline.
  2. Tune `max_depth` and `min_samples_leaf` with validation folds.
  3. Inspect train-versus-validation gaps and leaf sample counts.
  4. Evaluate a cost-complexity pruning path when useful.
  5. Choose the simplest tree whose validation performance is competitive.
Previous: RecursionNext: Regression Trees