The split criterion trains the tree; evaluation metrics decide whether it is useful.
Separate model fitting from business evaluation, tune complexity with validation folds, and reserve the test set for one final estimate.
Three different questions
| Stage | Question | Typical tool |
|---|---|---|
| Split selection | Which local question best improves node purity? | Entropy, Gini, or squared error. |
| Model selection | Which depth and leaf size generalize best? | Validation or cross-validation. |
| Business evaluation | Are the resulting errors acceptable? | Precision, recall, costs, MAE, calibration, fairness checks. |
Intuition: high information gain does not guarantee that the completed model meets the real-world objective.
Train, validation, and test
Data-use question
If we try twenty depths and keep the one with the best test score, is the test set still untouched?
fit candidate treesValidation folds
choose settingsRefit chosen setupTest data
evaluate onceMonitor after launch
K-fold cross-validation
| Symbol | Meaning |
|---|---|
| \(K\) | Number of validation folds. |
| \(\lambda\) | A candidate hyperparameter setting such as depth and leaf size. |
| \(Metric_k\) | Score on fold \(k\) after fitting on all other folds. |
Intuition: each row gets a turn as validation evidence, producing a less lucky estimate than one small validation split.
Use stratified folds for classification when class proportions should remain similar. Use ordinary K-fold for standard regression, and time-aware splits for temporal data.
Tune related controls together
Depth and leaf-size settings interact. A large `min_samples_leaf` may regularize an otherwise deep tree, so evaluating them independently can miss useful combinations.
Classification metrics follow error cost
Loan decision question
Is a false approval equally costly as a false decline?
| Metric | Useful when |
|---|---|
| Accuracy | Classes and mistake costs are reasonably balanced. |
| Precision | False positive decisions are costly. |
| Recall | Missing positive cases is costly. |
| F1 | A single balance between precision and recall is needed. |
| ROC-AUC | Ranking quality across thresholds matters and imbalance is not extreme. |
| PR-AUC | Positive-class retrieval matters in an imbalanced problem. |
| Expected cost | The organization can assign costs or utilities to each outcome. |
Class imbalance changes both fitting and evaluation
With class weights \(w_k\), class proportions inside a node become weighted proportions:
Intuition: a minority-class row can count as more than one unit when impurity and split improvements are calculated.
- Use stratified train/test and cross-validation splits.
- Compare baseline, class-weighted, and sampling strategies inside validation only.
- Report minority precision, recall, F1, PR-AUC, and confusion counts.
- Choose the probability threshold from business costs, not automatically 0.5.
Probability quality
Calibration question
If a leaf contains 2 positive rows and 0 negative rows, is 100% a reliable probability estimate?
It is the empirical leaf proportion, but its uncertainty is enormous. Deep trees often produce coarse and extreme probabilities.
| Check | Why |
|---|---|
| Calibration curve | Compares predicted probability with observed frequency. |
| Brier score | Measures squared probability error. |
| Log loss | Strongly penalizes confident wrong probabilities. |
| Leaf support | Shows how many examples produced the probability. |
Regression metrics
| Metric | Expression | Interpretation |
|---|---|---|
| MAE | \(\frac{1}{n}\sum|y_i-\hat y_i|\) | Average absolute miss in target units. |
| RMSE | \(\sqrt{\frac{1}{n}\sum(y_i-\hat y_i)^2}\) | Emphasizes larger misses. |
| \(R^2\) | \(1-\frac{SS_{res}}{SS_{tot}}\) | Improvement over predicting the target mean. |
Intuition: MAE asks, “How far away are predictions on average?” RMSE asks the same question but makes large misses hurt more. \(R^2\) compares the tree with the simple baseline that predicts the training-target mean for every row.
Inspect residuals by target range and important groups. One average metric can hide systematically poor regions.
Report variability, not only the mean
Cross-validation results should include fold-to-fold variation:
Intuition: two models with the same average may differ greatly in stability. A tree whose score changes sharply across folds may be sensitive to the sample.