Part 8

The split criterion trains the tree; evaluation metrics decide whether it is useful.

Separate model fitting from business evaluation, tune complexity with validation folds, and reserve the test set for one final estimate.

Three different questions

StageQuestionTypical tool
Split selectionWhich local question best improves node purity?Entropy, Gini, or squared error.
Model selectionWhich depth and leaf size generalize best?Validation or cross-validation.
Business evaluationAre the resulting errors acceptable?Precision, recall, costs, MAE, calibration, fairness checks.

Intuition: high information gain does not guarantee that the completed model meets the real-world objective.

Train, validation, and test

Data-use question

If we try twenty depths and keep the one with the best test score, is the test set still untouched?

Training data
fit candidate trees
Validation folds
choose settings
Refit chosen setupTest data
evaluate once
Monitor after launch
Repeatedly choosing settings from test performance leaks test information into model selection and makes the final score optimistic.

K-fold cross-validation

\[ CV(\lambda)=\frac{1}{K}\sum_{k=1}^{K}Metric_k(\lambda) \]
SymbolMeaning
\(K\)Number of validation folds.
\(\lambda\)A candidate hyperparameter setting such as depth and leaf size.
\(Metric_k\)Score on fold \(k\) after fitting on all other folds.

Intuition: each row gets a turn as validation evidence, producing a less lucky estimate than one small validation split.

Use stratified folds for classification when class proportions should remain similar. Use ordinary K-fold for standard regression, and time-aware splits for temporal data.

Tune related controls together

param_grid = { "max_depth": [2, 3, 5, None], "min_samples_leaf": [1, 5, 15], "min_samples_split": [2, 10, 30], "ccp_alpha": [0.0, 0.001, 0.01] }

Depth and leaf-size settings interact. A large `min_samples_leaf` may regularize an otherwise deep tree, so evaluating them independently can miss useful combinations.

Prefer the simplest setting whose cross-validation score is meaningfully competitive, especially when explanation stability matters.

Classification metrics follow error cost

Loan decision question

Is a false approval equally costly as a false decline?

MetricUseful when
AccuracyClasses and mistake costs are reasonably balanced.
PrecisionFalse positive decisions are costly.
RecallMissing positive cases is costly.
F1A single balance between precision and recall is needed.
ROC-AUCRanking quality across thresholds matters and imbalance is not extreme.
PR-AUCPositive-class retrieval matters in an imbalanced problem.
Expected costThe organization can assign costs or utilities to each outcome.

Class imbalance changes both fitting and evaluation

With class weights \(w_k\), class proportions inside a node become weighted proportions:

\[ p_{mk}=\frac{\sum_{i\in Q_m}w_{y_i}\mathbf{1}(y_i=k)}{\sum_{i\in Q_m}w_{y_i}} \]

Intuition: a minority-class row can count as more than one unit when impurity and split improvements are calculated.

Probability quality

Calibration question

If a leaf contains 2 positive rows and 0 negative rows, is 100% a reliable probability estimate?

It is the empirical leaf proportion, but its uncertainty is enormous. Deep trees often produce coarse and extreme probabilities.

CheckWhy
Calibration curveCompares predicted probability with observed frequency.
Brier scoreMeasures squared probability error.
Log lossStrongly penalizes confident wrong probabilities.
Leaf supportShows how many examples produced the probability.

Regression metrics

MetricExpressionInterpretation
MAE\(\frac{1}{n}\sum|y_i-\hat y_i|\)Average absolute miss in target units.
RMSE\(\sqrt{\frac{1}{n}\sum(y_i-\hat y_i)^2}\)Emphasizes larger misses.
\(R^2\)\(1-\frac{SS_{res}}{SS_{tot}}\)Improvement over predicting the target mean.

Intuition: MAE asks, “How far away are predictions on average?” RMSE asks the same question but makes large misses hurt more. \(R^2\) compares the tree with the simple baseline that predicts the training-target mean for every row.

Inspect residuals by target range and important groups. One average metric can hide systematically poor regions.

Report variability, not only the mean

Cross-validation results should include fold-to-fold variation:

\[ \text{mean score}\ \pm\ \text{standard deviation across folds} \]

Intuition: two models with the same average may differ greatly in stability. A tree whose score changes sharply across folds may be sensitive to the sample.

Previous: Regression TreesNext: Practice