Part 9

Trees need less numerical preprocessing, but they still need careful data engineering.

A reliable tree workflow protects the validation boundary, handles categories and missingness deliberately, inspects leaf support, and treats interpretation as evidence rather than causality.

End-to-end workflow

Workflow question

If trees do not usually need scaling, does that mean preprocessing cannot leak information?

Define target and costsSplit by real deployment unitFit preprocessing on trainTune with validationTest and monitor
  1. Define the prediction time, target window, and cost of each mistake.
  2. Split data by person, account, location, or time as deployment requires.
  3. Fit imputation and encoding only on training folds.
  4. Build a simple baseline and tune tree complexity.
  5. Evaluate metrics, calibration, fairness, stability, and latency.
  6. Document the final rules and monitor drift after deployment.

Why scaling is usually unnecessary

\[ x_j\le t \]

If income is converted from thousands of rupees to rupees, the ordering of values is unchanged and the equivalent threshold rescales with it.

\[ Income_{thousands}\le42.5\quad\Longleftrightarrow\quad Income_{rupees}\le42{,}500 \]

Intuition: a tree cares about which rows fall on each side of a threshold, not Euclidean distance or gradient size.

Scaling can still be needed elsewhere in a pipeline, such as when a tree is compared or stacked with scale-sensitive models.

Categorical features

StrategyBenefitRisk
One-hot encodingSafe for unordered categories and standard scikit-learn pipelines.Each split isolates one encoded category; high cardinality creates many columns.
Ordinal encodingCompact and suitable for truly ordered categories.Invents an order when used on nominal categories.
Native categorical tree libraryCan search category groupings more directly.Behavior differs across libraries and must be validated.
Target encodingCompact for high cardinality.Severe leakage risk unless learned within each training fold.

Missing values

Missingness question

Can “missing income” itself carry information, and when would using that information be dangerous?

Intuition: treat missingness as a data-generating event that needs investigation, not just an empty cell to fill.

A leakage-safe classification pipeline

numeric_pipe = Pipeline([ ("imputer", SimpleImputer(strategy="median")) ]) categorical_pipe = Pipeline([ ("imputer", SimpleImputer(strategy="most_frequent")), ("onehot", OneHotEncoder(handle_unknown="ignore")) ]) model = Pipeline([ ("prepare", ColumnTransformer([...])) , ("tree", DecisionTreeClassifier( criterion="entropy", random_state=42 )) ])

The complete pipeline is fitted inside each validation fold so imputers and encoders never learn from validation rows.

Interpret one tree carefully

InspectionQuestion to ask
Root and early splitsAre dominant rules plausible, stable, and available at prediction time?
Leaf sample countsAre important predictions supported by enough examples?
Class distributionHow uncertain is each leaf?
Train vs validation behaviorDo the same rules survive unseen data?
Subgroup errorsAre mistakes concentrated in protected or operationally important groups?

Feature importance is not causality

\[ Importance(j)\propto\sum_{t:\,\text{node }t\text{ splits on }j}\frac{N_t}{N}\Delta I_t \]

Intuition: a feature receives credit when the fitted tree uses it to reduce impurity for many rows.

Compare impurity importance with validation-set permutation importance and domain knowledge.

High-cardinality identifiers

Identifier question

Why might customer ID appear to create excellent training splits?

An identifier can isolate very small groups and memorize labels without learning reusable structure. Similar risks occur with post-outcome codes, timestamps that reveal the target period, and category levels seen only once.

Remove pure identifiers unless they have a defensible deployment role. Audit suspiciously strong features before celebrating model performance.

Computational behavior

At one node containing \(n\) rows and \(p\) features, sorting and scanning candidate numerical thresholds is roughly:

\[ O(p\,n\log n) \]

Intuition: each feature is ordered so adjacent values can become candidate boundaries, then the algorithm scans those boundaries. Across the complete tree, total fitting cost also depends on how many nodes grow and whether the tree is balanced. Worst-case analyses can be much higher; optimized implementations reduce practical cost by reusing sorted order and other cached work.

Prediction is simpler because one row follows only one root-to-leaf path:

\[ O(depth)\text{ per sample} \]

Intuition: trees can predict quickly because a row evaluates only the questions on its path, not every node.

Operational checks

Previous: EvaluationNext: Model Choice