Part 10

A forest can rank features, but a ranking is not a causal explanation.

Interpretation should answer what the fitted model relies on, how stable that evidence is, and whether the feature relationship is trustworthy outside the observed data.

Impurity-based importance

Importance question

If a feature appears near the root of many trees and creates large impurity reductions, what might its importance score look like?

Mean decrease in impurity (MDI) adds the weighted impurity reduction contributed by split nodes and averages across trees:

\[I_j=\frac{1}{B}\sum_{b=1}^{B}\sum_{t\in b:\ feature=j}\frac{N_t}{N}\left(C_t-\frac{N_L}{N_t}C_L-\frac{N_R}{N_t}C_R\right)\]

Intuition: a split counts more when it improves purity, and when many rows pass through it. After normalization, the feature scores typically sum to 1.

Impurity importance can favor continuous or high-cardinality features because they have more candidate split points. Treat it as a model summary, not proof that the feature is intrinsically important.

Permutation importance

Permutation question

If we randomly shuffle one feature in a validation set and the model score drops sharply, what does that suggest?

Permutation importance measures the performance decrease after breaking the relationship between one feature and the rows:

\[\operatorname{PI}_j=\operatorname{score}(X,y)-\operatorname{score}(\operatorname{permute}(X_j),y)\]

Use held-out data or cross-validation so the importance reflects generalization rather than memorization. Repeat permutations to see uncertainty.

CreditIncomeAgeNoise

Credit has the largest validation-score drop when shuffled. Noise has almost no effect. This is a more direct “what happens if the model loses access?” question.

Correlated features complicate importance

Correlation question

If CreditScore and a nearly identical credit-history feature are both available, what happens when only one is permuted?

The forest may recover the same signal from the other feature, so each individual permutation can look small even though the feature group matters. Consider grouped permutation, correlation clustering, or domain review.

Partial dependence, ICE, and SHAP can add views of feature effects, but all remain descriptions of the fitted model under the observed data distribution. Association is not causation.

From global importance to one prediction

Interpretation question

A feature is important globally. Does that automatically mean it caused this particular customer to be predicted as high risk?

Interpretation has several levels. A global ranking asks which features the forest uses across many rows. A partial-dependence or ICE plot asks how predictions change as a feature changes. A local explanation asks why one specific row received its prediction. These answer different questions and should not be treated as interchangeable.

GlobalWhat does the forest use overall?
RelationshipHow do predictions change?
LocalWhy this row?
All of these describe the fitted model. They do not prove that changing a feature in the real world will cause the prediction to change in the same way.

What is SHAP?

Local explanation question

If the average predicted probability is 0.20, which features could move one customer’s prediction to 0.65?

SHAP is a feature-attribution method based on Shapley values from cooperative game theory. Think of the prediction as a total amount of evidence. SHAP starts from a baseline prediction and distributes the difference among the features according to their contribution.

\[f(x)=\underbrace{\mathbb{E}[f(X)]}_{\text{baseline}}+\sum_{j=1}^{p}\underbrace{\phi_j}_{\text{contribution of feature }j}\]

For a binary classifier, we must state what the output means. We might explain the predicted probability of the positive class, or explain a score such as log-odds. The numbers can differ because the output space differs.

Worked example

Suppose the baseline positive-class probability is \(0.20\). For one customer:

  • long tenure contributes \(-0.05\)
  • high balance contributes \(+0.18\)
  • low credit score contributes \(+0.22\)
  • other features together contribute \(+0.10\)
\[0.20-0.05+0.18+0.22+0.10=0.65\]
basetenurebalancecreditevidence builds the prediction

Intuition: SHAP turns “the model predicted 0.65” into a receipt: this is the starting point, these features pushed the prediction up, and these features pushed it down.

Using SHAP with a Random Forest

Reading question

In a SHAP summary plot, what does a red point on the right side usually indicate?

With a fitted forest, shap.TreeExplainer(forest) computes tree-aware contributions. A summary plot shows global patterns across rows; a waterfall plot shows the contributions for one row. For binary classification, select and label the positive-class output explicitly, because SHAP values may be returned separately for each class.

explainer = shap.TreeExplainer(forest); shap_values = explainer.shap_values(X_valid); shap.summary_plot(shap_values[:, :, 1], X_valid, feature_names=feature_names)

In a summary plot, points farther right push the explained output higher; points farther left push it lower. Color commonly represents feature value, so a red point on the right means high values of that feature tend to increase the explained output for those rows.

SHAP explanations depend on the background data and feature-dependence assumptions. With strongly correlated features, credit may be shared across them in a way that is useful for explanation but not a unique causal decomposition.
Previous: TuningNext: Practice