Part 4

A practical tree compares categories and numeric thresholds in one search.

The synthetic loan example combines income, credit score, loan amount, employment type, and previous default. We will calculate the best root question without pretending the example is a real lending policy.

The mixed loan dataset

Feature question

Which feature types need a threshold, and which can create category branches?

IDIncome (₹k/month)Credit scoreLoan (₹L)EmploymentPrevious defaultApproved
13561012SalariedNoNo
24565015SalariedNoYes
3285908Self-employedYesNo
47072025SalariedNoYes
55568022Self-employedNoYes
63264010SalariedNoNo
78575030BusinessNoYes
84062018Self-employedNoNo
96070024SalariedNoYes
104866020BusinessNoYes
11306259SalariedNoYes
127569035BusinessYesNo

Educational synthetic data. Sensitive attributes are intentionally absent, but real-world fairness cannot be guaranteed merely by omitting them because other variables can act as proxies.

Parent entropy

There are 7 approvals and 5 declines.

\[ H(Approved)=-\frac{7}{12}\log_2\frac{7}{12}-\frac{5}{12}\log_2\frac{5}{12}\approx0.980 \]

Intuition: the labels are close to evenly mixed, so the root has high uncertainty.

Generate numeric thresholds

Threshold question

Why test 622.5 rather than testing every real number between credit scores 620 and 625?

Sort the unique observed values and test their midpoints:

\[ t_i=\frac{x_{(i)}+x_{(i+1)}}{2} \]

For 620 and 625:

\[ t=\frac{620+625}{2}=622.5 \]

Intuition: any threshold inside that empty gap sends exactly the same rows left and right. The midpoint is a convenient representative.

Work the split: CreditScore ≤ 622.5

ChildYesNoEntropy
Left: score ≤ 622.5030.000
Right: score > 622.5720.764
\[ H_{after}=\frac{3}{12}(0)+\frac{9}{12}(0.764)=0.573 \]
\[ IG=0.980-0.573=0.407\text{ bits} \]

Intuition: the question isolates three declines perfectly and leaves a mostly-approved group, so it removes substantial uncertainty.

622.5 Credit scoreIncome
Approved Declined Candidate split

Compare mixed candidates fairly

Root decision

Can information gain compare a numeric threshold with a categorical feature?

Yes. Every candidate ultimately partitions the same parent rows, so we compare their expected child impurity.

CandidateBest questionGain
Credit score≤ 622.50.407
Income≤ ₹42.5k/month0.334
Previous defaultNo / Yes0.245
Loan amount≤ ₹19L0.196
EmploymentBusiness / Salaried / Self-employed0.062

The root chooses `CreditScore ≤ 622.5` for this sample.

Multiway categories versus binary CART splits

A teaching ID3 tree may create one child per category. CART uses two children, so a categorical split conceptually divides categories into two groups.

\[ Employment\in\{\text{Salaried},\text{Business}\}\quad\text{versus}\quad Employment\in\{\text{Self-employed}\} \]

Intuition: both are category questions, but one produces many branches and the other always produces exactly two.

Scikit-learn's standard decision trees expect numeric input. One-hot or ordinal encoding changes which binary questions are available; encoding therefore belongs inside the model pipeline.

What happens next?

The left child is pure and becomes a Decline leaf. The right child still contains 7 approvals and 2 declines, so it recalculates every remaining candidate using only those nine rows.

CreditScore ≤ 622.5? ├── yes → Decline [0 Yes, 3 No] └── no → mixed [7 Yes, 2 No] calculate local gains again

Intuition: no global feature ranking is applied down the whole tree. The useful question can change from branch to branch.

Fairness and decision consequences

Deployment question

If a split increases validation accuracy, is it automatically acceptable for a lending decision?

No. Predictive usefulness is only one requirement. Real systems must examine legal permissibility, proxy discrimination, different error rates across groups, data provenance, human appeal, calibration, drift, and the harm caused by false approvals and false declines.

A transparent tree can still encode unfair historical patterns. Interpretability makes a rule visible; it does not make the rule fair.
Previous: Information GainNext: Recursion