A practical tree compares categories and numeric thresholds in one search.
The synthetic loan example combines income, credit score, loan amount, employment type, and previous default. We will calculate the best root question without pretending the example is a real lending policy.
The mixed loan dataset
Feature question
Which feature types need a threshold, and which can create category branches?
| ID | Income (₹k/month) | Credit score | Loan (₹L) | Employment | Previous default | Approved |
|---|---|---|---|---|---|---|
| 1 | 35 | 610 | 12 | Salaried | No | No |
| 2 | 45 | 650 | 15 | Salaried | No | Yes |
| 3 | 28 | 590 | 8 | Self-employed | Yes | No |
| 4 | 70 | 720 | 25 | Salaried | No | Yes |
| 5 | 55 | 680 | 22 | Self-employed | No | Yes |
| 6 | 32 | 640 | 10 | Salaried | No | No |
| 7 | 85 | 750 | 30 | Business | No | Yes |
| 8 | 40 | 620 | 18 | Self-employed | No | No |
| 9 | 60 | 700 | 24 | Salaried | No | Yes |
| 10 | 48 | 660 | 20 | Business | No | Yes |
| 11 | 30 | 625 | 9 | Salaried | No | Yes |
| 12 | 75 | 690 | 35 | Business | Yes | No |
Educational synthetic data. Sensitive attributes are intentionally absent, but real-world fairness cannot be guaranteed merely by omitting them because other variables can act as proxies.
Parent entropy
There are 7 approvals and 5 declines.
Intuition: the labels are close to evenly mixed, so the root has high uncertainty.
Generate numeric thresholds
Threshold question
Why test 622.5 rather than testing every real number between credit scores 620 and 625?
Sort the unique observed values and test their midpoints:
For 620 and 625:
Intuition: any threshold inside that empty gap sends exactly the same rows left and right. The midpoint is a convenient representative.
Work the split: CreditScore ≤ 622.5
| Child | Yes | No | Entropy |
|---|---|---|---|
| Left: score ≤ 622.5 | 0 | 3 | 0.000 |
| Right: score > 622.5 | 7 | 2 | 0.764 |
Intuition: the question isolates three declines perfectly and leaves a mostly-approved group, so it removes substantial uncertainty.
Compare mixed candidates fairly
Root decision
Can information gain compare a numeric threshold with a categorical feature?
Yes. Every candidate ultimately partitions the same parent rows, so we compare their expected child impurity.
| Candidate | Best question | Gain |
|---|---|---|
| Credit score | ≤ 622.5 | 0.407 |
| Income | ≤ ₹42.5k/month | 0.334 |
| Previous default | No / Yes | 0.245 |
| Loan amount | ≤ ₹19L | 0.196 |
| Employment | Business / Salaried / Self-employed | 0.062 |
The root chooses `CreditScore ≤ 622.5` for this sample.
Multiway categories versus binary CART splits
A teaching ID3 tree may create one child per category. CART uses two children, so a categorical split conceptually divides categories into two groups.
Intuition: both are category questions, but one produces many branches and the other always produces exactly two.
What happens next?
The left child is pure and becomes a Decline leaf. The right child still contains 7 approvals and 2 declines, so it recalculates every remaining candidate using only those nine rows.
Intuition: no global feature ranking is applied down the whole tree. The useful question can change from branch to branch.
Fairness and decision consequences
Deployment question
If a split increases validation accuracy, is it automatically acceptable for a lending decision?
No. Predictive usefulness is only one requirement. Real systems must examine legal permissibility, proxy discrimination, different error rates across groups, data provenance, human appeal, calibration, drift, and the harm caused by false approvals and false declines.