A decision tree is a flowchart learned from data.
Start with the prediction experience: one row enters at the root, answers a sequence of questions, and leaves with a class, probability, or number.
From a business rule to a learned rule
Opening question
For a loan application, which question would you ask first: income, credit score, employment type, or previous default?
A human can propose rules, but a decision tree learns questions from labelled examples. It searches for the question that produces the most useful separation of the target.
This loan dataset is deliberately small and synthetic. It teaches mechanics, not a deployable lending policy. Real lending requires legal, fairness, causal, calibration, and monitoring review.
Anatomy of a tree
| Part | Meaning |
|---|---|
| Root | The first question, seen by every row. |
| Internal node | A later question seen by a subset of rows. |
| Branch | One possible answer to a question. |
| Leaf | A final prediction after no more questions are asked. |
| Depth | Number of splits along the longest root-to-leaf path. |
A path is a conjunction of rules: first condition AND second condition AND so on.
A tree partitions feature space
Geometry question
If every numeric node asks one question such as \(x_j\le t\), what shape will its decision boundary have?
One numeric split creates an axis-aligned cut. Repeating these cuts creates rectangular regions. The overall boundary can be nonlinear even though every individual question is simple.
The top horizontal cut is the root. Only rows below it encounter the vertical second cut.
Intuition: a leaf is a region defined by all the answers along its path.
Classification and regression use the same structure
| Question | Classification tree | Regression tree |
|---|---|---|
| Target | A category such as Approve/Decline | A number such as house price |
| Split quality | Entropy or Gini reduction | Squared-error or variance reduction |
| Leaf output | Majority class and class proportions | Usually the mean target value |
| Prediction | Follow a path to a class leaf | Follow a path to a numeric leaf |
What training must learn
- Which feature should this node inspect?
- For a numeric feature, which threshold should it use?
- Which rows move into each child?
- Should each child split again or become a leaf?
- What class, probability, or number should each leaf predict?
The next page supplies the missing measurement: impurity.