Part 1

A decision tree is a flowchart learned from data.

Start with the prediction experience: one row enters at the root, answers a sequence of questions, and leaves with a class, probability, or number.

From a business rule to a learned rule

Opening question

For a loan application, which question would you ask first: income, credit score, employment type, or previous default?

A human can propose rules, but a decision tree learns questions from labelled examples. It searches for the question that produces the most useful separation of the target.

CreditScore ≤ 622.5? ├── yes → predict Decline └── no → PreviousDefault? ├── yes → predict Decline └── no → ask another question

This loan dataset is deliberately small and synthetic. It teaches mechanics, not a deployable lending policy. Real lending requires legal, fairness, causal, calibration, and monitoring review.

Anatomy of a tree

PartMeaning
RootThe first question, seen by every row.
Internal nodeA later question seen by a subset of rows.
BranchOne possible answer to a question.
LeafA final prediction after no more questions are asked.
DepthNumber of splits along the longest root-to-leaf path.
Credit score ≤ 622.5? root yesno Decline Default before? leafinternal node Decline Approve

A path is a conjunction of rules: first condition AND second condition AND so on.

A tree partitions feature space

Geometry question

If every numeric node asks one question such as \(x_j\le t\), what shape will its decision boundary have?

One numeric split creates an axis-aligned cut. Repeating these cuts creates rectangular regions. The overall boundary can be nonlinear even though every individual question is simple.

Feature 1Feature 2

The top horizontal cut is the root. Only rows below it encounter the vertical second cut.

\[ R_1=\{x:x_2\le t_2\ \text{and}\ x_1\le t_1\} \]

Intuition: a leaf is a region defined by all the answers along its path.

Classification and regression use the same structure

QuestionClassification treeRegression tree
TargetA category such as Approve/DeclineA number such as house price
Split qualityEntropy or Gini reductionSquared-error or variance reduction
Leaf outputMajority class and class proportionsUsually the mean target value
PredictionFollow a path to a class leafFollow a path to a numeric leaf

What training must learn

  1. Which feature should this node inspect?
  2. For a numeric feature, which threshold should it use?
  3. Which rows move into each child?
  4. Should each child split again or become a leaf?
  5. What class, probability, or number should each leaf predict?

The next page supplies the missing measurement: impurity.

Previous: OverviewNext: Impurity