A regression tree asks threshold questions and predicts a number in each leaf.
The structure is unchanged from classification. What changes is the definition of impurity and the value stored at a leaf.
From class purity to numeric compactness
Regression question
If a node contains house prices rather than class labels, what should “pure” mean?
A useful regression node contains target values that are close together. Squared error measures how spread out they are around the node mean.
Intuition: every house reaching leaf \(m\) receives the average training price of houses in that leaf.
Why the mean appears
The constant \(c\) minimizing squared error inside a leaf is the mean:
Intuition: if one number must represent every target in a group and large misses are penalized quadratically, the balancing point is their mean.
A small house-price dataset
Split question
Where would you split house area so that prices within each side are as similar as possible?
| Area (sq ft) | 600 | 700 | 800 | 900 | 1100 | 1300 | 1500 | 1700 |
|---|---|---|---|---|---|---|---|---|
| Price (₹L) | 30 | 34 | 38 | 43 | 55 | 65 | 72 | 82 |
The parent mean is:
Parent squared error
Intuition: this is the total squared miss if one number, ₹52.375L, predicts every house before any split.
Try Area ≤ 1000
| Child | Prices | Mean | SSE |
|---|---|---|---|
| Left | 30, 34, 38, 43 | 36.25 | 92.75 |
| Right | 55, 65, 72, 82 | 68.50 | 389.00 |
Intuition: two local averages represent the data far better than one global average.
A depth-1 regression tree is a two-step prediction function.
Choose the best regression split
Intuition: try every allowed feature-threshold pair and keep the one whose two local averages leave the smallest total squared error.
Maximizing \(\Delta SSE\) and minimizing child SSE are equivalent because parent SSE is fixed at the current node.
More depth creates a step function
Shape question
Can a standard regression tree produce a smooth rising line between training regions?
More leaves approximate a curve with finer steps; they do not make each leaf prediction smooth.
Mixed features in regression
A practical house model might use area, age, distance to metro, city zone, and furnishing status. Numeric and encoded categorical candidates compete using the same squared-error objective.
These values are leaf means, not coefficients.
Regression-tree criteria
| Criterion | Leaf center | Behavior |
|---|---|---|
| Squared error | Mean | Penalizes large residuals strongly. |
| Absolute error | Median | More robust to extreme target values, but slower to optimize. |
| Poisson deviance | Positive mean structure | Useful for non-negative count-like targets under suitable assumptions. |
Regression-tree limitations
Extrapolation question
If the largest training house is 1700 sq ft, what trend will a tree use for a new 2500 sq ft house?
A standard tree sends the house to an existing terminal region and returns that leaf's mean. It does not continue the upward trend beyond training data.