Part 7

A regression tree asks threshold questions and predicts a number in each leaf.

The structure is unchanged from classification. What changes is the definition of impurity and the value stored at a leaf.

From class purity to numeric compactness

Regression question

If a node contains house prices rather than class labels, what should “pure” mean?

A useful regression node contains target values that are close together. Squared error measures how spread out they are around the node mean.

\[ \hat y_m=\frac{1}{N_m}\sum_{i\in Q_m}y_i \]

Intuition: every house reaching leaf \(m\) receives the average training price of houses in that leaf.

Why the mean appears

The constant \(c\) minimizing squared error inside a leaf is the mean:

\[ c^*=\arg\min_c\sum_{i\in Q_m}(y_i-c)^2=\bar y_m \]

Intuition: if one number must represent every target in a group and large misses are penalized quadratically, the balancing point is their mean.

A small house-price dataset

Split question

Where would you split house area so that prices within each side are as similar as possible?

Area (sq ft)6007008009001100130015001700
Price (₹L)3034384355657282

The parent mean is:

\[ \bar y=\frac{30+34+38+43+55+65+72+82}{8}=52.375 \]

Parent squared error

\[ SSE(Q)=\sum_{i\in Q}(y_i-\bar y_Q)^2 \]
\[ SSE(parent)=2561.875 \]

Intuition: this is the total squared miss if one number, ₹52.375L, predicts every house before any split.

Try Area ≤ 1000

ChildPricesMeanSSE
Left30, 34, 38, 4336.2592.75
Right55, 65, 72, 8268.50389.00
\[ SSE_{after}=92.75+389.00=481.75 \]
\[ \Delta SSE=2561.875-481.75=2080.125 \]

Intuition: two local averages represent the data far better than one global average.

₹36.25L₹68.50L1000 sq ft Area (sq ft)Price (₹L)

A depth-1 regression tree is a two-step prediction function.

Choose the best regression split

\[ (j^*,t^*)=\arg\min_{j,t}\left[SSE(Q_L(j,t))+SSE(Q_R(j,t))\right] \]

Intuition: try every allowed feature-threshold pair and keep the one whose two local averages leave the smallest total squared error.

Maximizing \(\Delta SSE\) and minimizing child SSE are equivalent because parent SSE is fixed at the current node.

More depth creates a step function

Shape question

Can a standard regression tree produce a smooth rising line between training regions?

deeper tree: more, smaller constant regions

More leaves approximate a curve with finer steps; they do not make each leaf prediction smooth.

Mixed features in regression

A practical house model might use area, age, distance to metro, city zone, and furnishing status. Numeric and encoded categorical candidates compete using the same squared-error objective.

Area ≤ 1000? ├── yes: Furnished? │ ├── yes → predict ₹41L │ └── no → predict ₹34L └── no: DistanceToMetro ≤ 2.5 km? ├── yes → predict ₹76L └── no → predict ₹62L

These values are leaf means, not coefficients.

Regression-tree criteria

CriterionLeaf centerBehavior
Squared errorMeanPenalizes large residuals strongly.
Absolute errorMedianMore robust to extreme target values, but slower to optimize.
Poisson deviancePositive mean structureUseful for non-negative count-like targets under suitable assumptions.

Regression-tree limitations

Extrapolation question

If the largest training house is 1700 sq ft, what trend will a tree use for a new 2500 sq ft house?

A standard tree sends the house to an existing terminal region and returns that leaf's mean. It does not continue the upward trend beyond training data.

Regression trees are strong at interactions and local thresholds, but poor at smooth extrapolation. Plot predictions against important continuous features before deployment.
Previous: PruningNext: Evaluation