Watch two small trees repair one prediction.
We will predict delivery time from distance using four training rows. The numbers are intentionally small so every prediction, residual, split, and error can be checked by hand.
Initial-prediction question
If you must predict one delivery time for every distance, which constant minimizes squared error?
Step 1: begin with the mean
| Delivery | Distance \(x\) | Actual time \(y\) |
|---|---|---|
| A | 1 km | 2 min |
| B | 2 km | 8 min |
| C | 3 km | 10 min |
| D | 4 km | 12 min |
With squared error, the best constant is the mean:
The model predicts 8 minutes everywhere. Its residual is actual minus prediction:
| \(x\) | Actual | Prediction \(F_0\) | Residual | Meaning |
|---|---|---|---|---|
| 1 | 2 | 8 | -6 | Prediction is 6 too high |
| 2 | 8 | 8 | 0 | Exactly right |
| 3 | 10 | 8 | +2 | Prediction is 2 too low |
| 4 | 12 | 8 | +4 | Prediction is 4 too low |
The vertical arrows are the corrections the current model needs. At 1 km, the prediction must move down; at 3 and 4 km, it must move up.
Correction-tree question
Can a one-split tree group rows that need similar corrections?
Step 2: fit tree 1 to the residuals
The tree receives the original input \(x\), but its target is no longer the original delivery time \(y\). Its supervised training pairs are now:
The first number in each pair is the distance. The second is the residual that the current model wants corrected. Therefore, this tree learns a function from distance to required correction.
Before calculating
A depth-1 regression tree can make only one split. Which split appears to group similar residuals together?
How CART trains this one-split tree
- Generate candidate splits. Sort the distinct distances and test the midpoint between every adjacent pair:
\[1.5,\qquad2.5,\qquad3.5\]
- Divide the rows. Each candidate creates a left leaf and a right leaf.
- Calculate each leaf's prediction. A regression-tree leaf predicts the mean of the target values that reach it. Here those targets are residuals.
- Calculate split error. Measure the squared distance between every residual and its leaf mean:
\[\operatorname{SSE}_{\text{split}}=\sum_{i\in L}(r_i-\bar r_L)^2+\sum_{i\in R}(r_i-\bar r_R)^2\]
- Choose the smallest SSE. That split groups correction targets most effectively.
Evaluate the candidate \(x\leq1.5\)
The left leaf contains only residual \(-6\), so its mean prediction is \(-6\). The right leaf contains \(0,2,4\), so its prediction is:
Its total training error is:
Compare all possible splits
| Candidate split | Left residuals | Left prediction | Right residuals | Right prediction | Total SSE |
|---|---|---|---|---|---|
| \(x\leq1.5\) | \([-6]\) | \(-6\) | \([0,2,4]\) | \(2\) | 8 |
| \(x\leq2.5\) | \([-6,0]\) | \(-3\) | \([2,4]\) | \(3\) | 20 |
| \(x\leq3.5\) | \([-6,0,2]\) | \(-4/3\) | \([4]\) | \(4\) | \(104/3\approx34.67\) |
The smallest SSE is \(8\), so the trained stump is:
What has the tree learned? For a distance near 1 km, the current prediction should move down by 6. For a distance above 1.5 km, it should move up by 2. The tree predicts a correction, not delivery time itself.
Update question
Should the ensemble accept the tree's entire correction immediately?
Use learning rate \(\eta=0.5\), so the model accepts half of the correction:
| \(x\) | Actual | Old \(F_0\) | Tree correction \(h_1\) | Applied correction | New \(F_1\) | New residual |
|---|---|---|---|---|---|---|
| 1 | 2 | 8 | -6 | -3 | 5 | -3 |
| 2 | 8 | 8 | +2 | +1 | 9 | -1 |
| 3 | 10 | 8 | +2 | +1 | 9 | +1 |
| 4 | 12 | 8 | +2 | +1 | 9 | +3 |
What happened? One weak tree reduced MSE from 14 to 5. It made B temporarily worse, but improved the total objective. Boosting optimizes the overall loss, not every row independently at every round.
Second-round question
Should tree 2 fit the original targets again, or the remaining residuals \((-3,-1,1,3)\)?
Step 3: recalculate, then fit tree 2
The first tree changed the model, so we calculate new residuals. A one-split regression tree now chooses \(x\le2.5\):
The left residual mean is \((-3-1)/2=-2\). The right residual mean is \((1+3)/2=2\). Apply half again:
| \(x\) | Actual | Old \(F_1\) | Tree correction \(h_2\) | New \(F_2\) | Residual after tree 2 |
|---|---|---|---|---|---|
| 1 | 2 | 5 | -2 | 4 | -2 |
| 2 | 8 | 9 | -2 | 8 | 0 |
| 3 | 10 | 9 | +2 | 10 | 0 |
| 4 | 12 | 9 | +2 | 10 | +2 |
Adding shallow step functions creates a more detailed step function. The plotted values match the tables.
What students should notice before seeing any calculus
- The initial model is already a valid prediction.
- Each tree predicts a correction, not the final target by itself.
- The correction target changes after every tree.
- The learning rate shrinks every tree's contribution.
- The final model is the sum of the initial value and all tree contributions.