For regression, each tree predicts a number and the forest averages those numbers.
A regression tree creates piecewise-constant predictions. A forest averages many such step functions, usually making the result less sensitive to any one partition.
House-price example
Prediction question
Four trees predict 42, 48, 51, and 59 lakh for the same house. What should the forest report?
Intuition: every tree gets one equal-weight voice. The forest does not choose the tree with the smallest training error for this row.
What the forest can and cannot do
| It can | It cannot guarantee |
|---|---|
| Capture nonlinear interactions without polynomial features. | A smooth linear trend or smooth derivative everywhere. |
| Reduce sensitivity to one tree's partition. | Reliable extrapolation beyond the training range. |
| Model mixed relationships after suitable data preparation. | Freedom from leakage, noisy targets, or biased data. |
Extrapolation intuition: outside the observed feature range, leaves tend to reuse values learned inside the training range. A forest is excellent at interpolation but should be treated carefully beyond the data it has seen.
Regression metrics and uncertainty clues
Metric question
If one business error is twice as costly as another, is RMSE alone enough to choose a forest?
Use MAE for typical absolute error, RMSE when large misses deserve extra penalty, and \(R^2\) as a comparison with the target-mean baseline. The spread of tree predictions can be a useful diagnostic of model disagreement, but it is not automatically a calibrated prediction interval.