The gradient tells every prediction how to lower the loss.
Residual fitting looks like a clever regression trick. The gradient reveals the deeper rule: calculate the direction that most quickly reduces the chosen loss, then train a tree to imitate that direction from the input features.
One-row question
If the actual value is 12 and the current prediction is 8, should the prediction move upward or downward?
A derivative is local advice
Consider one row with loss \(L(y,F)\), where \(F\) is the current prediction. The derivative \(\partial L/\partial F\) tells us how the loss changes when the prediction increases slightly.
Positive derivative
Increasing the prediction increases loss, so move the prediction downward.
Negative derivative
Increasing the prediction decreases loss, so move the prediction upward.
That is why we use the negative gradient:
\(r_{im}\) is called a pseudo-residual. It is the desired direction of movement for row \(i\) at boosting round \(m\).
Plain language: for every training row, ask the loss, “Which way should this prediction move?” The answers become the target for the next tree.
Connection question
Why did ordinary residuals appear in our squared-error example?
For squared error, the negative gradient is the residual
Choose the convenient half-squared loss:
Differentiate with respect to the current prediction \(F\):
Reverse the sign:
This is exactly the residual used on the previous page. For the delivery with \(y=12\) and \(F=8\), the pseudo-residual is \(12-8=+4\), so the next learner should push its prediction upward.
Function question
Ordinary gradient descent updates a parameter such as a slope. What is Gradient Boosting updating?
Gradient descent in prediction-function space
In linear regression, we might update coefficients. In Gradient Boosting, the object being improved is the entire prediction function \(F(x)\). For the training rows, the negative gradient gives one desired movement per current prediction.
A tree cannot store one unrelated correction for every possible future row. It must learn a rule from features. We therefore fit a regression tree to the pairs
The tree groups rows with similar correction directions into leaves. A future observation receives the correction associated with the leaf it reaches.
The approximation step
The exact negative-gradient values exist only at the training rows. The tree approximates the mapping from input features to those values.
This is why the algorithm can make corrections for new observations. It learned regions in feature space rather than memorizing a list of row IDs.
Generalization question
If residuals belong specifically to squared loss, how can boosting optimize another loss?
Change the loss, and the correction signal changes
| Problem | Loss idea | Negative-gradient signal | What it emphasizes |
|---|---|---|---|
| Squared-error regression | Penalize squared deviations | \(y-F\) | Large errors strongly |
| Absolute-error regression | Penalize absolute deviations | Direction toward \(y\), usually \(\operatorname{sign}(y-F)\) away from zero | More resistance to large outliers |
| Binary classification with log loss | Penalize wrong probabilities | \(y-p\) on the raw-score scale | Probability errors |
The important generalization: Gradient Boosting does not mean “always fit residuals.” It means “fit the negative gradient of the selected loss.” Residuals are the squared-error special case.
For non-smooth losses such as absolute error, a subgradient or algorithm-specific treatment is used at points where the ordinary derivative is not unique.