Part 3

The gradient tells every prediction how to lower the loss.

Residual fitting looks like a clever regression trick. The gradient reveals the deeper rule: calculate the direction that most quickly reduces the chosen loss, then train a tree to imitate that direction from the input features.

One-row question

If the actual value is 12 and the current prediction is 8, should the prediction move upward or downward?

A derivative is local advice

Consider one row with loss \(L(y,F)\), where \(F\) is the current prediction. The derivative \(\partial L/\partial F\) tells us how the loss changes when the prediction increases slightly.

Positive derivative

Increasing the prediction increases loss, so move the prediction downward.

Negative derivative

Increasing the prediction decreases loss, so move the prediction upward.

That is why we use the negative gradient:

\[r_{im}=-\left.\frac{\partial L(y_i,F(x_i))}{\partial F(x_i)}\right|_{F=F_{m-1}}\]

\(r_{im}\) is called a pseudo-residual. It is the desired direction of movement for row \(i\) at boosting round \(m\).

Plain language: for every training row, ask the loss, “Which way should this prediction move?” The answers become the target for the next tree.

Connection question

Why did ordinary residuals appear in our squared-error example?

For squared error, the negative gradient is the residual

Choose the convenient half-squared loss:

\[L(y,F)=\frac12(y-F)^2\]

Differentiate with respect to the current prediction \(F\):

\[\frac{\partial L}{\partial F}=\frac12\cdot2(y-F)(-1)=F-y\]

Reverse the sign:

\[-\frac{\partial L}{\partial F}=y-F\]

This is exactly the residual used on the previous page. For the delivery with \(y=12\) and \(F=8\), the pseudo-residual is \(12-8=+4\), so the next learner should push its prediction upward.

If squared loss is written without the factor \(1/2\), the negative gradient is \(2(y-F)\). The factor changes the scale, not the direction, and can be absorbed into the update size.

Function question

Ordinary gradient descent updates a parameter such as a slope. What is Gradient Boosting updating?

Gradient descent in prediction-function space

In linear regression, we might update coefficients. In Gradient Boosting, the object being improved is the entire prediction function \(F(x)\). For the training rows, the negative gradient gives one desired movement per current prediction.

A tree cannot store one unrelated correction for every possible future row. It must learn a rule from features. We therefore fit a regression tree to the pairs

\[(x_i,r_{im})\]

The tree groups rows with similar correction directions into leaves. A future observation receives the correction associated with the leaf it reaches.

current predictionmove toward lower lossprediction Floss

The approximation step

The exact negative-gradient values exist only at the training rows. The tree approximates the mapping from input features to those values.

\[h_m\approx-\nabla_F L\big|_{F_{m-1}}\]

This is why the algorithm can make corrections for new observations. It learned regions in feature space rather than memorizing a list of row IDs.

Generalization question

If residuals belong specifically to squared loss, how can boosting optimize another loss?

Change the loss, and the correction signal changes

ProblemLoss ideaNegative-gradient signalWhat it emphasizes
Squared-error regressionPenalize squared deviations\(y-F\)Large errors strongly
Absolute-error regressionPenalize absolute deviationsDirection toward \(y\), usually \(\operatorname{sign}(y-F)\) away from zeroMore resistance to large outliers
Binary classification with log lossPenalize wrong probabilities\(y-p\) on the raw-score scaleProbability errors

The important generalization: Gradient Boosting does not mean “always fit residuals.” It means “fit the negative gradient of the selected loss.” Residuals are the squared-error special case.

For non-smooth losses such as absolute error, a subgradient or algorithm-specific treatment is used at points where the ordinary derivative is not unique.

Previous: RegressionNext: Classification