Bagging made trees independent. Boosting makes them collaborate in sequence.
Students already know how Random Forest averages many independently trained trees. We now change one design choice: after each tree, inspect the current ensemble and give the next tree a new correction problem. Gradient Boosting appears naturally once we ask how to define that correction for regression and classification.
Opening thought
Would you rather build one extremely complicated model, or repeatedly improve a simple model by fixing what it still gets wrong?
The whole story in one equation
The current model \(F_{m-1}\) already knows something. A new weak learner \(h_m\) supplies a correction. The learning rate \(\eta\) controls how much of that correction enters the ensemble.
Read it as a sentence: new prediction equals old prediction plus a cautious correction.
Teaching order
From bagging to boosting
Begin with familiar Random Forest ideas, change one assumption, and discover sequential correction.
A complete regression story
Start from the mean and add two shallow trees using a checked four-row example.
Why “gradient”?
Move from ordinary residuals to negative gradients and arbitrary differentiable losses.
A classification story
Follow raw scores, sigmoid probabilities, log loss, and a correction tree by hand.
The complete algorithm
Initialization, pseudo-residuals, tree regions, leaf values, learning rate, and model behavior.
What this foundation deliberately postpones
AdaBoost's explicit sample reweighting and weighted voting, followed by XGBoost, LightGBM, and CatBoost, belong after this foundation. First make the shared boosting idea and Gradient Boosting's loss-based corrections completely stable.