Gradient boosting gets described often and explained rarely. Most summaries stop at “it combines many weak models,” which is true and tells you almost nothing about why it works or how to use it well.
The core idea: fit the errors, not the target
Gradient boosting builds models one at a time, sequentially. Each new model doesn’t try to predict the original target directly. It tries to predict the residual errors the previous models got wrong. Add that correction to the running total, repeat, and the combined prediction gradually improves with each round.
Why “gradient” is in the name
Each new model is trained to approximate the negative gradient of the loss function with respect to the current predictions, which is a precise way of saying: it points in the direction that would most reduce the current error. This connects gradient boosting directly to gradient descent, the same underlying optimization idea applied in function space rather than over a model’s parameters directly.
Why this is a form of ensemble method
Gradient boosting is a specific kind of ensemble method, in the boosting family specifically, which trains models sequentially so each one targets the previous ones’ mistakes. That’s distinct from bagging, which trains many models independently on different data samples and averages them. Boosting’s sequential dependency is exactly what lets it specifically target errors the ensemble is still making, rather than just reducing variance through averaging.
Why individual trees are kept weak on purpose
Gradient boosting typically uses shallow decision trees as the base learner, deliberately weak on their own. A single shallow tree underfits badly. Hundreds of them, each correcting the last, combine into something far stronger than any one of them. Using strong, complex trees from the start tends to overfit quickly instead, since each one would already be chasing noise rather than leaving genuine signal for the next round to correct.
The hyperparameters that actually matter
- Learning rate scales down each new model’s contribution. A smaller learning rate needs more rounds but generalizes better; a larger one trains faster but risks overshooting.
- Number of estimators sets how many sequential models get added. Too few underfits; too many, especially combined with a high learning rate, overfits.
- Tree depth controls how complex each individual weak learner is allowed to be, directly trading off underfitting against overfitting at the base-learner level.
Finding the right combination is exactly what hyperparameter tuning methods like grid search or Bayesian optimization are built for, since these three interact rather than tuning cleanly in isolation.
A quick checklist
- Are your base learners deliberately kept shallow and weak, or are they already complex on their own?
- Have you tuned learning rate and number of estimators together, rather than one at a time in isolation?
- Are you monitoring validation performance across boosting rounds to catch overfitting as more trees get added?
- Would a bagging approach suit this problem better, if reducing variance matters more than sequentially correcting bias?
FAQ
Is XGBoost the same thing as gradient boosting?
XGBoost is a specific, heavily optimized implementation of gradient boosting, with added regularization and engineering improvements, not a different algorithm.
Why does gradient boosting often outperform random forests?
It directly targets remaining errors each round rather than averaging independent trees, which can capture more signal when tuned well, though it’s also more prone to overfitting if not regularized carefully.
Does gradient boosting work for both classification and regression?
Yes. The loss function changes to match the task, but the sequential error-correction mechanism is the same underlying idea either way.

