What is ensemble learning
The idea behind ensemble learning is simple: instead of relying on a single model, train a group of models and combine their outputs. A single model tends to fail in certain ways, and different models do not make exactly the same mistakes, so combining them tends to be more stable and more accurate overall.
Why it works
From a bias-variance perspective, ensembles can lower variance and sometimes bias as well: averaging many weak models that are similar in distribution but reasonably independent sharply reduces variance, while boosting, which corrects mistakes step by step, reduces bias. The key prerequisite is diversity among base models; if every model makes the same error, ensembling cannot help.
Three core paradigms
Bagging: sample the data with replacement, train many homogeneous models, then vote or average. A representative is Random Forest; training is parallel and mainly reduces variance.
Boosting: train models sequentially, with each new model focused on correcting the previous ones' errors. Representatives include AdaBoost, Gradient Boosting (GBDT), XGBoost, LightGBM and CatBoost; it mainly reduces bias.
Stacking: use the outputs of several heterogeneous models (linear models, tree models, neural networks) as new features and train a meta-learner to combine them, usually with cross-validation to avoid overfitting.
Random Forest and Gradient Boosting
Random Forest additionally samples random subsets of features for each tree, boosting diversity further; it trains fast, needs little tuning and is robust. Gradient Boosting frames boosting as gradient descent in function space, fitting each tree to the residual (more precisely, the negative gradient); it is often more accurate but more prone to overfitting and slower to train. XGBoost and LightGBM, using histograms, approximate splits and regularization, turned it into the workhorse for tabular data in competitions and industry.
When to use it and what to watch
On structured tabular data, gradient boosting methods are often the first choice; for noisy settings that value robustness, Random Forest is more forgiving. Points to watch: more base models is not always better (diminishing returns and rising cost); boosting is sensitive to outliers and noise; Stacking can overfit the validation set without cross-validation; and production deployment must trade off inference latency against interpretability.
Sources: compiled from publicly available academic papers and industry materials, including Breiman's work on Bagging and Random Forests, Freund and Schapire's work on AdaBoost, Friedman's paper on gradient boosting, and the official documentation of open-source projects such as XGBoost and LightGBM.