An ensemble combines several models to produce a better prediction than any of them alone.
Why Ensembles Work
Different models make different mistakes. When their errors aren't perfectly correlated, combining them cancels some of the errors out — the same reason averaging many opinions often beats a single expert.
Bagging
Bootstrap aggregating trains many copies of a model on random samples of the data and averages them. It mainly reduces variance, making unstable models such as deep decision trees much more reliable. The random forest is the best-known example.
Boosting
Models are trained sequentially, each focusing on the errors of the ensemble so far. Boosting mainly reduces bias and often achieves the best accuracy on tabular data. Examples: AdaBoost, gradient boosting, XGBoost, LightGBM, CatBoost.
Stacking
Several different models (say, a linear model, a random forest and a gradient booster) make predictions, and a meta-model learns how best to combine them. Out-of-fold predictions must be used to train the meta-model to avoid leakage.
Simple Averaging
Averaging the predicted probabilities of a few diverse, well-performing models is an easy and effective ensemble.
Trade-offs
Ensembles are more accurate but slower, larger and harder to explain. Consider whether the gain justifies the added complexity in production.