Skip to content

Multi-Armed Bandits

A lightweight alternative to A/B testing that learns which option works best while still serving users.

Editorial team 1 min read

The multi-armed bandit problem: choose repeatedly between options with unknown payoffs, balancing learning with earning.

The Name

Imagine slot machines ("one-armed bandits") with different payout rates. You want to find the best while losing as little as possible along the way.

Common Strategies

  • Epsilon-greedy: usually pick the best-known option; occasionally explore at random.
  • Upper confidence bound (UCB): favour options with high potential given uncertainty.
  • Thompson sampling: choose options in proportion to the probability that each is best — effective and widely used.

Versus A/B Testing

A/B tests split traffic evenly and decide at the end, maximising learning. Bandits shift traffic to winners during the experiment, reducing the cost of showing worse options — but give less precise estimates.

Good Uses

  • Headline and creative optimisation.
  • Recommendations and promotions.
  • Situations with many options or short-lived content.

Contextual Bandits

Use information about the user or situation to choose, personalising decisions.

Cautions

Bandits can lock onto early winners if rewards change over time; use methods that handle non-stationary rewards.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026