Multi-Armed Bandits and Thompson Sampling
By the end of this lesson, you should be able to: explain Thompson sampling in three sentences, say which objective a bandit optimises and which it sacrifices, name the two conditions it needs, and decide between a bandit and an experiment on the question rather than on fashion.
Four promo banners
Meridian tests three promo banners against control. Only one of them does anything:
| Arm | True conversion rate |
|---|---|
| Control | 0.1840 |
| A | 0.1840 |
| B | 0.1840 |
| C | 0.1941 |
An A/B test splits 240,000 users four ways and reads the result at the end. A bandit reallocates as it goes: the more an arm looks like a winner, the more traffic it gets.
Thompson sampling is the version worth knowing, and it's three sentences. Keep a posterior for each arm, exactly the Beta from lesson 1. To serve a user, draw one sample from each arm's posterior and serve whichever arm's sample came out highest. Update that arm's posterior with what happened.
That's it. There's no exploration parameter to tune. An arm that might be best gets traffic in proportion to the probability that it is best, which falls out of the sampling rather than being engineered.
What it bought
Same 240,000 users, one run of each:
| Arm | A/B pulls | A/B conversions | Bandit pulls | Bandit conversions |
|---|---|---|---|---|
| Control | 60,000 | 10,857 | 22,500 | 4,221 |
| A | 60,000 | 10,947 | 34,500 | 6,501 |
| B | 60,000 | 11,097 | 14,000 | 2,607 |
| C | 60,000 | 11,719 | 169,000 | 32,848 |
| Total | 240,000 | 44,620 | 240,000 | 46,177 |
The bandit sent 70% of traffic to the winning arm and earned 1,557 extra conversions during the test.

Averaged over 400 runs:
| Conversions | Picked the winner | SE on a losing arm | |
|---|---|---|---|
| A/B, equal split | 44,742 | 100.0% | 0.00158 |
| Thompson sampling | 46,294 | 99.8% | 0.00489 |
+1,552 conversions, +3.47%. Both find the winner essentially always.