Designing a revenue A/B test when whales dominate

Variance & ratio metricsHard

Problem. You're A/B testing a new store layout and the success metric is bookings per player. A tiny fraction of players (whales) generate most revenue, so the metric is wildly noisy and the test "never reaches significance." How do you design and analyze it?

Before you reveal: say your answer out loud, as if you were in the real interview — get your reasoning across clearly first. There is no single correct answer: reading what the interviewer is really after and defending your own thinking is what makes an answer strong.