A safety filter should cut a rare harm — but prevalence is too rare to move in the test

Metric designHard

Problem. A product team is A/B testing a new filter meant to reduce how often users are exposed to a rare policy-violating category. They come to you: prevalence is so low that the per-arm confidence intervals overlap and the result looks "not significant," even though they're pretty sure the filter works. Design a metric that can actually detect the change.

Before you reveal: say your answer out loud, as if you were in the real interview — get your reasoning across clearly first. There is no single correct answer: reading what the interviewer is really after and defending your own thinking is what makes an answer strong.