You sliced 20 cohorts and one is significant — do you believe it?

Multiple testingHard

Problem. An experiment was flat overall, but you sliced it across ~20 cohorts (country × user-state × app type, as Pinterest's dashboard lets you) and one cohort shows a significant save lift at p < 0.05. Is it a real win in that cohort?

Before you reveal: say your answer out loud, as if you were in the real interview — get your reasoning across clearly first. There is no single correct answer: reading what the interviewer is really after and defending your own thinking is what makes an answer strong.