A feature only fires on 3% of queries — how do you not wash out the effect?

Experiment designHard

Problem. You're A/B testing a new weather answer box that only renders on weather-intent queries — about 3% of traffic. On the full population the success metric barely moves and "isn't significant." How do you design and analyze this so the real effect isn't diluted away?

Before you reveal: say your answer out loud, as if you were in the real interview — get your reasoning across clearly first. There is no single correct answer: reading what the interviewer is really after and defending your own thinking is what makes an answer strong.