It wins on offline evals but loses online — which do you trust?

Decide from resultsMedium

Problem. A new model beats the current one on your offline eval suite (math, coding, chat quality), but in the live A/B its online helpfulness proxies and retention are flat-to-down. Which do you trust, and what do you do?

Before you reveal: say your answer out loud, as if you were in the real interview — get your reasoning across clearly first. There is no single correct answer: reading what the interviewer is really after and defending your own thinking is what makes an answer strong.