How big must the gold set be to compare two categorization models?
Power & sample sizeHardProblem. To compare the current and candidate categorization models, the team reviews a sample of transactions by hand (a gold set). Reviews are expensive. You want to detect a 1-point difference in accuracy weighted by how customers use each category. The team also wants reliable accuracy for rare categories. How many transactions do you need, and how should you sample them?
Before you reveal: say your answer out loud, as if you were in the real interview — get your reasoning across clearly first. There is no single correct answer: reading what the interviewer is really after and defending your own thinking is what makes an answer strong.