Is the LLM auto-labeler good enough to ship?

Measurement & estimationHard

Problem. An LLM auto-labeler is proposed to replace or assist human annotation on a task. It scores 95% accuracy on a sample. How do you decide whether — and how — to deploy it?

Before you reveal: say your answer out loud, as if you were in the real interview — get your reasoning across clearly first. There is no single correct answer: reading what the interviewer is really after and defending your own thinking is what makes an answer strong.