Amazon — Prime Video — A/B Test & Causal Inference Questions
Role context: Data Scientist, Prime Video Personalization & Discovery · Est. study time: 60 min · 5 questions
How experimentation works for a recommender
Testing personalization is subtler than a normal UX A/B:
- The core outcome — a satisfying watch — is fuzzy, so the OEC must be engagement + satisfaction + retention, not clicks/starts or raw hours (Prime Video uses qualified view days and capped view time for exactly this reason).
- Offline ranking metrics (NDCG) screen models but don't always predict online engagement.
- Feedback loops and position bias — recommendations generate the clicks that train the next model, and top items get clicked because they're on top.
- Short-term engagement can diverge from long-term retention, and new titles/users are cold-start.
For the fundamentals — p-values, power, error types, distributions — see the Probability & Statistics section.
Each answer is a coaching walkthrough: a Sample answer (clarify → approach → a simulated back-and-forth → a clear call), then a Deep dive with illustrative example, then a Grading rubric.