Amazon — Prime Video — A/B Test & Causal Inference Questions

Role context: Data Scientist, Prime Video Personalization & Discovery · Est. study time: 60 min · 5 questions

How experimentation works for a recommender

Testing personalization is subtler than a normal UX A/B:

  • The core outcome — a satisfying watch — is fuzzy, so the OEC must be engagement + satisfaction + retention, not clicks/starts or raw hours (Prime Video uses qualified view days and capped view time for exactly this reason).
  • Offline ranking metrics (NDCG) screen models but don't always predict online engagement.
  • Feedback loops and position bias — recommendations generate the clicks that train the next model, and top items get clicked because they're on top.
  • Short-term engagement can diverge from long-term retention, and new titles/users are cold-start.

For the fundamentals — p-values, power, error types, distributions — see the Probability & Statistics section.

Each answer is a coaching walkthrough: a Sample answer (clarify → approach → a simulated back-and-forth → a clear call), then a Deep dive with illustrative example, then a Grading rubric.

Questions (5)