OpenAI — Integrity Measurement — A/B Test & Causal Inference Questions

Role context: Data Scientist, Integrity Measurement (Applied Foundations) · Est. study time: 45 min · 3 questions

Where this role meets A/B testing

This is a measurement role, not a growth-experimentation role — but building metrics that can carry a goal or an A/B test when prevalence itself cannot is squarely part of it. In practice that means you show up to experiments in three ways, and the three questions below map to them one-to-one:

  • You design the metric for a safety intervention's own A/B test, when the harm is too rare for prevalence to move.
  • You own the safety guardrail on other teams' experiments (a ranking or product change that shouldn't quietly increase harmful exposure).
  • You read an enforcement experiment and make the ship call, balancing harm reduction against over-blocking good users.

You don't need to be the world's expert on CUPED or sequential testing for this role, but you do need to speak the language of experimentation fluently enough to collaborate with the teams that run tests, and to design metrics they can actually act on. For the fundamentals — p-values, power, error types — see the Probability & Statistics section.

Each answer is a coaching walkthrough: a Sample answer (clarify the setup → lay out the approach → a simulated back-and-forth with the interviewer → a clear call), then a Deep dive with illustrative example carrying the real math, then a Grading rubric.

Questions (3)