Microsoft — Bing — A/B Test & Causal Inference Questions

Role context: Senior Data Scientist, Bing Growth & Experimentation (Microsoft AI) · Est. study time: 70 min · 7 questions

Experimentation at Bing

Bing is one of the homes of trustworthy online experimentation — CUPED and much of Microsoft's experimentation platform were invented here — and the role runs 1000s of controlled experiments. So this section focuses on the experiment-design and causal-inference depth this role lives in — triggered analysis, sizing tiny effects, variance reduction, org-scale error control, and causal impact without an RCT. (For the probability and statistics fundamentals — p-values, error types, distributions — see the Probability & Statistics section.) The questions lean toward Bing's signatures:

  • Triggered analysis & dilution — most features fire on a small slice of queries, so naive all-up analysis washes the effect out.
  • Tiny effects on enormous bases — you size for sub-percent MDEs and lean hard on variance reduction (CUPED).
  • Org-scale error control — 1000s of experiments × dozens of metrics, watched continuously, demands FDR and sequential/always-valid methods.
  • Causal impact without an RCT — default/market changes can't be A/B'd, so you forecast the counterfactual.

Each answer is a coaching walkthrough: a Sample answer (clarify the setup → lay out the approach → a simulated back-and-forth with the interviewer → a clear call), then a Deep dive with illustrative example carrying the real math and charts, then a Grading rubric.

Questions (7)