Plaid — Network Health & ML Products — A/B Test & Causal Inference Questions

Role context: Senior Data Scientist, Embedded Insights (central ML team, first data scientist) · Est. study time: 75 min · 7 questions

Experimentation in this domain

A central ML team on a data network runs three kinds of studies, and each has its own traps.

  • Product experiments on Link. Users can retry and can link in several apps, so the unit is the user, not the session or the attempt. Customers' integrations differ, so results need to be checked per customer and per platform, and a sample ratio mismatch in one integration is common.
  • Model evaluation. Two models are usually scored on the same data, so comparisons are paired. Gold sets are sampled by category, so they need weights. Credit labels exist only for approved borrowers, which biases any comparison of credit models.
  • Network launches. Features like automatic repair roll out by eligibility or in waves across institutions, with no holdout. Difference-in-differences and event studies do the work, and staggered timing has its own pitfalls.
  • Monitoring at scale. Thousands of institutions are checked every day. Without false-discovery control and honest variance for small institutions, alerts become noise.

This role weights model evaluation (paired tests, gold-set sizing, selection-biased labels), monitoring at scale (multiple testing), and causal estimates for network launches, plus the core design of Link experiments. It skips heavy variance-reduction work, since Link traffic is large and the hard problems here are bias, not noise.

Questions (7)