Plaid — Fraud (Plaid Protect) — A/B Test & Causal Inference Questions

Role context: Data Scientist, Fraud Data team (Plaid Protect) · Est. study time: 75 min · 7 questions

Experimentation in this domain

Experiments on a fraud product break most of the assumptions a normal A/B test leans on.

  • The outcome is rare and late. Fraud that gets through might be half a percent of users, and it takes 30 to 90 days to be labeled. Power is about fraud cases, not users, and every readout waits for labels to mature.
  • You can't observe what you block. A blocked user never shows whether they were fraud, so the usual way to measure precision fails. Backtests, audits, and small random allow-throughs fill the gap.
  • Labels are owned by customers. They're late, partial, and defined differently by each customer. Measuring recall well means estimating the fraud nobody labeled.
  • Friction has to be held fixed. A model that flags more people always "catches more fraud". Comparisons only mean something at the same step-up rate.
  • Rings and adaptation. Fraud rings span users in both arms, and fraudsters change tactics once blocked, so effects leak and decay.
  • The most common test is a comparison on the same users. Both models score every user, so tests should be paired, and the users where the models disagree carry almost all the information.

This fraud-analytics role weights model comparison (paired tests and live champion/challenger), rare-event power, label-quality measurement, and interference through rings, plus no-holdout causal estimates for customer go-lives. It skips variance reduction with pre-period covariates, which helps less when the outcome is a rare event most users never have.

Questions (7)