Plaid — Network Health & ML Products — Product Case Questions

Role context: Senior Data Scientist, Embedded Insights (central ML team, first data scientist) · Est. study time: 60 min · 5 questions

How to approach product cases here

Every case is the same chain: understand what the network, the partner team or the customer needs, turn it into a data problem, pick the metric or method, name the bias and the trade-off, and land on a decision.

Three facts about Plaid's network and ML layer sit under most cases:

  • It's a network. The same person, account and institution appear across many apps. That creates real network effects (faster linking for returning users, repairs through other apps) and also means changes ripple across apps and averages hide institution-level problems.
  • Models are judged in someone else's metric. A central team's model creates value in a partner team's product or a customer's decision. Offline scores are evidence, not value.
  • Labels are slow and biased. Credit labels take a year and exist only for approved borrowers; categories are labeled on samples; fraud labels come from customers.

The metrics that matter: healthy active Items, Link conversion, refresh success, 90-day survival, repairs, and for models, value at the operating point plus stability, fairness and adoption.

Traps specific to this domain:

  • Explaining a network-wide move without checking whether one big institution or a mix shift drives it.
  • Celebrating an offline accuracy gain without weighting it by what customers use.
  • Counting every automatic repair, or every adopter's improvement, as caused by the feature.
  • Measuring a central team by models shipped instead of value delivered.

This role weights Measure success (twice, since defining value for models and for the team itself is the core of being the first data scientist), with one Diagnose, one Launch or not, and one Measure impact. It skips Forecasting.

Questions (5)