Pinterest — Trust & Safety — Product Case Questions

Role context: Data Scientist, Trust & Safety (prevalence measurement) · Est. study time: 60 min · 5 questions

How to approach product cases here

Every case is one move in the same chain: understand the goal, break it into a data problem, choose the metric or method, name the bias and the trade-off, land on a decision and an action. Walk it out loud and survive the follow-ups.

Five things make Trust & Safety cases at Pinterest distinctive. Internalize these and most questions become mechanical.

1. Prevalence is exposure-weighted. It is the share of views that went to violating content, not the share of content that violates. One Pin seen ten million times outweighs ten thousand Pins seen once. Candidates who define it the content-weighted way have already signalled they have not done the reading.

2. Reports and prevalence are different objects. Reports measure who complained. Prevalence measures what was seen. Pinterest's own published work finds the two correlate weakly, and for the hardest policies the exposed users are often seeking the content and will never report. Never treat a report trend as a harm trend.

3. The score prioritizes; it never labels. You use a risk model to decide what gets reviewed, then reweight by the inverse sampling probability. If the model's output becomes the label, prevalence measures your own thresholds instead of reality, and retuning the model "improves" safety with nothing changing. This is the central trap of the role.

4. Prevalence without a false-positive guardrail is an instruction to delete the platform. The metric falls just as happily when you remove legitimate content. It cannot tell the difference, which is exactly why the counter-metric is mandatory rather than nice to have.

5. Half of every answer is about whether the number can be trusted. Confidence interval, effective sample size, minimum detectable effect, labeler drift, prompt version. Part of this role is advocating for decision quality before a metric reaches executive leadership, and that means being able to say "we cannot detect that" with arithmetic behind it.

The traps: reading a metric move as a platform move when the instrument shifted, accepting a target with no constraint, and quoting an agreement rate for a rare event.

This is a measurement science role, so the five below weight measurement design and estimation (two of the five), plus diagnosis, measuring success without a randomized test, and causal impact. There is deliberately no growth or launch-or-not case: this team is not asked whether to ship a feature, it is asked whether the number is real and what it means. That skip is itself the signal. If you only have time for two, do bc1 and bc2.

Questions (5)