Figma — Design Platform & Growth — A/B Test & Causal Inference Questions

Role context: Data Scientist, embedded across Product / Marketing / Finance / Platform · Est. study time: 75 min · 7 questions

Experimentation in this domain

Experimentation at Figma has one defining feature, and almost every question in this section comes back to it: the unit of randomization is a team, not a person.

The reason is the product. Figma is multiplayer. If you show a new comment button to one person in a file and not to their teammate, the two of them are working in the same document with different interfaces. That is confusing for them and it is fatal for you, because the treated person's behavior changes the control person's experience. Figma's own experiments randomize at the Professional and Organization team level so collaborators get a consistent experience, and Figma states the consequence directly: they need larger test groups than typical because of correlated teammate behavior.

That sentence contains the whole methodology of this role. Here is what follows from it:

  • Interference is the default, not an edge case. People inside a team affect each other. Randomizing by user violates the assumption that one person's treatment does not affect another person's outcome, and it biases effects toward zero while making your p-values lie.
  • Clustering destroys power, and you have to quantify how much. Teammates behave alike, so 10 people in a team carry much less information than 10 independent people. The design effect is how you put a number on that, and it routinely means you need two to three times more users than a naive calculation says.
  • The effects you are chasing are small. Figma's published wins include a share modal redesign that moved invites by about 2%. Detecting 2% while paying a cluster penalty is genuinely hard, which is why variance reduction is not a nice-to-have here. CUPED is what makes the test possible at all.
  • Team sizes are wildly unequal. Teams range from 2 people to thousands. A handful of enormous orgs landing in one arm can decide your result by themselves. Stratify or cap, or accept that your estimate is a lottery.
  • The metrics are skewed and they are ratios. Comments per file, response time, edits per user: all heavy-tailed, all with random denominators. Means and naive variances both mislead.
  • Some of the biggest questions cannot be randomized at all. You cannot randomize a procurement decision, an enterprise price change, or whether a team "collaborates." For those you fall back to quasi-experiments, and the sharpest tool available is that Figma already randomizes nudges, which turns a prompt into an instrument.

This role is a broad product and growth role, so the set below covers the full experiment-design core (design, power, test selection, and reading a result) and then weights depth toward variance reduction, causal inference without a clean experiment, and multiple testing, which are the three places this job actually gets hard.

Questions (7)