Whatnot — Product Analytics — A/B Test & Causal Inference Questions

Role context: Data Scientist, Product Analytics (live commerce marketplace) · Est. study time: 75 min · 7 questions

Experimentation in this domain

Experimentation at Whatnot is harder than at almost any product you have tested before, and the reason is a single structural fact: it is a two-sided marketplace with live auctions and ephemeral inventory.

That produces one dominant problem, and most of this section is about it. Buyer-randomized tests are biased upward. If treated buyers get better show recommendations, they take those shows' finite attention and inventory from control buyers. The measured gap is treatment against a degraded control, not against the status quo. This is not a nuance to mention at the end. It means the number on your scorecard is wrong, in a predictable direction, on essentially every allocation change the feed team ships.

Five more things shape everything below:

  • You cannot randomize a show. A show happens once, live, to whoever is in the room. There is no persistent page to split traffic on and no way to re-run it.
  • The inventory is fixed per show, so treatment changes the price. Routing more bidders into a room raises the clearing price for everyone in it, including control buyers. Interference runs through the price, which is unusual and nasty.
  • The objective is dual. Whatnot ranks on the likelihood a buyer will watch or purchase. Two metrics can move in opposite directions, and you need a decision rule agreed before you look.
  • Everything is heavy-tailed. GMV per seller, price per item, viewers per show. A single $20,000 coin auction can move a category's daily GMV and your p-value.
  • The population is concentrated. A small number of sellers is most of the GMV, so seller-side tests have far less power than the raw count suggests.

The set below covers the interference problem twice (recognizing it, then designing around it), then test selection under heavy tails, reading a dual-objective result, a genuine natural experiment, power under concentration, and heterogeneous effects on the seller pipeline.

Questions (7)