Distributions: Why User Data Lies
By the end of this lesson, you should be able to: say why user metrics are almost never symmetric, explain what the mean is hiding without reaching for a formula, know which summary to use for which question, and give a better answer than "drop the outliers" when somebody suggests it.
The average user doesn't exist
Basket had 36,781 active users over the eleven weeks. The average one placed 1.94 orders.
Now go looking for that person.

The median user placed 1 order. The mean sits at 1.94, which isn't the middle of anything. 61.2% of Basket users placed fewer orders than the average user. Three in five of your customers are below average, and nothing has gone wrong.
It gets worse the further you look. Here are three ordinary Basket metrics:
| Metric | Median | Mean | 90th percentile | 99th percentile | Max |
|---|---|---|---|---|---|
| Orders | 1 | 1.94 | 5 | 13 | 27 |
| GMV | $52 | $104 | $284 | $711 | $1,751 |
| Active days | 4 | 6.5 | 16 | 38 | 66 |
Look at active days. The mean is 63% higher than the median. If somebody asks how often a Basket user shows up and you answer "about a week in eleven weeks", you have described a person at roughly the 70th percentile and called them typical.
And 33.4% of these users, nearly one in three, placed no order at all. They searched, they browsed, they left. They are in every average you compute.
Why this shape keeps happening
This isn't bad luck with one dataset. Almost every user-level metric you'll ever meet comes out this shape, and there are two reasons.
It's bounded at zero and unbounded above. Nobody places negative orders. But somebody can place 27. A distribution that can't go left and can go a long way right has to lean right.
Engagement is multiplicative, not additive. To place a lot of orders you need several things to be true at once. You have to remember the app exists, have a reason to shop, find the delivery window convenient, have the budget, and not have a competitor's promotion in your inbox that morning. Each of those is a probability, and they multiply rather than add. Multiply enough small factors together and you get a long right tail. That's a mathematical near-certainty, not a quirk of Basket.
This is why "the distribution is roughly normal" is a claim you should never make about a user metric without looking. It almost never is, and the one time you assume it will be the time it costs you.
When an interviewer gives you an average, ask for the median before you do anything else. It costs one sentence and it's the fastest way to show you know user data is skewed. If the two are far apart, everything downstream changes.