Segment Effects and Interaction Tests
By the end of this lesson, you should be able to: tell a real heterogeneous effect from one manufactured by searching, run the interaction test that actually licenses the claim, decide when a segment finding is worth acting on, and explain why the honest version of this analysis is mostly decided before launch.
A result that says nothing happened
Basket rebuilt its new-user onboarding and tested it on 64,208 users.
| Overall lift | +3.12% |
| 95% interval | [−0.19%, +6.42%] |
| p | 0.064 |
Just misses. The team files it under "no effect" and moves on.
Now split by how long the user has been signed up, which is what the onboarding is about:
| Group | Users | Lift | 95% interval | p |
|---|---|---|---|---|
| First month | 28,589 | +8.01% | [+2.86%, +13.16%] | 0.002 |
| Everyone else | 35,619 | −0.49% | [−4.79%, +3.82%] | 0.825 |
The onboarding works, strongly, on exactly the people it was designed for. It does nothing for anybody else, which isn't surprising since they'd already been onboarded. The two effects average to something that looks like a null result.

Now the hard part
That was the easy example, because the segment was obvious, named in advance and mechanically sensible. Most segment findings arrive the other way round: somebody scans a breakdown and reports the cell that looks best.
To see what that produces, take Basket's checkout redesign. Its effect is the same for every user by construction, with no heterogeneity anywhere. Slice it 21 ways across country, platform, channel, tenure and past orders:
| True effect, identical for everyone | +3.51% |
| Segments tested | 21 |
| Best-looking segment | country=AU, +9.97% (p = 0.025) |
| Worst-looking segment | pre-orders=0, −3.28% |
| Apparent spread | 13.3 percentage points |
Australia looks like it responds three times better than average. It doesn't. Every user in that experiment got an identical effect, and the entire 13.3-point spread is sampling noise in smaller cells.
A segment table will always produce a best cell and a worst cell. The spread between them isn't evidence of anything, because a spread that size is what noise looks like when you slice a fixed sample twenty-one ways.
Five of those 21 segments came back significant, and that is not a false-positive story. The experiment has a real +3.51% effect, so any segment with enough users detects it. The illusion isn't which segments are significant. It's the variation between them.