Course outline

Segment Effects and Interaction Tests

By the end of this lesson, you should be able to: tell a real heterogeneous effect from one manufactured by searching, run the interaction test that actually licenses the claim, decide when a segment finding is worth acting on, and explain why the honest version of this analysis is mostly decided before launch.

A result that says nothing happened

Basket rebuilt its new-user onboarding and tested it on 64,208 users.

Overall lift+3.12%
95% interval[−0.19%, +6.42%]
p0.064

Just misses. The team files it under "no effect" and moves on.

Now split by how long the user has been signed up, which is what the onboarding is about:

GroupUsersLift95% intervalp
First month28,589+8.01%[+2.86%, +13.16%]0.002
Everyone else35,619−0.49%[−4.79%, +3.82%]0.825

The onboarding works, strongly, on exactly the people it was designed for. It does nothing for anybody else, which isn't surprising since they'd already been onboarded. The two effects average to something that looks like a null result.

Two panels. On the left, three confidence intervals: overall straddling zero near 3%, first-month users clearly above zero near 8%, and everyone else sitting on zero. On the right, a scatter of 21 segment estimates from a single experiment, spread from about minus 3% to plus 10% around a dashed line marking the one true effect, with five points coloured to show they reached significance.
Left, a real heterogeneous effect. Right, twenty-one segments of an experiment where every single user received the same effect.

Now the hard part

That was the easy example, because the segment was obvious, named in advance and mechanically sensible. Most segment findings arrive the other way round: somebody scans a breakdown and reports the cell that looks best.

To see what that produces, take Basket's checkout redesign. Its effect is the same for every user by construction, with no heterogeneity anywhere. Slice it 21 ways across country, platform, channel, tenure and past orders:

True effect, identical for everyone+3.51%
Segments tested21
Best-looking segmentcountry=AU, +9.97% (p = 0.025)
Worst-looking segmentpre-orders=0, −3.28%
Apparent spread13.3 percentage points

Australia looks like it responds three times better than average. It doesn't. Every user in that experiment got an identical effect, and the entire 13.3-point spread is sampling noise in smaller cells.

A segment table will always produce a best cell and a worst cell. The spread between them isn't evidence of anything, because a spread that size is what noise looks like when you slice a fixed sample twenty-one ways.

Five of those 21 segments came back significant, and that is not a false-positive story. The experiment has a real +3.51% effect, so any segment with enough users detects it. The illusion isn't which segments are significant. It's the variation between them.