Course outline

Segmentation: Choosing the Cut

By the end of this lesson, you should be able to: explain why the segmentation you choose determines the conclusion you reach, pick a cut from the decision rather than from what's available, avoid manufacturing findings by slicing until something appears, and say what a segment-level result can and can't support.

Four cuts, four different answers

Last lesson decomposed Basket's 1.78-point conversion decline by acquisition channel and found the mix explained 122% of it. Nothing was broken; the population changed.

That was one choice out of many. Here's the identical metric over the identical period, decomposed four ways:

Horizontal paired bars for four segmentations. Only the acquisition channel cut shows a large negative between effect and a positive within effect; country, platform and tenure all show the within effect carrying almost everything.
The same −1.78 pp decline under four segmentations. Only one of them finds the mix.
SegmentationWithin (rates moved)Between (mix moved)Sum
Acquisition channel+0.396−2.176−1.780
Country−1.780+0.000−1.780
Platform−1.870+0.090−1.780
How long they've been a user−1.622−0.159−1.781

Every row closes to the same total, with residuals around 1e-15. All four are correct arithmetic on the same data.

And they say completely different things.

Cut by country and the mix effect is literally zero: no country changed its share, so the story is "every country got worse". Cut by platform and it's the same story. Cut by tenure and 91% of the decline is segments performing worse.

Cut by channel and there's no performance problem at all. Every channel improved.

Three of these four cuts would send you looking for a defect that doesn't exist.

The data can't choose for you

This is worth sitting with, because it's uncomfortable and it's the actual job.

There's no statistical procedure that reads the four rows above and returns "channel is the right cut". Each decomposition is valid. Each closes exactly. The one that happens to be correct is correct because of something outside the data: we know a marketing campaign changed the channel mix in mid-April. The arithmetic didn't tell us that. Lesson 5 did.

So the honest description of what shift-share does: it tests a hypothesis you brought with you. It can't generate one.

That inverts the usual instinct. Most people open a dataset and start slicing to see what turns up. The better order is:

  1. List the things that could plausibly have changed. A release, a campaign, a pricing change, a seasonal shift, a competitor, a market launch.
  2. For each one, name the segmentation that would show it. A release shows up in version. A campaign shows up in channel. A market launch shows up in country.
  3. Then decompose, on those cuts, in that order.

The cut is a hypothesis in disguise. Write down which one you're testing before you run it.

In an interview, say what you'd expect to see before you cut. "If it's the campaign, I'd expect a large between effect on channel and near-zero on country." Predicting the shape first is the difference between testing an idea and fishing, and interviewers can hear it immediately.