Churn: Defining It Before Predicting It
By the end of this lesson, you should be able to: explain why churn has to be inferred rather than observed for most products, build the curve that prices every candidate threshold, say why that curve usually won't choose for you, and tie the threshold to the decision it exists to serve.
Nobody tells a grocery app they've left
Subscription products have it easy. Somebody clicks cancel, a row changes, and churn is a fact you can look up.
Basket has no cancel button. A user who is done with Basket does exactly the same thing as a user who is on holiday, or who ordered a big shop last week, or who is busy: nothing. Churn isn't observed, it's inferred from silence, and the whole difficulty is deciding how much silence counts.
Start with how noisy that signal is. Across 275,571 active days from 39,001 Basket users, the gap between one active day and the next:
| Median | Mean | 90th percentile | 99th percentile |
|---|---|---|---|
| 3 days | 6.0 days | 14 days | 42 days |
One user in ten routinely goes two weeks between visits and comes back perfectly happily. Any rule that calls two weeks of silence "churned" will label a large group of ordinary customers as lost.
Seven defensible thresholds, an eleven-fold spread
So pick a threshold. Here's what each reasonable choice reports, on exactly the same users:

| Inactive for | Reported churn | Users called churned |
|---|---|---|
| 7 days | 58.8% | 22,946 |
| 14 days | 37.8% | 14,755 |
| 21 days | 26.2% | 10,203 |
| 28 days | 18.6% | 7,257 |
| 35 days | 13.5% | 5,283 |
| 42 days | 9.8% | 3,829 |
| 56 days | 5.3% | 2,086 |
Eleven times. You could walk into a board meeting and report a churn crisis or a healthy business, with the same query, the same data, and no dishonesty at all.
This is lesson 1 for the third time, and by now the pattern should be automatic: when a metric requires a threshold, the threshold is part of the metric. Say it out loud or your number means nothing.