Course outline

Reading a Result: P-Values and Confidence Intervals

By the end of this lesson, you should be able to: say what a confidence interval covers and what it doesn't, state a p-value in a sentence that's true, explain why a real effect often fails to reach significance, and read a result in a way that survives someone pushing back.

Two ideas, measured instead of defined

The textbook definitions of these two objects are correct and almost impossible to hold onto. So instead of restating them, run them.

Take Basket's control arm, plant a +5.5% effect on a copy of it, and run the experiment a thousand times with 9,000 users per arm.

Every experiment gets a different answer

True effect+5.5%
Smallest estimate across 1,000 runs−3.50%
Largest estimate+16.60%
Mean estimate+5.471%

The same feature, the same effect, the same sample size. One team would have reported a 3.5% loss and another a 16.6% gain.

The mean lands on +5.471% against a truth of +5.500%, so nothing is broken. The method is unbiased and any single run of it is still not the truth. That distinction is the whole of statistical thinking: you never get to see the effect, only one draw from a distribution centred on it.

Two panels. On the left, forty horizontal confidence intervals from repeated experiments on the same effect, most crossing a dashed vertical line marking the truth and a couple coloured differently because they miss it. On the right, a histogram of measured lifts from experiments with no effect at all, centred on zero and spreading from about minus 9 to plus 10, with dashed lines marking the outer 5%.
Left, forty experiments on one true effect. Right, the range of results produced by nothing at all.