
Causal Inference
Measure an effect when you cannot randomise. Twenty-two lessons on one company's data, where the true answer is planted and every method is scored against it.
22 lessons in 6 modules, all running on one product's data and building on each other. Start at the top.
Module 1 · What Causal Means
The number you can never see, where the gap comes from, and which variables to condition on.
- 1Potential Outcomes and the CounterfactualPotential outcomes, the fundamental problem of causal inference, and the difference between ATE, ATT and a naive comparison. One planted 6% effect produces three different correct answers to 'what is the effect'.
- 2Selection BiasWhere selection bias comes from, how to decompose an observed gap into effect plus bias, and what randomisation actually removes. 87% of one observed difference is not the product.
- 3Causal Graphs: Confounders, Mediators, CollidersCausal graphs, and the three structures that decide what to adjust for. Two stories with identical correlation matrices to three decimals need opposite adjustments.
- 4The Back-Door CriterionWhich variables to condition on and which to leave alone. Six defensible adjustment sets on one dataset give six different answers, and nothing in the data says which is right.
Module 2 · Adjusting for What You Can See
Regression, propensity scores, matching and weighting, in increasing order of honesty.
- 5Regression AdjustmentRegression as a causal estimator, what the coefficient rests on, and the two separate jobs a control variable does. Adding four covariates moved one estimate from +60.62% to +17.55%.
- 6Propensity ScoresThe propensity score, why calibration matters more than discrimination, and the diagnostic that catches a bad model. Two models with the same AUC, one leaving an effective sample of three teams.
- 7Matching and Inverse-Probability WeightingMatching and IPW, how to read a balance table, and why perfect covariate balance still does not guarantee the right answer.
- 8Doubly Robust EstimationCombining an outcome model with a treatment model so that either one being right is enough, and why the two chances are not the same size.
Module 3 · Using Time
Before-and-after and why it fails, then the difference-in-differences family and its assumption.
- 9Before-and-After and Interrupted Time SeriesWhy a before-and-after comparison fails, how to decompose what it measures, and interrupted time series done properly. Three quarters of one +4.79% reading arrived on its own.
- 10Difference-in-DifferencesBorrowing a counterfactual from untreated units, three equivalent ways to compute the estimate, and inference when only one unit is treated.
- 11The Parallel Trends AssumptionThe one assumption DiD rests on, the event-study plot, how to test it, and how often that test actually catches a violation.
- 12Staggered RolloutsWhy two-way fixed effects can return the wrong sign when units are treated at different times, and the group-time estimators that fix it.
- 13Fixed EffectsWhat a fixed effect absorbs, what it cannot, and why a time-varying confounder survives every level of dummies you add. Pooled 0.953, city 0.451, city and week 0.247, truth 0.08.
Module 4 · One Treated Unit
Building a control group out of other markets, and deciding whether to believe it.
- 14Synthetic ControlBuilding a comparison unit as a weighted average of untreated ones, why the weights are constrained the way they are, and when the method works.
- 15Inference for Synthetic ControlPlacebo tests across every donor, backdating, and leave-one-out checks: how to decide whether a synthetic control result is real.
- 16Counterfactual ForecastingEstimating an effect with no control group anywhere, by forecasting the counterfactual from the unit's own history, and backtesting before you trust it.
Module 5 · Manufacturing Randomness
Geo holdouts, instruments, encouragement designs and thresholds: where a design beats an adjustment.
- 17Geo Experiments and IncrementalityGeo holdout design, why assignment matters more than analysis, and the gap between attributed and incremental. 127 attributed orders per thousand, 51 actually caused.
- 18Instrumental VariablesInstrumental variables and two-stage least squares, what makes an instrument valid, whose effect you are estimating, and what a weak instrument costs.
- 19Encouragement Designs and LATERandomising a nudge instead of the treatment: intention-to-treat, the complier average causal effect, and the assumption with no diagnostic.
- 20Regression DiscontinuityUsing a threshold somebody else already set, choosing a bandwidth, and the assumption that actually gets violated in practice.
Module 6 · Who, and How Sure
Who the effect is biggest for, how wrong it could be, and choosing a method you can defend.
- 21Heterogeneous Treatment Effects (CATE)S-, T- and X-learners for estimating who the effect is biggest for, validating a model of something you never observe, and what a targeting rule is worth.
- 22Sensitivity Analysis and Choosing a MethodHow strong an unmeasured confounder would have to be to overturn a result, and a decision procedure for choosing among the methods in this track.