Crash Courses
Master the concepts behind the answers. Each track runs in order, from the idea it is built on to the version you would actually defend in an interview — start at the top, or pick out the one method you need. Concept links in the practice answers bring you here.
Advanced Experimentation
10 coursesBayesian A/B Testing
Prior to posterior on a conversion rate, the beta-binomial worked end to end, and P(B beats A). The posterior is honest at every moment, but monitoring it daily and stopping at 95% fires 19.7% of the time on two identical arms, because the stopping rule breaks the same way it does for a p-value.
Master it →Expected Loss and the Decision Rule
Turning a posterior into a decision: expected loss, the region of practical equivalence, and why a probability of winning is not a reason to ship. Meridian's rebooking test is 99.4% likely to be better and only 59% likely to clear its own break-even.
Master it →Multi-Armed Bandits and Thompson Sampling
Epsilon-greedy and Thompson sampling, and the trade they actually make. Thompson beat an equal split by 3.47% on the same 240,000 users, and ended with each losing arm estimated three times more noisily, because it stopped sending them traffic.
Master it →Stratified and Blocked Randomisation
Balancing covariates by construction rather than by luck, plus blocking and re-randomisation. Across 4,000 re-randomisations, simple assignment was unbiased and widely scattered; stratifying cut that scatter by 56%, and a single experiment needs the scatter, not the average.
Master it →Cluster Randomisation and the Design Effect
Randomising groups instead of users: the intra-cluster correlation, the design effect, and what your effective sample size really is. An ICC of 0.01 across 3,000 users per city gives a design effect of 31, and a test that ignores it fires on a true null 71.3% of the time.
Master it →Switchback Experiments
Turning the whole system on and off over time when a change touches a shared marketplace. Meridian's surge change truly cuts ETA 6.2%; a user-level A/B measures 3.47% and tightens confidently around the wrong number, while a switchback recovers 7.04%.
Master it →Interleaving for Ranking Evaluation
Blending two rankings into one list so each user compares them directly. Between-session variation on Meridian's data is 21 times the gap between the rankers, which is the setup interleaving exists for, and it still came out level with an A/B test on clicks.
Master it →Long-Term Holdouts and Effect Decay
Measuring what a two-week test cannot see, using a holdout you keep running. The same feature reads +10.2% in week one and +2.5% in week twenty-six, both correct, and nothing in the short readout tells you which one you are looking at.
Master it →Meta-Analysis Across Experiments
Learning from the whole corpus rather than one test: empirical priors, the winner's curse, and a win-rate reality check. Across 220 experiments with known truths, the average significant winner reported +4.60% against a true +2.86%, and shrinking toward the corpus prior cut the error by 43%.
Master it →Building an Experimentation Platform
Metric governance, the guardrail catalogue and the gates between a readout and a launch, which is what an interviewer means by 'how would you build this'. 220 correct experiments against 20 metrics each produce hundreds of false significant readings a year by arithmetic alone.
Master it →Causal Inference
27 coursesPotential Outcomes and the Counterfactual
Potential outcomes, the fundamental problem of causal inference, and the difference between ATE, ATT and a naive comparison. One planted 6% effect produces three different correct answers to 'what is the effect'.
Master it →Causal Impact (Counterfactual Forecasting)
Estimate the effect of a change you couldn't randomize by forecasting what the metric would have done anyway, then reading the gap.
Master it →Selection Bias
Where selection bias comes from, how to decompose an observed gap into effect plus bias, and what randomisation actually removes. 87% of one observed difference is not the product.
Master it →Causal Graphs: Confounders, Mediators, Colliders
Causal graphs, and the three structures that decide what to adjust for. Two stories with identical correlation matrices to three decimals need opposite adjustments.
Master it →The Back-Door Criterion
Which variables to condition on and which to leave alone. Six defensible adjustment sets on one dataset give six different answers, and nothing in the data says which is right.
Master it →Media Mix Modeling (MMM)
Estimate what each marketing channel actually contributes to sales, with carryover and diminishing returns baked in, and calibrate it with experiments.
Master it →Heterogeneous Treatment Effects (CATE)
A flat average effect can hide a change that helps some users and hurts others. Estimate who it helps, then target them.
Master it →Regression Adjustment
Regression as a causal estimator, what the coefficient rests on, and the two separate jobs a control variable does. Adding four covariates moved one estimate from +60.62% to +17.55%.
Master it →Incrementality Testing (Ghost Ads, PSA & Geo Lift)
Measure the conversions an ad actually caused — not the ones that would have happened anyway — using ghost-ad/PSA holdouts and geo lift when you can't cleanly A/B.
Master it →Propensity Scores
The propensity score, why calibration matters more than discrimination, and the diagnostic that catches a bad model. Two models with the same AUC, one leaving an effective sample of three teams.
Master it →Cold-Start Markets
Every causal method you know needs history for the treated unit. A brand-new city has none. Here is what to do instead, starting with the design that avoids the problem entirely.
Master it →Matching and Inverse-Probability Weighting
Matching and IPW, how to read a balance table, and why perfect covariate balance still does not guarantee the right answer.
Master it →Doubly Robust Estimation
Combining an outcome model with a treatment model so that either one being right is enough, and why the two chances are not the same size.
Master it →Before-and-After and Interrupted Time Series
Why a before-and-after comparison fails, how to decompose what it measures, and interrupted time series done properly. Three quarters of one +4.79% reading arrived on its own.
Master it →Difference-in-Differences
Borrowing a counterfactual from untreated units, three equivalent ways to compute the estimate, and inference when only one unit is treated.
Master it →The Parallel Trends Assumption
The one assumption DiD rests on, the event-study plot, how to test it, and how often that test actually catches a violation.
Master it →Staggered Rollouts
Why two-way fixed effects can return the wrong sign when units are treated at different times, and the group-time estimators that fix it.
Master it →Fixed Effects
What a fixed effect absorbs, what it cannot, and why a time-varying confounder survives every level of dummies you add. Pooled 0.953, city 0.451, city and week 0.247, truth 0.08.
Master it →Synthetic Control
Building a comparison unit as a weighted average of untreated ones, why the weights are constrained the way they are, and when the method works.
Master it →Inference for Synthetic Control
Placebo tests across every donor, backdating, and leave-one-out checks: how to decide whether a synthetic control result is real.
Master it →Counterfactual Forecasting
Estimating an effect with no control group anywhere, by forecasting the counterfactual from the unit's own history, and backtesting before you trust it.
Master it →Geo Experiments and Incrementality
Geo holdout design, why assignment matters more than analysis, and the gap between attributed and incremental. 127 attributed orders per thousand, 51 actually caused.
Master it →Instrumental Variables
Instrumental variables and two-stage least squares, what makes an instrument valid, whose effect you are estimating, and what a weak instrument costs.
Master it →Encouragement Designs and LATE
Randomising a nudge instead of the treatment: intention-to-treat, the complier average causal effect, and the assumption with no diagnostic.
Master it →Regression Discontinuity
Using a threshold somebody else already set, choosing a bandwidth, and the assumption that actually gets violated in practice.
Master it →Heterogeneous Treatment Effects (CATE)
S-, T- and X-learners for estimating who the effect is biggest for, validating a model of something you never observe, and what a targeting rule is worth.
Master it →Sensitivity Analysis and Choosing a Method
How strong an unmeasured confounder would have to be to overturn a result, and a decision procedure for choosing among the methods in this track.
Master it →Concepts
6 coursesSurrogate Metrics & the OEC
Pick the one metric you'll decide by before the test starts, and a validated early signal that stands in for a long-term outcome you can't wait months to observe.
Master it →Customer Lifetime Value (LTV)
How much a customer is worth over their whole relationship with a product, not just today, and the number that decides how much you can spend to acquire them.
Master it →Price Elasticity of Demand
How much quantity moves when you change price, and why 'raise the price, raise the revenue' is often wrong.
Master it →Prevalence Estimation (Design-Based Sampling)
How to estimate how much of a rare harm is on a platform, using probability sampling and an unbiased estimator instead of a convenience sample or a model threshold.
Master it →Label-Error Correction (Rogan–Gladen)
An imperfect labeler biases a prevalence estimate, and at low base rates a small false-positive rate can dominate. Here's the one-line correction and how to use it.
Master it →LLM-as-Judge for Measurement
Use a large language model as a scalable labeler for measurement, with the golden-set governance and drift monitoring that make its labels trustworthy.
Master it →Experimentation
20 coursesRandomisation and What an Experiment Proves
What randomisation buys you, and why the two obvious alternatives fail. On a feature whose true effect is +3%, before-and-after says +19% and adopters-versus-non-adopters says +74%. Randomising says +3.45%, and it is the only one of the three that stays right over 200 repeats.
Master it →Pre-Registration and the Design Doc
What to decide before launch and why it has to be beforehand. One experiment with no effect at all has 363 defensible ways to analyse it, five of which return p below 0.05. The pre-registered analysis says p = 0.559, and nothing else is allowed to decide the launch.
Master it →Choosing the Randomisation Unit
Picking the unit to randomise, and the separate decision of which unit to analyse. Splitting by session lets 82.2% of users see both variants; splitting by user but analysing by session pushes the false positive rate to 12.1%. Both mistakes are invisible in the readout.
Master it →The OEC and Guardrail Metrics
Choosing the one metric a decision turns on, and the guardrails that stop it being gamed. Revenue needs 13.8 weeks to resolve a 2% effect where sessions per user needs 0.7, and the most sensitive candidate is disqualified outright for conditioning on an outcome.
Master it →Power and Sample Size
Sizing a test, and turning users into days. Detecting a 10% lift on Basket's checkout takes two days; detecting a 1% lift takes seven months. The gap between those two answers is most of experiment design, and most tests are sized by nobody.
Master it →A/A Tests and Platform Validation
Testing nothing against nothing to check the platform itself. Over 4,000 null experiments a healthy platform declares a winner 5.3% of the time and spreads its p-values evenly; the wrong analysis unit gives 6.7%, and a leaky assignment table 8.7%.
Master it →Sample Ratio Mismatch
The chi-square check on the user counts, and why it invalidates a result outright rather than widening an interval. A shelf that read +8.15% against a true +5.00% was caught by the counts, and three of the six standard balance checks never noticed.
Master it →Experiment Data Quality
Symmetric faults barely matter and asymmetric ones destroy the answer. A logging gap hitting both arms moves the read from +3.51% to +3.86%; the same gap hitting treatment alone turns a real +5.5% into -6.21% with p below 0.0001, and the split check does not notice.
Master it →Reading a Result: P-Values and Confidence Intervals
What the two numbers actually mean, measured rather than defined. A thousand experiments on the same true +5.5% returned estimates from -3.50% to +16.60%; the 95% interval covered the truth 96.3% of the time and the test called it significant in only 39.3%.
Master it →Choosing the Test
Metric shape to test, and why the choice matters less than people think. Five of six tests hold their 5% error rate, including the ones the warnings are about. What breaks a readout is analysing a different unit than you randomised.
Master it →Ratio Metrics and the Delta Method
Standard errors when the denominator is random too. On a CTR of 97,658 clicks over 1,156,407 impressions the naive standard error is 0.93 times the truth, pushing the A/A false positive rate to 6.3%. The delta method matches the bootstrap to three decimals.
Master it →Heavy Tails and Capping
What a long tail costs a standard error, and what capping buys and costs. Basket's top 1% of users hold 17.4% of revenue and 39.9% of the variance; capping at the 99th percentile cuts the standard error 11% and stops measuring 4.6% of the money.
Master it →CUPED and Variance Reduction
Buying precision with a column you already have. Pre-period spend correlates with the test metric at 0.32, and spending that correlation cuts variance by exactly 10.5% without touching the estimate. The derivation, the traps, and whether 10% is worth it.
Master it →Sequential Testing and Peeking
What checking every day costs, and the corrections that make it valid. A null experiment watched daily for three weeks crosses 5% significance 19% of the time, because the error rate you designed for only applies if you look once.
Master it →Multiple Testing: FWER and FDR
Bonferroni, Benjamini-Hochberg, and which error rate you actually want to control. Twenty secondary metrics generated with no effect at all produced exactly one winner at p = 0.032, which is precisely what twenty tests at 5% predicts.
Master it →Novelty and Primacy Effects
Telling a fading effect from a real one, and why running longer does not fix it. A banner lifted daily orders 13.15% on day one and 0.55% by week three, and every stopping point along the way says ship. The steady-state effect is nothing.
Master it →Interference and SUTVA
When your control group is not untreated: effect leakage and shared resources. An invite feature with a true +8.00% measures +4.68% because treatment users invite control users, and on a courier network speeding up one order slows another.
Master it →Triggered Analysis and Dilution
Analysing the users who actually saw the feature, and the much worse mistake sitting next to it. A feature reaching 12.2% of assigned users reads +2.41% on everyone and +24.25% on those who triggered, and the gap is exactly the trigger rate.
Master it →Segment Effects and Interaction Tests
The test that licenses a claim about a segment, and why eyeballing segments does not. An onboarding change reads +3.12% overall and +8.01% for first-month users, but a segment search on a perfectly uniform experiment produces an apparent 13.3% spread.
Master it →Shipping the Decision
One experiment read start to finish, applying every check in the track. The split is clean, the primary metric reads +1.54% with p = 0.30, support contacts are up 23.65%, and the feature genuinely works. The correct decision is still not to ship.
Master it →Product Analytics
20 coursesDefine the Metric Before You Move It
Four teams counted daily active users on the same table and got four different answers, the biggest 52% above the smallest. Here's where the gap comes from and how to write a definition nobody can argue with.
Master it →Which Metric Should This Business Care About
Basket's last quarter was either a 72% triumph or a 26% decline, depending on which number you put on the slide. Two tests separate a real North Star from one you can simply buy.
Master it →The Growth Equation
One identity ties acquisition, retention, churn and resurrection together. Split Basket's growth with it and the headline story changes completely: the business isn't growing faster, it's leaking faster.
Master it →Distributions: Why User Data Lies
Basket's average user places 1.94 orders. Three in five users place fewer than that, and the average describes almost nobody. Here's why user data always comes out this shape, and what it breaks.
Master it →Acquisition: Where Users Come From
Basket doubled its weekly signups. It also tripled the size of its worst channel and starved its best one. Here's how to compare channels honestly, starting with the tenure trap that makes every naive comparison wrong.
Master it →Activation: The First Session
Basket's most impressive activation metric predicts 87.0% retention and reaches 2.1% of new users. Here's how to find a real activation moment, and why the best-looking one is almost always a trap.
Master it →Retention Curves
The same Basket cohort retains at 11%, 42% or 70% on day 30, depending which standard definition you use. Then the cohort triangle shows something the aggregate curve hides completely.
Master it →Churn: Defining It Before Predicting It
Seven defensible inactivity thresholds give Basket seven churn rates, from 5% to 59%. The data can price each choice for you, but it can't make it, and pretending otherwise is the most common mistake here.
Master it →Resurrection and Win-Back
Basket brings back nearly twice as many users as it acquires, and each one is worth half as much. Resurrection is the largest term nobody budgets for, and the only one measured without a denominator.
Master it →Funnel Analysis and Drop-off
Basket's worst funnel step loses 44% of users and is perfectly healthy. The step worth fixing lost 14 points and hides in a segment you only find by splitting two dimensions at once.
Master it →Additive and Multiplicative Decomposition
Basket's GMV grew $123,483. Two of the three standard ways to split that between users, frequency and basket size don't even add up, and the piece they leave behind is 17% of the change.
Master it →Mix Shift and Simpson's Paradox
Every one of Basket's five acquisition channels improved its conversion rate. The blended rate fell 1.78 points. Nothing is broken, nobody made a mistake, and the dashboard is red.
Master it →Segmentation: Choosing the Cut
The same 1.78-point decline, decomposed four ways. Three of them say the segments got worse. One says nothing got worse at all. Every one closes exactly, and the data won't tell you which is right.
Master it →Forecasting and Baselines
Basket's Saturday runs 37% above its Tuesday. Until you know what a number should have been, you can't say it fell, and a one-day drop under 5% here means nothing at all.
Master it →Sizing the Opportunity
Basket has two problems worth $79k and $76k a week. They are the same size and they are not remotely the same opportunity, and knowing why is the difference between an analyst and someone who gets listened to.
Master it →Data Quality: Doubt the Instrument First
Basket lost 72% of its iOS search events for three weeks and nobody noticed, because the number still went up. The diagnostic that catches it takes one query and works even when nothing looks wrong.
Master it →Why Did It Move?
Basket's checkout conversion fell 4.5 points. The answer isn't one thing, it is three, they pull in different directions, and running the steps in the wrong order gets you the wrong answer twice.
Master it →Attribution Models
Five standard attribution models, scored against a known truth. They disagree with each other by under one point and they are all wrong by fifty. The argument the industry has been having isn't the one that matters.
Master it →Monetization: LTV, Payback and the Formula That Lies
The textbook LTV formula says a Basket user is worth $121 for life. They hit $119 in eight weeks and were still spending $21 a fortnight. The formula isn't slightly off, it is structurally wrong.
Master it →Growth Loops and Virality
Basket's viral coefficient is 0.066, which sounds like failure and is worth 7% free growth. It also varies twelve-fold by acquisition channel, which is why the April campaign halved it.
Master it →Statistics
27 coursesPopulation and Sample
Parameter versus statistic, and where sampling error comes from. Pantry's average order value is computed from every order in the week, by a query with no bug in it, and still misses the true number by 24 cents.
Master it →Entity Resolution & Probabilistic Record Linkage
Decide which identifiers belong to the same person — deterministic vs probabilistic matching, match scores, the base-rate trap, and turning pairwise links into clean clusters.
Master it →Describing a Distribution
Centre, spread and shape: mean versus median, variance and standard deviation, quantiles, and when the 68-95-99.7 rule does not hold. On Pantry's order values it covers 89.3%, not 68%.
Master it →Probability Rules
The axioms, unions and intersections, independence, complements, and counting. Worked on a rare checkout bug that still has a 91.9% chance of hitting somebody in any given week.
Master it →Conditional Probability
Conditional probability and the law of total probability, and why P(A given B) and P(B given A) answer different questions. Two such numbers about the same 337 Pantry users differ by a factor of 63.
Master it →Bayes' Theorem
Bayes' theorem, the base-rate trap, and stacking evidence with odds. A filter that catches 94% of fraud and clears 97% of honest accounts is still wrong about 91% of the accounts it flags.
Master it →Random Variables
PMF, PDF and CDF, expectation and variance, and the linearity rule that holds without conditions. Expectation always adds; variance only adds when the parts are independent.
Master it →Joint Distributions and Correlation
Joint, marginal and conditional distributions, covariance, correlation and R-squared, and why uncorrelated is not independent. Two Pantry columns that determine each other exactly have a correlation of 0.001.
Master it →Discrete Distributions
Bernoulli, binomial, geometric, negative binomial and Poisson, and the variance-to-mean check that tells them apart. Assuming Poisson for Pantry's session counts understates the tail by a factor of 16.
Master it →Continuous Distributions
Uniform, exponential and normal, standardising and z-scores, and the memorylessness that makes ten days of waiting worth exactly nothing.
Master it →The Normal Family: Chi-Square, t and F
Where the chi-square, t and F distributions come from and why each one exists. A nominal 95% interval built with 1.96 covers 81% at n=3, and the t distribution is the exact repair.
Master it →Choosing a Distribution
How the families connect, how to pick one, and heavy tails and the lognormal. Pantry's order values fail every normality check; their logarithm passes to four decimals.
Master it →Sampling Distributions and Standard Error
The sampling distribution of a statistic, and the difference between a standard deviation and a standard error. On Pantry they are 158 times apart, and that factor is exactly the square root of n.
Master it →The Central Limit Theorem
What the CLT promises, how large n has to be before it delivers on a skewed metric, and where it fails entirely. Pantry's order values need n=500, not the n=30 of the rule of thumb.
Master it →Estimators and Maximum Likelihood
Bias, variance, consistency and mean squared error, then maximum likelihood. Why a constant that ignores your data can beat the sample mean at small n, and why that stops being true as n grows.
Master it →The Bootstrap
Resampling your own data to get a standard error for a statistic that has no formula, percentile intervals, and the cases where the method breaks.
Master it →Confidence Intervals
What 95% confidence actually means, shown by coverage simulation, z versus t, and why the proportion interval everyone learns under-covers.
Master it →Hypothesis Testing and P-Values
Null and alternative, test statistic, rejection region, and what a p-value is and is not. Under the null a p-value is uniform, and that one fact explains the 5% false positive rate.
Master it →Type I, Type II and Power
The two error types, the two overlapping distributions and the areas under them, and how n, effect size and alpha move those areas. At sample sizes teams actually use, 61% of genuine wins come back non-significant.
Master it →Comparing Two Groups
Two means with the pooled and Welch t-tests, two proportions, and paired data. On unbalanced groups the pooled test fires more than half the time when nothing is happening.
Master it →Chi-Square Tests for Categorical Data
Goodness-of-fit, independence and homogeneity, plus Fisher's exact test and McNemar. Plain chi-square holds its error rate well below the textbook warning, and the usual corrections overshoot.
Master it →Nonparametric and Permutation Tests
Permutation tests and Mann-Whitney, what each one actually tests, and how a rank test and a t-test can both be right and still disagree.
Master it →Bayesian Inference
Prior to posterior, the beta-binomial conjugate, and a credible interval versus a confidence interval. A p-value of 0.16 and a 92% posterior probability that treatment wins are not in conflict.
Master it →An Experiment as a Two-Sample Problem
The sampling distribution of a difference, and where an experiment's standard error comes from. Measuring a difference costs four times the traffic of measuring a level.
Master it →Choosing a Test for an Experiment
Metric shape to test, the CLT argument that licenses a t-test on revenue, and where that argument runs out. The test name barely matters; the row you analyse doubles the error rate.
Master it →Alpha, Power and Sample Size
Alpha, power, MDE and n as one set of four dials that trade against each other, and why the MDE is set by finance rather than by statistics.
Master it →Probability & Statistics: The One-Page Recap
Every formula, distribution and decision rule from the twenty-five lessons, on one page, with the measured numbers that make each one concrete. Built to be read the night before an interview.
Master it →