A full read of a classic landing-page experiment (Udacity "Analyze A/B Test
Results" dataset: 294,478 page views, Jan 2–24 2017, users randomised to the old
or new page with a binary converted outcome). The deliverable isn't the
two-proportion test; it's the checks around it that decide whether that number
can be trusted: a sample-ratio-mismatch gate, a power calculation, a
novelty/primacy check, segment reversals, and a guardrail.
Recommendation: do not ship the new page. Conversion was 12.04% (control) vs 11.88% (treatment), an absolute lift of −0.16 pp (95% CI −0.39 to +0.08), z = −1.31, p = 0.19. No evidence the new page helps, the point estimate is slightly negative, and the test had the traffic to detect a lift of 0.34 pp, a little over half the 0.6 pp the team said it would care about. Keep the old page.
| Step | Result | Read |
|---|---|---|
| Data cleaning | 3,893 rows (1.3%) had group / landing_page disagreeing; 1 duplicate user → 290,584 clean |
reported, not silent |
| Sample ratio mismatch | 145,274 vs 145,310; χ² p = 0.95 | randomisation intact; run before any outcome analysis |
| Power | achieved ~145k/arm → detectable effect ±0.34 pp at 80% power, vs a pre-registered ±0.6 pp of interest | a genuine null, not an underpowered shrug |
| Primary test | −0.16 pp [−0.39, +0.08], z = −1.31, p = 0.19 | not significant; CI comfortably spans 0 |
| Novelty / primacy | cumulative lift settles near −0.16 pp within days and stays flat | no decay or ramp; a longer test won't change the call |
| Segments (3 countries × 7 weekdays) | 3 slices flip sign vs the aggregate; largest is Saturday −0.75 pp (raw p = 0.02), not significant after Bonferroni | no credible subgroup where the new page wins |
| Guardrail | page-delivery error rate 1.31% vs 1.33% (p = 0.56) | the ~1.3% mis-delivery is balanced across arms, so it doesn't bias the comparison |
analysis/charts/ (regenerated by make analysis):
01_srm_check.png: observed vs expected arm sizes02_power_curve.png: required n vs MDE, with the pre-registered and achieved points marked03_primary_effect.png: conversion by arm (95% Wilson CI) and the effect size04_novelty_check.png: daily and cumulative lift across the test window05_segment_effects.png: forest plot of per-segment lift against the aggregate
- SRM gate first. χ² goodness-of-fit of arm sizes against the intended 50/50 split, checked before any conversion analysis: if randomisation is broken, nothing downstream means anything. Flag threshold p < 0.01.
- Power framed as pre-registration. Per-arm sample size for an 80%-power, α = 0.05 two-proportion test via the arcsine (Cohen's h) effect size, and its inverse (smallest detectable lift at the achieved n). Baseline 12.04%, MDE +0.6 pp set before looking at results.
- Primary test. Pooled two-proportion z-test, cross-checked against a χ² contingency test (z² = χ² for a 2×2), with a Wald CI on the absolute lift and Wilson intervals on each arm.
- Novelty / primacy. Daily conversion by arm plus the cumulative lift, to distinguish a decaying novelty bump or a slow-learning ramp from a flat effect.
- Segments. The primary test re-run within each country and weekday; segments with only one arm are dropped; weekday p-values Bonferroni-corrected.
- Guardrail. Page-delivery error rate (wrong page for the assigned arm) by arm (instrumentation only; see caveats).
Udacity "Analyze A/B Test Results" landing-page experiment: one row per page
view, user_id / timestamp / group / landing_page / converted, plus a per-user
country lookup. Both CSVs are committed under data/raw/; src/fetch_data.py
re-downloads them from the
nirupamaprv/Analyze-AB-test-Results
mirror and verifies sha256. data/processed/results.json holds the frozen
headline numbers the regression test checks against.
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
make all # runs src/run_analysis.py, then the test suitemake analysis writes the five charts and data/processed/results.json;
make test runs 19 tests, including an end-to-end check that the committed
headline numbers still reproduce on the real dataset. Verified from a fresh
clone on Python 3.9 and 3.13.
src/experiment.py statistics as plain, tested functions; no I/O, no plots
src/run_analysis.py headless pipeline: 7 steps -> 5 charts + results.json
src/fetch_data.py re-download the raw CSVs (also committed) with sha256 checks
tests/ synthetic-data unit tests per function + a frozen-value
regression test replaying the pipeline on the real data
data/raw/ ab_data.csv, countries.csv (committed)
data/processed/ results.json: frozen headline numbers for the regression test
analysis/charts/ generated figures
- The guardrail is instrumentation-only. This dataset has no engagement or revenue field (time-on-page, bounce, pages/session, order value), so a real ship decision's guardrail can't be computed here, only whether the ~1.3% page-mis-delivery bug lands disproportionately on one arm (it doesn't).
- Segment cuts are exploratory. Country and weekday slices are Bonferroni-corrected and should be read as hypotheses, not findings.
- The null depends on the power framing. This test is well-powered for the +0.6 pp effect the team cared about, which is what makes "no effect" informative rather than inconclusive; a smaller effect of interest would be a different memo.
- No CUPED / regression adjustment. The dataset carries no pre-period covariate to reduce variance with, so the estimator is the plain difference in proportions.