Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Landing-page A/B test: did the new page lift conversion?

A full read of a classic landing-page experiment (Udacity "Analyze A/B Test Results" dataset: 294,478 page views, Jan 2–24 2017, users randomised to the old or new page with a binary converted outcome). The deliverable isn't the two-proportion test; it's the checks around it that decide whether that number can be trusted: a sample-ratio-mismatch gate, a power calculation, a novelty/primacy check, segment reversals, and a guardrail.

Recommendation: do not ship the new page. Conversion was 12.04% (control) vs 11.88% (treatment), an absolute lift of −0.16 pp (95% CI −0.39 to +0.08), z = −1.31, p = 0.19. No evidence the new page helps, the point estimate is slightly negative, and the test had the traffic to detect a lift of 0.34 pp, a little over half the 0.6 pp the team said it would care about. Keep the old page.

Findings at a glance

Step Result Read
Data cleaning 3,893 rows (1.3%) had group / landing_page disagreeing; 1 duplicate user → 290,584 clean reported, not silent
Sample ratio mismatch 145,274 vs 145,310; χ² p = 0.95 randomisation intact; run before any outcome analysis
Power achieved ~145k/arm → detectable effect ±0.34 pp at 80% power, vs a pre-registered ±0.6 pp of interest a genuine null, not an underpowered shrug
Primary test −0.16 pp [−0.39, +0.08], z = −1.31, p = 0.19 not significant; CI comfortably spans 0
Novelty / primacy cumulative lift settles near −0.16 pp within days and stays flat no decay or ramp; a longer test won't change the call
Segments (3 countries × 7 weekdays) 3 slices flip sign vs the aggregate; largest is Saturday −0.75 pp (raw p = 0.02), not significant after Bonferroni no credible subgroup where the new page wins
Guardrail page-delivery error rate 1.31% vs 1.33% (p = 0.56) the ~1.3% mis-delivery is balanced across arms, so it doesn't bias the comparison

Figures

analysis/charts/ (regenerated by make analysis):

  • 01_srm_check.png: observed vs expected arm sizes
  • 02_power_curve.png: required n vs MDE, with the pre-registered and achieved points marked
  • 03_primary_effect.png: conversion by arm (95% Wilson CI) and the effect size
  • 04_novelty_check.png: daily and cumulative lift across the test window
  • 05_segment_effects.png: forest plot of per-segment lift against the aggregate

Method

  • SRM gate first. χ² goodness-of-fit of arm sizes against the intended 50/50 split, checked before any conversion analysis: if randomisation is broken, nothing downstream means anything. Flag threshold p < 0.01.
  • Power framed as pre-registration. Per-arm sample size for an 80%-power, α = 0.05 two-proportion test via the arcsine (Cohen's h) effect size, and its inverse (smallest detectable lift at the achieved n). Baseline 12.04%, MDE +0.6 pp set before looking at results.
  • Primary test. Pooled two-proportion z-test, cross-checked against a χ² contingency test (z² = χ² for a 2×2), with a Wald CI on the absolute lift and Wilson intervals on each arm.
  • Novelty / primacy. Daily conversion by arm plus the cumulative lift, to distinguish a decaying novelty bump or a slow-learning ramp from a flat effect.
  • Segments. The primary test re-run within each country and weekday; segments with only one arm are dropped; weekday p-values Bonferroni-corrected.
  • Guardrail. Page-delivery error rate (wrong page for the assigned arm) by arm (instrumentation only; see caveats).

Data

Udacity "Analyze A/B Test Results" landing-page experiment: one row per page view, user_id / timestamp / group / landing_page / converted, plus a per-user country lookup. Both CSVs are committed under data/raw/; src/fetch_data.py re-downloads them from the nirupamaprv/Analyze-AB-test-Results mirror and verifies sha256. data/processed/results.json holds the frozen headline numbers the regression test checks against.

Reproduce

python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
make all          # runs src/run_analysis.py, then the test suite

make analysis writes the five charts and data/processed/results.json; make test runs 19 tests, including an end-to-end check that the committed headline numbers still reproduce on the real dataset. Verified from a fresh clone on Python 3.9 and 3.13.

Repo layout

src/experiment.py     statistics as plain, tested functions; no I/O, no plots
src/run_analysis.py   headless pipeline: 7 steps -> 5 charts + results.json
src/fetch_data.py     re-download the raw CSVs (also committed) with sha256 checks
tests/                synthetic-data unit tests per function + a frozen-value
                      regression test replaying the pipeline on the real data
data/raw/             ab_data.csv, countries.csv (committed)
data/processed/       results.json: frozen headline numbers for the regression test
analysis/charts/      generated figures

Caveats

  • The guardrail is instrumentation-only. This dataset has no engagement or revenue field (time-on-page, bounce, pages/session, order value), so a real ship decision's guardrail can't be computed here, only whether the ~1.3% page-mis-delivery bug lands disproportionately on one arm (it doesn't).
  • Segment cuts are exploratory. Country and weekday slices are Bonferroni-corrected and should be read as hypotheses, not findings.
  • The null depends on the power framing. This test is well-powered for the +0.6 pp effect the team cared about, which is what makes "no effect" informative rather than inconclusive; a smaller effect of interest would be a different memo.
  • No CUPED / regression adjustment. The dataset carries no pre-period covariate to reduce variance with, so the estimator is the plain difference in proportions.

About

Landing-page A/B test analysis with the checks that decide whether the result is trustworthy: SRM gate, power, effect-size CI, novelty/primacy, segment reversals, guardrail. Well-powered null (-0.16 pp, p=0.19).

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages