Skip to content

Fix #1060: make the uplift multiplier bootstrap seedable - #1061

Open
Hari20032005 wants to merge 1 commit into
py-why:mainfrom
Hari20032005:fix/uplift-bootstrap-random-state
Open

Hari20032005 wants to merge 1 commit into
py-why:mainfrom
Hari20032005:fix/uplift-bootstrap-random-state

Conversation

@Hari20032005

Copy link
Copy Markdown

Fixes #1060.

calc_uplift drew the multiplier bootstrap behind the uniform confidence bands from the global numpy stream, so DRTester.evaluate_uplift() and evaluate_all() returned different bands on every call for identical inputs — and no caller could pin them.

The change

econml/validate/utils.py:

w = check_random_state(random_state).normal(0, 1, size=(n, n_bootstrap))

random_state is threaded through calc_uplift, evaluate_uplift and evaluate_all, normalized with sklearn.utils.check_random_state — the convention already used in _ortho_learner.py:741 and orf/_ortho_forest.py:234.

evaluate_uplift builds one RandomState and passes that instance into each per-treatment calc_uplift call, rather than passing the raw seed down. That keeps each treatment's multipliers independent — as they are today — while making the whole call reproducible from a single seed. evaluate_all does the same across its qini and toc calls.

The parameter defaults to None, so existing callers see no behavior change.

Effect

Three calc_uplift calls on identical inputs, before:

run 0: coeff=0.2579597804  stderr=0.0137873008  uniform_crit=2.845789  one_side_crit=2.633597
run 1: coeff=0.2579597804  stderr=0.0137873008  uniform_crit=2.808732  one_side_crit=2.609976
run 2: coeff=0.2579597804  stderr=0.0137873008  uniform_crit=2.884547  one_side_crit=2.595785

The point estimate and standard error were already deterministic — they never touch the bootstrap — which isolates the multiplier draws as the only source of the drift. With a seed passed, all three lines are identical.

Tests

test_uplift_random_state in econml/tests/test_drtester.py asserts that:

  • the same seed gives identical bands,
  • a different seed gives different bands, so the test cannot be satisfied by hard-coding a constant,
  • the point estimates are unchanged across seeds,
  • evaluate_all is reproducible end to end.

econml/tests/test_drtester.py is 6 passed locally on Python 3.13 (numpy 2.5.2, scikit-learn 1.9.0), and ruff check is clean on the three changed files.

🤖 Generated with Claude Code

calc_uplift drew its multiplier bootstrap from the global numpy stream, so
DRTester.evaluate_uplift() and evaluate_all() returned different uniform
confidence bands on every call for identical inputs, with no way for a
caller to pin them.

Thread a random_state through calc_uplift, evaluate_uplift and
evaluate_all, normalized with sklearn's check_random_state. evaluate_uplift
builds one RandomState and shares it across the per-treatment calls, so each
treatment still draws independent multipliers while the whole call is
reproducible from a single seed. Defaults to None, so existing callers are
unaffected.

Add a regression test.

Signed-off-by: Hari20032005 <hariharan944212005@gmail.com>
@Hari20032005
Hari20032005 force-pushed the fix/uplift-bootstrap-random-state branch from c804626 to 1452617 Compare September 4, 2026 07:18

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DRTester uplift confidence bands are not reproducible: multiplier bootstrap uses the global numpy stream

1 participant