Repository navigation
EconML MetaLearner Estimators and Categorical Variables - Error: None of Index are in the Columns #1318
|
I have been trying to use EconML MetaLearner Estimators with DoWhy, but I get the following error message when estimating the effect. This only happens when I have categorical variables as confounders. With continous variables, it all works as expected. ""None of [Index(['Factor2', 'Factor1'], dtype='object')] are in the [columns]"" This is based on simulated data Datanp.random.seed(42) num_rows = 1000 data = { DAGEstimate
Full Error Message`--------------------------------------------------------------------------- File \AppData\Local\Programs\Python\Python312\Lib\site-packages\dowhy\causal_model.py:361, in CausalModel.estimate_effect(self, identified_estimand, method_name, control_value, treatment_value, test_significance, evaluate_effect_strength, confidence_intervals, target_units, effect_modifiers, fit_estimator, method_params) File \AppData\Local\Programs\Python\Python312\Lib\site-packages\dowhy\causal_estimator.py:758, in estimate_effect(data, treatment, outcome, identifier_name, estimator, control_value, treatment_value, target_units, effect_modifiers, fit_estimator, method_params) File \AppData\Local\Programs\Python\Python312\Lib\site-packages\dowhy\causal_estimators\econml.py:244, in Econml.estimate_effect(self, data, treatment_value, control_value, target_units, **_) File \AppData\Local\Programs\Python\Python312\Lib\site-packages\dowhy\causal_estimators\econml.py:327, in Econml.effect(self, df, *args, **kwargs) File \AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\core\frame.py:4108, in DataFrame.getitem(self, key) File \AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\core\indexes\base.py:6200, in Index._get_indexer_strict(self, key, axis_name) File \AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\core\indexes\base.py:6249, in Index._raise_if_missing(self, key, indexer, axis_name) KeyError: "None of [Index(['Factor2', 'Factor1'], dtype='object')] are in the [columns]"` |
Replies: 1 comment
|
It's a DoWhy bug in the EconML wrapper, not something in your data. EconML metalearners take a single Workaround: encode before building the model and pass the dummy columns as confounders. df = pd.get_dummies(data, columns=["Factor1", "Factor2"], drop_first=True, dtype=int)
confounders = [c for c in df.columns if c.startswith(("Factor1_", "Factor2_"))]
model = CausalModel(data=df, treatment="Treatment", outcome="Outcomes", common_causes=confounders)
estimand = model.identify_effect()
estimate = model.estimate_effect(
estimand,
method_name="backdoor.econml.metalearners.SLearner",
target_units="ate",
method_params={"init_params": {"overall_model": LGBMRegressor(n_estimators=500, max_depth=10)},
"fit_params": {}},
)If you keep a DAG instead of |

It's a DoWhy bug in the EconML wrapper, not something in your data. EconML metalearners take a single
X, so DoWhy moves the common causes into the effect modifiers and one-hot encodes them (econml.py#L147).estimate_effectthen passes that encoded frame (Factor1_B,Factor1_C, ...) toeffect(), which selects the original names['Factor1', 'Factor2']from it and raises theKeyError(L244, L327). Numeric columns aren't encoded, which is why continuous confounders work.LinearDMLhits the same error when a categorical column is an effect modifier.Workaround: encode before building the model and pass the dummy columns as confounders.