Causal ML & Double Machine Learning — From Association to Effect
Hub: correlation is not causation. Defend ATE, ATT, and CATE first, then identification, then estimators. Predictive ML is the wrong tool for treatment decisions; this cluster maps graphs, DML, CATE forests, causal time series, and health ethics.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Stakeholder asks: should we do X?
Prefer
Causal frame (estimand, then identify, then estimate)
A treatment decision changes the world the model was trained on. You need P(Y | do(T)), not a better holdout AUC.
- Name ATE, ATT, or CATE before naming a library.
- Draw a DAG; DML only inherits a valid adjustment set.
- RCTs remain gold-standard identification when they are ethical and feasible.
Alternative
Train a high-AUC model and treat the high-risk group
Prediction answers who looks like yesterday’s treated. It does not answer what happens if you intervene.
- AUC scores association under current assignment.
- Selection (healthier patients already in the pathway) masquerades as effect.
- Intervening shifts covariates and often the outcome mechanism.
Interview order of operations
Estimand first. Estimator last. Jumping to a GBM on T is the trap.
- 1
Name the decision
Should we do X, for whom, vs who will churn under the current policy? - 2
Write the estimand
ATE / ATT / CATE. If they asked for a forecast, stop — that is prediction. - 3
Identify
DAG, backdoor, overlap. Unmeasured confounding → sensitivity or a better design. - 4
Pick the estimator
DML for a low-dimensional ATE; forests/metalearners for CATE; ITS/DiD/SC for a shock over time. - 5
Validate before acting
Overlap, sensitivity, and the health checklist. Do not ship a CATE from a feature-importance plot.
Overview
Interviewers will ask: our model predicts well — why isn’t that enough for a treatment decision? Prediction answers “what will happen.” Causal ML answers “what would happen if we intervene.” Correlation is not causation.
This hub stays under AI / ML. It does not create a machine-learning topic, and it does not duplicate Structured Outputs. Live AI/ML today is still mostly LLM systems; this cluster is the missing causal track: graphs, Double Machine Learning (DML), heterogeneous effects, causal time series, and health ethics.
You should be able to:
- Defend estimands before estimators.
- Point at the five sibling pages without folding them into this one.
- Send a pure forecast question to a model-choice / forecasting lesson — without re-teaching ARIMA, Prophet, or GBM.
Association vs causation
| Lens | Formal target | Interview one-liner |
|---|---|---|
| Association | P(Y given X=x) — observational | Prediction under current assignment |
| Intervention | P(Y given do(X=x)) — we set treatment | What happens if we set treatment |
| Counterfactual | Y_i(1) minus Y_i(0) for the same unit | Usually unobservable; we estimate averages |
Potential outcomes:
- ATE =
E[Y(1) − Y(0)]— population average effect - ATT =
E[Y(1) − Y(0) | T=1]— effect on the treated - CATE =
E[Y(1) − Y(0) | X=x]— effect at covariatesx(personalization)
When ML prediction is the wrong tool: pricing experiments, drug dosing, hospital staffing policy, feature launches with selection bias — any decision that changes the world the model was trained on.
Health / policy / product examples
- Health: Does a new care pathway reduce 30-day readmission, or do healthier patients select into it? Depth: graphs and health ethics.
- Policy: Did a vaccine mandate change infection rates beyond seasonal trends? Depth: causal time series.
- Product: Did recommending a plan increase adherence, or did engaged users already click?
Predictive ML vs causal ML vs RCTs
| Approach | Wins | Loses |
|---|---|---|
| Predictive ML | Accuracy, mature tooling, clear holdout metrics | Breaks under intervention; confuses selection with effect |
| Causal ML | Targets a treatment effect; flexible ML for nuisances (DML) | Needs identification, positivity/overlap, harder validation |
| RCTs | Gold-standard identification | Cost, speed, ethics, generalizability |
Rule of thumb: stakeholder asks “should we do X?” → causal frame. Stakeholder asks “who will churn?” → prediction may suffice.
Architecture (decision paths)
Single-column chooser. Identification before any fit. Unmeasured confounding does not get a cleverer GBM.
Decisions
- ?
1 Decision or prediction?
- predictionPredictive ML
- intervene2 Draw DAG
- 2
Predictive ML
- 3
2 Draw DAG
- next3 Identified?
- ?
3 Identified?
- noSensitivity / redesign
- yes4 Personalize?
- 5
Sensitivity / redesign
- ?
4 Personalize?
- yes5 CATE forests
- no6 Policy over time?
- 7
5 CATE forests
- ?
6 Policy over time?
- yes7 ITS / DiD / SC
- no8 ATE via DML
- 9
7 ITS / DiD / SC
- 10
8 ATE via DML
Lesson map
Causal ML & Double Machine Learning — From Association to Effect
Hub: correlation is not causation. Defend ATE, ATT, and CATE first, then identification, then estimators. Predictive ML is the wrong tool for treatment decisions; this cluster maps graphs, DML, CATE forests, causal time series, and health ethics.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB q["1 Decision or prediction?"] d["2 Draw DAG"] i["3 Identified?"] h["4 Personalize?"] q -->|intervene| d d -->|2 Draw DAG to 3 Identified?| i i -->|yes| h
- Graphs & backdoor: identification
- ATE with ML nuisances: DML
- For whom: CATE / forests
- Shocks over time: ITS / SC / DiD
- Fail closed: health validation
Sandbox: prediction vs causal gate (Python)
Educational router. predictive-ml is the forecasting / model-choice sibling — not this cluster.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same gate (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
This cluster (five siblings)
- Causal graphs & identification — DAGs, confounders, backdoor.
- Double Machine Learning — nuisances, orthogonalization, cross-fitting.
- Heterogeneous treatment effects — CATE, causal forests, personalized medicine.
- Causal time series — ITS, synthetic control, DiD over time.
- Causal ML in health — outcomes, bias, ethics, validation.
Forecasting answers capacity and “what next under the old policy.” Causal time series answers the shock. Do not bake off ARIMA vs Prophet vs GBM here.
Pitfalls
Product wants to “use the churn model to decide who gets a retention discount.” Write ATE vs ATT vs CATE for that sentence. What changes if the discount is the intervention vs if you only rank who looks likely to churn under the current policy? Where does a DAG enter before any fit?
Interview Q&A
Why isn’t a high-AUC model enough for treatment?
Answer
AUC scores association under current assignment. Intervening changes the covariate distribution and often the outcome mechanism. You need an identified effect, not a ranking of who already looks like the treated.
ATE vs ATT vs CATE?
Answer
ATE is the population average effect. ATT conditions on being treated. CATE conditions on covariates for personalization. Different decisions need different estimands — say which one before naming DML or a forest.
When does predictive ML still win?
Answer
Pure forecasting, triage risk scores that do not change care pathways, or when an RCT already identified the effect and you only need targeting proxies. Capacity “how many beds next Tuesday” is a forecast, not a mandate effect.
Biggest interview trap?
Answer
Fitting Y ~ T + X with a black-box ML model and reading the T coefficient as causal without identification. Regularization that helps prediction can wreck that coefficient. Depth: DML.
How does this relate to forecasting docs?
Answer
Forecasting predicts the future under status quo dynamics. Causal time series estimate counterfactuals after a shock. Cross-link a Choosing ML Algorithms / forecasting lesson lightly. Do not re-teach ARIMA, Prophet, or GBM here.
Do we need an RCT every time?
Answer
RCTs identify by design. Observational causal ML (backdoor + DML, DiD, synthetic control) is complementary when trials are slow, unethical, or too narrow — and it does not replace confirmatory trials for high-stakes treatment claims. Depth: health.
What is Double Machine Learning in one sentence?
Answer
Use flexible ML for nuisance functions, then an orthogonal score plus cross-fitting so a low-dimensional effect (often ATE) stays √n-consistent. It does not invent identification. Depth: DML.
When do I want CATE instead of ATE?
Answer
When the decision is who to treat, not only whether the average effect is positive. Requires overlap inside subgroups and a policy evaluation — not a hunt for a star leaf. Depth: CATE.
Prediction vs counterfactual after a Tuesday launch?
Answer
A dashboard that “forecasts well” can still mis-attribute seasonality to the launch. You need ITS / DiD / synthetic control. Depth: causal TS.
Where does health data silently break this story?
Answer
Confounding by indication, immortal time, informative missingness, positivity holes, and deploying a CATE without prospective validation. Depth: health.