Causal ML in Health — Outcomes, Bias, Ethics & Validation
Health causal ML is a safety rail: endpoints, selection bias, confounding by indication, fairness, and validation before acting on an estimate. The bravest senior answer is often not deploying a CATE yet.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Rich EHR, aggressive ML
Prefer
Estimand, DAG, overlap, sensitivity, silent eval
Observational causal ML complements trials for surveillance and hypothesis generation. High-stakes treatment claims still need confirmatory evidence.
- Write ATE/ATT/CATE and the decision before fitting.
- Clinician DAG review — disagreement is a stop, not an average of graphs.
- Kill criteria if calibration or fairness degrade after a protocol change.
Alternative
Unvalidated black-box uplift on EHR as causal CDS
Interview fail if sold as a treatment effect. More columns do not causalize missing labs or immortal time.
- Informative missingness is not MCAR imputation.
- Sicker patients get the drug — confounding by indication.
- TRIPOD+AI is for **prediction** reporting; do not launder it as causal identification.
Health validation gate
The bravest senior answer is often: we should not deploy causal ML here yet.
- 1
Write estimand + decision
ATE/ATT/CATE linked to an action a clinician would take. - 2
Clinician DAG review
Adjustment set justified. Disagreement → do not average and ship. - 3
Estimate + diagnostics
DML / CATE / ITS plus overlap and sensitivity (E-value / bounds). - 4
Checklist
Negative controls, multiplicity, equity, monitoring plan, legal pathway. - 5
Prospective silent mode
Fail in silent eval → do not deploy. Human override stays on.
Overview
Health data looks rich (EHR, claims, registries) and invites aggressive ML — then quietly violates causal assumptions. Senior interviews probe immortal time bias, informative missingness, confounding by indication, and whether you would deploy an observational CATE model without prospective validation.
This page is the cluster’s safety rail: when not to ship, how to validate, and how ethics constrain personalization.
EHR pitfalls (name them)
- Informative missingness: labs missing because the clinician did not order them — missingness correlates with severity. “Impute and predict” is not causal repair.
- Immortal time bias: follow-up time before treatment eligibility counted as treated survival — artificially favors treatment.
- Confounding by indication: sicker patients receive therapy; naive comparisons look harmful or magical. Depth: DAGs.
- Competing risks & censoring: death precludes readmission; naive binary outcomes mislead.
- Practice-pattern drift: coding and protocols change; time is a confounder. Depth: causal TS.
- Selection into the EHR: insured, urban, portal users overrepresented — transportability fails.
RCTs vs observational causal ML
| Design | Wins | Loses |
|---|---|---|
| RCTs | Randomization balances observed and unobserved confounders in expectation | Cost, ethics, eligibility limits, Hawthorne effects, slow |
| Observational DML / forests | Scale, real-world populations, heterogeneity | Unmeasured confounding, positivity holes, data quality |
Interview answer: observational methods complement RCTs for surveillance and subgroup hypothesis generation; they do not replace confirmatory trials for high-stakes treatment claims.
Sensitivity analysis
Ask: how strong must an unmeasured confounder be to explain away the effect? (E-value, Rosenbaum bounds.)
Report ranges, not a single heroic ATE. If conclusions flip under mild hidden bias, do not deploy as clinical decision support.
When not to deploy
- No credible identification story, or DAG disagreement among clinicians
- Overlap failure for the actionable subgroup
- Effect size smaller than measurement error / coding noise
- Equity harm: CATE systematically under-treats a protected group due to biased labels
- No monitoring plan for distribution shift after protocol change
- Legal / regulatory pathway unclear for the claim type
Validation checklist (senior interview)
- Estimand written (ATE/ATT/CATE) and decision linked
- DAG reviewed with domain experts; adjustment set justified
- Positivity diagnostics plotted; trimming policy pre-specified
- Pre-analysis plan for primary outcome; multiplicity controlled for subgroups
- Negative controls / placebo outcomes where possible
- Sensitivity to unmeasured confounding reported
- External / temporal validation; ideally prospective silent mode
- Clinician override UX and adverse-event monitoring
- Documentation of known EHR biases for this outcome
- Kill criteria if calibration or fairness metrics degrade
Comparative: evidence tiers
| Tier | Role |
|---|---|
| RCT meta-analysis | Highest for average effects when available |
| Well-identified observational (DiD, SC, DML with rich Z) | Hypothesis + real-world evidence |
| Predictive risk scores | Triage under current care — not automatic treatment effects |
| Unvalidated black-box uplift on EHR | Interview fail if sold as causal |
TRIPOD+AI is a prediction reporting standard. Use it to contrast “this is a risk score” vs “this is a treatment effect.” Do not cite it as identification.
Architecture (health validation gate)
Decisions
- 1
1 Write estimand
- next2 Clinician DAG review
- 2
2 Clinician DAG review
- next3 Fit DML / CATE / ITS
- 3
3 Fit DML / CATE / ITS
- next4 Sensitivity plus overlap
- 4
4 Sensitivity plus overlap
- next5 Checklist pass?
- ?
5 Checklist pass?
- yes6 Silent prospective eval
- noDo not deploy
- 6
6 Silent prospective eval
- fail silentDo not deploy
- pass7 Monitor with kill switch
- 7
Do not deploy
- 8
7 Monitor with kill switch
Lesson map
Causal ML in Health — Outcomes, Bias, Ethics & Validation
Health causal ML is a safety rail: endpoints, selection bias, confounding by indication, fairness, and validation before acting on an estimate. The bravest senior answer is often not deploying a CATE yet.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB e["1 Write estimand"] g["2 Clinician DAG review"] d["3 Fit DML / CATE / ITS"] s["4 Sensitivity plus overlap"] e -->|1 Write estimand to 2 Clinician DAG review| g g -->|2 Clinician DAG review| d d -->|3 Fit DML / CATE / ITS| s
Sandbox: immortal-time flagger (Python)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Deploy checklist scorer (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
A forest says patients in zip code Z have negative CATE for a care-management program. Labels are 30-day readmission from an EHR that under-codes uninsured follow-up. Overlap in Z is thin. Write three reasons you do not withhold the program from Z, and one design (RCT, silent eval, or coarsened ATE) you would run instead.
Interview Q&A
Immortal time bias fix?
Answer
Landmark analysis, time-varying exposure, or clone-censor-weight designs. Never assign pre-treatment survival to the treated. If labeled_as_treated_from_zero and time_to_treatment > 0, you are in the trap.
Can you causalize with more features?
Answer
More X helps only if they block backdoors. Proxies ≠ magic; extra variables can open colliders. Depth: graphs.
When is predictive risk OK clinically?
Answer
When it triggers assessment under existing protocols, not automatic treatment changes without identification. Triage scores are not CATEs.
How to talk unmeasured confounding?
Answer
Sensitivity / E-values; design alternatives (RCT, natural experiments, DiD / SC). If mild hidden bias flips the sign, do not ship CDS.
Ethics of CATE personalization?
Answer
Avoid amplifying disparities from biased outcomes. Require equity review and human override. Tiny slices with no overlap are not precision medicine. Depth: CATE.
Informative missingness vs MCAR imputation?
Answer
A lab is missing because nobody ordered it — that missingness is severity. Multiple imputation that assumes MCAR (or even MAR given the wrong set) does not repair the causal problem.
Competing risks in one sentence?
Answer
Death precludes readmission. A naive binary “readmitted in 30 days” treats the dead as successes. Use competing-risk or time-to-event estimands that match the decision.
What is a negative control here?
Answer
An outcome or exposure that should show a null if your design works (placebo diagnosis, future treatment). If the negative control “fires,” stop claiming the primary effect.
Silent mode vs A/B test?
Answer
Silent mode scores the observational policy without changing care — watches calibration and fairness. An RCT still identifies treatment. Silent mode is a gate, not a substitute trial.
Kill criteria examples?
Answer
Propensity collapse after a protocol change, fairness gaps widening for a protected group, calibration drift, or E-values that no longer bound a clinically relevant confounder. Write them before deploy.
Go Deeper
- Hernán & Robins — open book
- Immortal time bias review
- E-value calculator
- DoWhy / PyWhy
- TRIPOD+AI (prediction reporting — contrast with causal claims)
- Hub: Causal ML
- Prev: Causal time series