Causal Time Series — ITS, Synthetic Control & Diff-in-Diff Over Time
Interrupted time series, synthetic control, and DiD over time estimate counterfactuals after a shock. Parallel trends and pre-fit beat a low forecast MAPE. Prediction under status quo is a different question — do not re-teach ARIMA, Prophet, or GBM.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
A feature ships on Tuesday. Did it move the metric?
Prefer
Counterfactual design (ITS / DiD / synthetic control)
Build what would have happened without the shock. Identification leads; the smoother is downstream.
- ITS: level/slope change at a clear interrupt, with seasonality.
- DiD: change vs controls under parallel trends.
- Synthetic control: weighted donors that match the pre-path.
Alternative
A more accurate forecast of next week
MAPE under status quo is a planning tool. It is not P(Y | do(launch)).
- Seasonality and concurrent shocks get attributed to your launch.
- ARIMA / Prophet / GBM bake-offs belong on a forecasting lesson.
- A dashboard that “looks accurate” can still be causally wrong.
Shock evaluation
If pre-trends fail, you do not get to claim causal with a better smoother.
- 1
Mark t*
Clear intervention time. Co-occurring shocks are the ITS killer. - 2
Ask for controls
None → ITS with caveats. One treated + donors → SC. Many treated → DiD (staggered-aware if timing differs). - 3
Pre-trends / pre-fit
Event-study plots; SC pre-MSPE. Parallel trends are not testable post. - 4
Placebos and sensitivity
Placebo-in-space, placebo-in-time. Concurrent shock → do not claim.
Overview
Hospitals roll out a sepsis bundle; states change Medicaid rules; a feature ships on Tuesday. Stakeholders want the effect of the shock — not a better forecast of the next week under the old policy.
Causal time series methods (interrupted time series, difference-in-differences, synthetic control) build counterfactual trajectories. Pure forecasting (ARIMA / Prophet / GBM) answers a different question. Cross-link a Choosing ML Algorithms sibling; do not re-teach those bake-offs here.
Hub reminder: if they asked “should we do X?”, you are in this cluster. If they asked “how many admissions next Tuesday under current protocol?”, you are in forecasting.
Prediction ≠ counterfactual
| Target | Formal | Use |
|---|---|---|
| Forecast | Expected Y at t+h given history, under status quo | Capacity, staffing, inventory |
| Counterfactual | Expected Y0 at t+h given history, without the intervention after t* | Policy / launch impact |
Confusing them ships dashboards that look “accurate” while mis-attributing seasonality or concurrent shocks to your launch.
Interrupted time series (ITS)
Single series; model level and/or slope change at intervention.
Needs: clear interrupt time, enough pre/post points, careful seasonality and autocorrelation, no co-occurring shocks.
Health: hospital-wide protocol change with monthly infection rates.
Weakness: one series cannot easily separate the intervention from everything else that moved that month — prefer controls when available.
Difference-in-differences (DiD)
Compare change over time in treated units vs change in controls.
Parallel trends: without treatment, treated and control gaps would stay constant. Pre-period event studies / plots defend it; it is not testable post. Sensitivity to violations is part of the answer.
Two-way fixed effects variants; staggered adoption needs care (heterogeneous timing can make TWFE weight oddly). Interview awareness beats implementing every modern estimator on the whiteboard. Health policy: states adopting a mandate in different years.
Synthetic control
Build a weighted combination of donor units that matches the pre-treatment trajectory of the treated unit; project those weights as the post counterfactual.
Great for one treated region / hospital with many donors.
Assumptions: stable weights, no interference, good pre-fit. Placebo-in-space / placebo-in-time for inference intuition.
Comparative: causal TS vs forecasting
| Method | Pros | Cons |
|---|---|---|
| ITS | Works with one series | Fragile to concurrent shocks |
| DiD | Uses controls; transparent | Parallel trends; staggered pitfalls |
| Synthetic control | Interpretable weights; good narrative | Needs suitable donors; extrapolation risk |
| Forecasting models | Accuracy under status quo | Not an effect estimator unless embedded in a causal design |
Rule: “impact of policy” → identification first. “Census next month” → forecasting sibling.
Architecture (shock evaluation)
Decisions
- 1
1 Mark intervention time
- next2 Controls available?
- ?
2 Controls available?
- no3 ITS with seasonality
- yes3 One treated unit?
- 3
3 ITS with seasonality
- next6 Pre-trends / pre-fit
- ?
3 One treated unit?
- yes4 Synthetic control
- no4 Staggered adoption?
- 5
4 Synthetic control
- next6 Pre-trends / pre-fit
- ?
4 Staggered adoption?
- yes5 Staggered-aware DiD
- no5 Classic DiD
- 7
5 Staggered-aware DiD
- next6 Pre-trends / pre-fit
- 8
5 Classic DiD
- next6 Pre-trends / pre-fit
- 9
6 Pre-trends / pre-fit
- next7 Placebos ok?
- ?
7 Placebos ok?
- yes8 Report with caveats
- noDo not claim causal
- 11
8 Report with caveats
- 12
Do not claim causal
Lesson map
Causal Time Series — ITS, Synthetic Control & Diff-in-Diff Over Time
Interrupted time series, synthetic control, and DiD over time estimate counterfactuals after a shock. Parallel trends and pre-fit beat a low forecast MAPE. Prediction under status quo is a different question — do not re-teach ARIMA, Prophet, or GBM.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB i["1 Mark intervention time"] c["2 Controls available?"] its["3 ITS with seasonality"] u["3 One treated unit?"] i -->|1 Mark intervention time| c c -->|no| its c -->|yes| u
Four methods sharing a pre-fit node is still one TB chooser — no side-by-side subgraphs.
Sandbox: ITS level-shift sketch (Python)
Toy mean shift ignoring seasonality — teaching only. A real ITS needs seasonality and autocorrelation.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Method picker (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
A hospital changes a bundle in March. Infection rates drop in April. Flu season also ends in April. No sister hospital. Is ITS enough? What concurrent-shock check do you demand? If a sister site exists, what would parallel trends look like in the pre year?
Interview Q&A
Why not just compare post vs pre?
Answer
Trends, seasonality, and concurrent events confound. You need a model (ITS) or controls (DiD / SC). A raw mean shift is a descriptive, not an effect.
How do you defend parallel trends?
Answer
Pre-period event studies / plots. The assumption is not testable post. Sensitivity to violations belongs in the answer, not a p-value theater.
Synthetic control vs DiD?
Answer
SC constructs a bespoke twin for one treated unit from donor weights. DiD uses group averages under parallel trends. SC wants good pre-fit and suitable donors; DiD wants a credible control group.
Staggered adoption issue?
Answer
Two-way FE can weight timing oddly under heterogeneous effects. Mention modern DiD estimators at intuition level. Do not pretend you implemented every paper on the whiteboard.
Forecasting cluster link?
Answer
Use forecasting for capacity/planning under no policy change. Use this page for intervention effects. Do not bake off ARIMA vs Prophet vs GBM here.
When is ITS acceptable?
Answer
Clear t*, enough pre/post, modeled seasonality, and a serious look at concurrent shocks. Prefer controls. Hospital-wide protocols often look like ITS because there is one series.
What is placebo-in-space?
Answer
Pretend a donor was treated and rerun SC. If donors also show huge “effects,” your design is picking up common shocks, not the intervention.
What is pre-MSPE in synthetic control?
Answer
How well the weighted donors match the treated unit before t*. Bad pre-fit → do not trust the post gap. Good pre-fit is necessary, not sufficient.
Can I use a GBM forecast as the counterfactual?
Answer
Only inside a design that says why that forecast is Y(0) (for example, a control series or a pre-only model with no concurrent shock). A low MAPE on the treated series after launch is circular.
How does this touch health ethics?
Answer
Policy rollouts on EHR time series still need outcome definitions, competing risks, and a “do not claim” gate. Depth: health.