Distributed systems
Part 6 of 6 · Sagas & Distributed TransactionsSaga Failure Modes — Poison Steps, Partial Failure & Reconciliation
Poison steps, partial failure, dead-letter quarantine, and idempotent reconciliation for sagas that miss their SLA.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Sagas fail in structured ways. A step is retryable, a step is poisonous, compensation fails, a message is duplicated, or the process sits past its SLA. The design is detection, quarantine, and reconciliation. It is not an infinite retry loop.
Catalog
| Mode | Symptom | First response |
|---|---|---|
| Transient step error | 5xx or a timeout | Bounded retry with jitter, inside the deadline |
| Business reject | 4xx such as card declined | Compensate, then terminal FAILED |
| Poison payload | The same failure forever | Dead-letter, fix data or code |
| Partial completion | Some steps done, saga not terminal | Continue or compensate by policy |
| Compensation failure | Undo throws or keeps failing | COMPENSATION_FAILED and page |
| Duplicate delivery | The same event twice | Idempotency key or inbox |
| Lost event | Saga stuck in WAITING | Timeout, then probe and reconcile |
| Lease skew | Two drivers | Version or fencing check; reject the stale writer |
| Zombie | Past deadline and still RUNNING | Sweeper |
Flow
- 1
1. Classify the step error
- next2. Transient retry stays inside budget
- 2
2. Transient retry stays inside budget
- next3. Business reject starts compensation
- 3
3. Business reject starts compensation
- next4. Poison payload goes to quarantine
- 4
4. Poison payload goes to quarantine
- next5. Failed undo pages reconciliation
- 5
5. Failed undo pages reconciliation
- next6. Sweeper checks invariants
- 6
6. Sweeper checks invariants
Lesson map
Quarantine, then probe
The undo is quarantined. The sweeper is watching. Payments still shows the unfinished capture.
Architecture. Sweeper Watching. Saga log COMPENSATION_FAILED. Payments Undo failed
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB sweeper["Sweeper Watching"] saga["Saga log COMPENSATION_FAILED"] payments["Payments Undo failed"] payments -->|Declined| saga saga -->|Compensate| payments payments -->|Poison undo| saga sweeper -->|Probe| payments sweeper -->|Audited repair| saga sweeper -->|Force complete| saga
Play is one path: a decline, a poison undo, then a probe. It is not a path every error walks. A decline does not retry five times first. A poison message does not compensate until you know the product policy says “cancel.” A sweeper still runs for sagas that never raised, because a lost event looks like silence.
Poison steps
A poison step will not succeed without a person or a deploy. Typical causes: a schema bug, a permanently invalid SKU, a missing payment-provider config, a payload that fails validation on every attempt.
Detect it with:
- Attempt count at or above N with the same error signature.
- An error taxonomy you labeled non-retryable (most 4xx except a timeout you chose to retry).
- A dependency whose breaker has been open long enough that the saga deadline will hit anyway. Breaker behavior stays on Resilience Patterns.
Action: stop the forward retry. Move the message to a dead-letter or quarantine store. Optionally compensate earlier steps if the product says the order should cancel. Do not block a whole consumer group on one poison record. Bus-level dead-letter and consumer lag are a different lesson; here the requirement is that one bad saga id cannot freeze every other saga.
Duplicates are not poison. They are the at-least-once case, absorbed by the idempotency key from the compensation lesson and by the inbox in Transactional Outbox & Inbox Patterns.
Partial failure is a product state
Users and downstream systems will observe RESERVED without PAID. Hide that and they refresh into a contradiction. Show it.
Engineering invariants are queries, not slogans:
- No shipment without a captured payment.
- No double capture for one saga id.
- A released hold does not stay reserved.
Reconcilers run those queries continuously for money and stock, and on a schedule for everything else. A nightly job that nobody reads is not a control.
Reconciliation
| Pattern | How | Use when |
|---|---|---|
| Timeout sweeper | Scan non-terminal sagas older than the SLA | Always |
| Probe participant | Read inventory or payment by saga id | A lost event is plausible |
| Ledger compare | Saga log versus provider or inventory reports | Money and stock |
| Change-stream assist | Look for participant rows the saga does not know | You already capture changes |
| Manual tool | Force compensate or force complete, with an audit | Heuristic cases only |
Flow
- 1
1. Sweeper finds a non-terminal saga
- next2. Probe participant by saga id
- 2
2. Probe participant by saga id
- next3. Compare the local log to the ledger
- 3
3. Compare the local log to the ledger
- next4. Use a change stream only as an input
- 4
4. Use a change stream only as an input
- next5. Apply an idempotent repair and audit it
- 5
5. Apply an idempotent repair and audit it
If the business already tails the database log, an orphan row is a useful signal. The capture pipeline itself is Change Data Capture. Lag, schema breaks, and tombstones stay on CDC Failure Modes. This page only consumes a fact those systems already emit: the participant table and the saga log disagree.
Reconciliation must be idempotent. Running the sweeper twice posts one reversing entry, not two. Every repair emits an audit event that names the saga id, the invariant, and the actor (job or human).
Do not auto-complete to clear a dashboard
Why people do it. The stuck count drops. A user gets unblocked.
Why it is dangerous. A guessed “complete” can ship without payment or capture twice. You launder an unknown into a success.
Prefer force-compensate into a safe state, or a dual-control human approval for force-complete when the probes prove every forward effect is durable. Silent healing is how invariants rot.
Chaos belongs in the test plan: kill the orchestrator mid-step, drop the completion event, delay the compensation. The allowed outcomes are a terminal state or COMPENSATION_FAILED with an alert. A saga that stays RUNNING with no sweeper is a failed test.
Classify the error (run this)
Business declines compensate immediately. A 408 is a retry even though it is a 4xx. A 400 that has already been seen six times is poison. A 503 is still a retry until the attempt cap or the saga deadline says otherwise.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The numbers N = 5 and N = 8 are teaching thresholds. Production sets them from the deadline and the error budget, not from a magic constant you copy into every service.
Invariant check (run this)
Shipped without paid is illegal in every state, including RUNNING. The query is the reconciler’s unit test.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
What is the difference between retry and reconcile?
Answer
Retry attempts the same step again soon, inside a budget. Reconcile compares authoritative systems later and repairs drift. Retry cannot see a lost event that nobody will redeliver. The sweeper can.
How do you stop one poison saga from stalling the rest?
Answer
Handle the error per saga id, dead-letter after N attempts, and alert. Do not loop forever on that record in the consumer, and do not pause the whole partition for the lifetime of a bad payload. Consumer-group mechanics stay in the messaging cluster; the saga rule is isolation of the poison id.
Compensation failed. What do you do?
Answer
Page with the saga id, the partial undo log, and a runbook that finishes the remaining undos or posts a ledger adjustment. Do not flip the saga to FAILED while the undo is incomplete. COMPENSATION_FAILED is the honest state from the state-machine lesson.
Can reconciliation complete a saga automatically?
Answer
Only when the invariants prove every forward effect is durable and unique. Prefer an explicit policy over a job that marks COMPLETED because the row is old. Force-complete is a dual-control action with an audit.
How do you detect a duplicate capture?
Answer
The payment provider’s idempotency key rejects the second capture. A ledger compare of captures versus orders catches the case where two different keys were used by mistake. The compensation lesson owns the key. This page owns the compare.
Why not retry until it works?
Answer
You amplify the outage, hide poison, and outlive the user’s patience. The saga deadline exists so “later” has an end. Past that end, compensate or quarantine.
What should chaos tests assert?
Answer
Kill the orchestrator mid-step, drop the completion event, and delay compensations. The saga reaches a terminal state or COMPENSATION_FAILED with an alert. “It usually finishes if you wait” is not an assertion.
What do you show leadership?
Answer
Terminal success rate, p95 time-to-terminal, stuck count past SLA, compensation-failure count, and dollars at risk in non-terminal payment states. A green consumer lag chart does not answer those.
Pitfalls
Pick checkout. Write the query that finds shipped-without-paid and the query that finds captured-without-order. For each hit, say whether the repair is compensate, probe-and-complete, or a human. Name the audit fields the repair writes. Say which existing change stream, if any, could feed the query without a new connector project.