Distributed systems
Part 5 of 6 · Sagas & Distributed TransactionsSaga State Machines — Timeouts, Retries & Deadlines
Saga state machines with per-step timers, whole-saga deadlines, bounded retries, and durable state.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Treat every saga as an explicit state machine. Illegal transitions are bugs you can reject. Operators can see the state. Timers attach to states instead of to a thread blocked in an RPC.
Canonical states
| State | Meaning |
|---|---|
PENDING or NEW | Accepted, not started |
RUNNING | A worker is driving a forward step |
WAITING | Parked on a timer, a person, or an external ack |
COMPENSATING | Undoing steps that already committed |
COMPLETED | Terminal success |
FAILED | Terminal business failure after compensations finished |
COMPENSATION_FAILED | Undo itself failed; a person or reconciler must finish |
TIMED_OUT | The saga deadline fired; policy chooses compensate or escalate |
FAILED means the business process stopped cleanly after undo. It does not mean “we are still trying.” COMPENSATION_FAILED means the undo did not finish. Mixing those two labels hides money.
Flow
- 1
1. NEW accepted
- next2. RUNNING a forward step
- 2
2. RUNNING a forward step
- next3. WAITING on a timer or person
- 3
3. WAITING on a timer or person
- next4. Resume into RUNNING
- 4
4. Resume into RUNNING
- next5. COMPLETED when every step is ok
- 5
5. COMPLETED when every step is ok
Lesson map
States, not a blocked thread
The row is WAITING. The worker is parked. Payment is not in a call.
Architecture. Saga record WAITING. Worker Parked. Payment Needs a person
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB saga["Saga record WAITING"] worker["Worker Parked"] payment["Payment Needs a person"] worker -->|Forward| payment saga -->|Timer| worker worker -->|Undo| payment saga -->|Deadline| worker payment -->|Undo failed| saga
Flow
- 1
1. Step failure or saga deadline
- next2. Enter COMPENSATING
- 2
2. Enter COMPENSATING
- next3. Undo completed steps in reverse
- 3
3. Undo completed steps in reverse
- next4. Undos ok means terminal FAILED
- 4
4. Undos ok means terminal FAILED
- next5. Poison undo is COMPENSATION_FAILED
- 5
5. Poison undo is COMPENSATION_FAILED
TIMED_OUT can be its own recorded state before the policy moves the saga into COMPENSATING. Some teams skip the extra label and compensate directly. Either way, the deadline is a transition you can query, not a log line. WAITING is not RUNNING: nothing should be inside the payment RPC while the row says waiting. A human approval is WAITING plus a long timer and an escalation, never a blocked thread.
Timeouts versus retries
| Knob | Too aggressive | Too loose | Guidance |
|---|---|---|---|
| Step timeout | False failure while the work still succeeds | Holds resources and the user | Start from p99 plus a margin; cap by remaining deadline |
| Step retries | Amplifies an outage | Gives up on a blip | Bounded, jittered, only on idempotent steps |
| Saga deadline | Cancels a healthy long flow | Leaves zombie sagas | Set it from the product SLA |
| Compensation retry | Hammers a broken undo | Sticks money or stock | Separate backoff, then page a human |
Circuit breakers and load shedding protect callers. Saga deadlines protect the business process. Use both. When a breaker is open, the saga should not burn its retry budget against it. The breaker states live on Resilience Patterns — Circuit Breakers, Bulkheads & Load Shedding. How long a hop may wait, and how nested calls shrink the budget, lives on Timeouts, Budgets & Deadline Propagation. Why unsynchronized retries stampede lives on Retry Storms, Backoff & Jitter. This lesson only decides the saga transition.
Deadline propagation
Pass deadline_unix_ms, or the remaining budget, into each activity. If the remainder is smaller than the minimum useful work, do not start the RPC. Fail into compensate. A server that keeps working after the saga has moved on is ghost work: the late capture races the refund from the previous lesson.
Choreography does not get a free pass. Each handler still needs a timer or a shared timeout event. A correlation read model can show WAITING, but some participant has to notice that the event never arrived. The orchestration lesson is why that timer is easier next to a coordinator.
Durable state
- Primary key:
saga_id. - Optimistic version or etag on every transition. A lost update is two drivers applying two decisions.
- A history or event log so you can audit and replay. The current row is a projection of that history.
- A worker lease so two orchestrators do not both run the same step. Fencing for stale holders is the locks cluster; do not re-derive it here. Store the version on the saga row and reject a write from an old lease.
The store belongs to the orchestrator or the workflow engine. In a choreographed flow the closest equivalent is a correlation read model. It is still durable. Memory on one pod is not a saga state machine.
Deadline-aware runner (run this)
Time is synthetic: each successful step subtracts its cost from the remaining budget. Reserve is 0.2, pay is 0.5, ship is 0.4. A 1.0 budget finishes reserve and pay, then refuses ship. A 0.6 budget finishes reserve and refuses pay.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The abort does not magically refund. The caller marks COMPENSATING and runs the undos for the OK steps. Starting ship with 0.3 left when ship needs 0.4 is how you get a timeout and a late success together.
Retry budget (run this)
Jitter here is deterministic so the snippet is testable: (attempt * 17) mod base. Waits stop when the next delay would pass the deadline. Exponential growth without that cap is how a saga outlives the customer.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Production jitter is random, not this modulo. The property to keep is the cap: no attempt starts when the remaining saga budget is gone. Backoff formulas beyond that belong to the retry-storms lesson.
Interview Q&A
Why model the saga as a state machine?
Answer
Illegal transitions become rejected writes. Operators see one current state. Timers are transitions out of WAITING or RUNNING, not forgotten setTimeout handles on a dead process. You can also answer “where is order X?” without grepping six services.
The timeout fired and the step succeeded later. What now?
Answer
That is a dual outcome. The forward step must be idempotent and must check saga state or version before it applies side effects. A late success after the saga has started compensating should not capture again; it should no-op or it should trigger the matching undo if the capture really landed. “The timeout means the call failed” is the bug.
What is wrong with exponential backoff and no jitter?
Answer
Retries line up and hit the dependency together. Jitter spreads them. The saga deadline is still a hard stop even if the next backoff would have been polite. Unbounded backoff is a zombie with a calendar invite.
Where does saga state live?
Answer
In a durable store the orchestrator owns: a database row or the workflow engine’s history. Choreography usually keeps a correlation read model fed by the events. Either store needs a version. A process-local map resets on deploy and splits brain when two replicas run.
How do WAITING and RUNNING differ?
Answer
RUNNING means a worker is actively driving a step. WAITING means the machine is parked on an external signal or a timer. Human approval is WAITING with a long timer and an escalation, not a thread blocked for a day.
How do human approvals fit?
Answer
Transition to WAITING, persist the timer, and leave. On timeout, escalate or compensate according to policy. Do not hold a worker or a database connection until the person clicks.
How do circuit breakers fit?
Answer
An open breaker can fail the step call fast. The saga then chooses retry later, compensate, or shed. Do not loop retries through an open breaker. Breaker internals stay on the resilience hub.
Which metric catches deadline bugs?
Answer
A histogram of saga age by state, plus counts of TIMED_OUT and COMPENSATION_FAILED. If age grows in RUNNING while step timeouts do not fire, the timer is not attached to the state machine.
Pitfalls
Write the states for reserve, pay, and a one-day fraud review. Put the saga deadline at 15 minutes for pay and one day for review. Say which timeout compensates, which one escalates to a person, and what a late payment capture must check before it writes.