Distributed systems
Part 3 of 6 · Sagas & Distributed TransactionsOrchestration vs Choreography — Central Coordinator vs Event Dance
Central coordinator versus an event dance for sagas, with decision guidance and pointers to outbox, inbox, and deadlines.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Both styles are sagas. The difference is where the script lives. Choose from branching, ownership, and how painful invisible coupling will be.
Orchestration, one column
A coordinator (a workflow engine or a saga service) decides the next command, records the state, and starts compensations. Participants do domain work. They do not know the whole checkout.
Flow
- 1
1. Orchestrator stores saga state
- next2. Command inventory to reserve
- 2
2. Command inventory to reserve
- next3. Command payments to capture
- 3
3. Command payments to capture
- next4. Command shipping to schedule
- 4
4. Command shipping to schedule
- next5. On failure compensate in reverse
- 5
5. On failure compensate in reverse
Lesson map
Who owns the next step
The orchestrator has the saga row and the hold. Payments and shipping have not run. The script is in one place.
Architecture. Orchestrator RUNNING. Inventory Reserved. Payments Not asked. Shipping Idle
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB orchestrator["Orchestrator RUNNING"] inventory["Inventory Reserved"] payments["Payments Not asked"] shipping["Shipping Idle"] orchestrator -->|Reserve| inventory orchestrator -->|Capture| payments orchestrator -->|Schedule| shipping orchestrator -->|Release| inventory inventory -->|StockReserved| payments payments -->|PaymentCaptured| shipping payments -->|Unpaid ship| shipping payments -->|PaymentFailed| inventory
Choreography, one column
Each service subscribes to a fact and emits the next fact. Inventory does not call payments. Payments learn that stock was reserved because an event said so.
Flow
- 1
1. OrderPlaced event
- next2. Inventory reserves and emits
- 2
2. Inventory reserves and emits
- next3. Payments react to StockReserved
- 3
3. Payments react to StockReserved
- next4. Shipping reacts to PaymentCaptured
- 4
4. Shipping reacts to PaymentCaptured
- next5. PaymentFailed releases the hold
- 5
5. PaymentFailed releases the hold
The Failure view on the map is the missed branch: payment declined, and shipping still ran. On the decline path the orchestrator releases the hold and does not schedule shipping. Both reactions have to exist either way.
Side by side
| Dimension | Orchestration | Choreography |
|---|---|---|
| Flow visibility | The script is in one place | The script is scattered across handlers |
| Coupling | Services depend on the orchestrator contract | Services depend on event contracts |
| Branching and timeouts | Natural in a workflow definition | Easy to miss a rare path |
| Deploy independence | Flow changes often ship in the orchestrator | New handlers can ship without a central deploy |
| Debugging | Replay the saga state | Trace across topics |
| Main risk | The orchestrator becomes a god service | Cycles and event spaghetti |
| Typical tools | Temporal, Cadence, Camunda, Step Functions | A log or bus plus a schema owner |
Where should the happy path live?
Prefer
Match the tool to the branching
Long-running, branching, or approval-heavy flows want a coordinator. Two to four stable steps with a schema owner can stay an event dance.
- One screen can answer where order X is.
- Compensation order stays explicit.
- Notifications and analytics can still be events that do not block the saga.
Alternative
Pick a style because the last system used it
A central workflow service with no domain left in the participants is a distributed monolith. A topic graph nobody can draw is an outage waiting for a rare branch.
- Sync RPC orchestration fails the request when payments are down.
- Choreography without a compensation for the rare branch ships the unpaid order.
- Versioned events become a product whether you staffed that or not.
Decision guidance
Prefer orchestration when:
- There are many steps, parallel forks, human approvals, or long waits.
- You need one SLA view: “where is order X?”
- Compensations are order-sensitive and easy to get wrong in handlers.
Prefer choreography when:
- You have two to four stable steps and a mature schema contract.
- Teams already own their topics and will not operate a workflow cluster for this flow.
- The latency path is “react when the event arrives,” not “poll a coordinator.”
Hybrid, on purpose: the orchestrator owns the spine (reserve, pay, ship). Choreography owns fan-out that must not block the saga: email, analytics, search. If the email topic is down, the order still reaches a terminal state.
Messaging and time, by reference
- Publish “step completed” with the transactional outbox so the local commit and the event cannot diverge. Do not add a second publish after commit.
- Consume with an inbox or idempotency key so at-least-once delivery does not run the step twice. That mechanism is the same outbox lesson.
- Put step and saga deadlines on the activities. The budget rules live on Timeouts, Budgets & Deadline Propagation. Jitter and retry storms live on Retry Storms, Backoff & Jitter. Breakers that fail a call fast live on Resilience Patterns. The saga still decides retry versus compensate.
Synchronous orchestration couples availability: if payments are down, the user request fails even though inventory could have reserved. Asynchronous orchestration with a durable workflow survives a coordinator restart. The user sees a processing state. That state has to be honest. The state-machine lesson names it.
A managed engine (Step Functions, Temporal, a similar workflow product) is still a coordinator. You trade license and platform skill for a durable timer and a replay log. It is not a way to skip compensations.
What each style costs
Orchestration. Flows are testable, compensations are visible, audits have a single history. The platform has to run. The temptation is to put every business rule in the orchestrator until domain services are thin RPC wrappers.
Choreography. Runtime coupling is loose and the system already speaks events. “Who owns the happy path?” gets blurry. Event versioning is a product. A cycle (A emits what B emits what A consumes) becomes a storm.
A tracker topic is a middle ground: participants still decide their reactions, and a read model keyed by correlation id shows progress without commanding every step. That is soft orchestration. It helps the dashboard. It does not invent missing compensations.
Orchestrator transitions (run this)
The table is the script. Unknown events leave the state unchanged so a duplicate does not jump the saga. pay_fail lands in CANCELLED. The reserve undo still has to run; this snippet only shows the state the coordinator would record before it calls that undo. The next lesson is the undo itself.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Choreography router (run this)
The map is the contract. If a new event has no entry, you found a missing path before production did.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
What is an orchestrated saga?
Answer
A coordinator invokes participants, synchronously or asynchronously, records saga state, and triggers compensations when a step fails or a deadline fires. Participants own their local transactions. The coordinator owns the sequence.
What is a choreographed saga?
Answer
Participants subscribe to domain events and emit the next events. There is no process that commands every step. The script is the union of the handlers. Someone still has to own the event names and the missing branches.
Which one makes timeouts easier?
Answer
Orchestration. Timers sit next to saga state, so a deadline is a transition. Choreography needs a per-service timer or a shared timeout event that every participant honors the same way. The state-machine lesson is the lifecycle; the resilience timeout lesson is how remaining budget shrinks.
How do you avoid event spaghetti?
Answer
Give event schemas an owner, ban undocumented topics, and add a check that flags cycles. When a flow grows forks and compensations, promote that flow to an orchestrator instead of adding another handler “just for this case.”
Should the orchestrator call participants with sync RPC?
Answer
Only for short, highly available steps. Sync couples the user request to every callee. A durable workflow plus async commands survives process death. The user-facing state should say processing, not success, until the saga is terminal.
Is a managed workflow engine cheating?
Answer
No. Step Functions, Temporal, and similar products are coordinators with a durable log and timers. You still design compensations, idempotency keys, and versioned activity contracts. You are buying operations, not skipping the saga.
Can choreography still show where the order is?
Answer
Yes. A tracker or read model consumes the same events and keys them by correlation id. That is observation, not control. Do not let the tracker start sending commands or you have built an orchestrator without saying so.
What breaks first as choreography grows?
Answer
Versioned contracts, and compensation paths for branches nobody tested. The happy path is in production. The “payment captured, ship call lost” path is a stuck saga. Failure modes in the last lesson are how you detect that.
Pitfalls
Take a flow with a human approval and a refund branch. List the steps. Mark which ones a single team can explain in one sitting. If the answer is “no one team,” sketch an orchestrator state list. If the answer is “four events we already version,” sketch the handler map and the one failure event each handler must understand.