Data engineering
Part 5 of 6 · Workflow OrchestrationTemporal Patterns in Production - Sagas, Child Workflows, Continue-As-New, Versioning & Worker Scaling
Sagas with compensations (register compensation first, non-retryable errors, runnable trip saga), child workflows and parent close policy, continue-as-new and history limits (runnable), versioning with patching vs Worker Versioning (runnable patched-marker and unsafe-change demo), human-in-the-loop with signals/updates and timers, idempotent activities, worker scaling and sticky queues, persistence and visibility stores and history shards; shipping-a-change decision chart.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Why register a compensation before calling the step?
Answer
A step can succeed remotely and still time out locally, and that booking still needs cancelling.
L2
Which saga failures should be non-retryable?
Answer
Business failures such as no cars available; backoff will not fix them.
L3
Child workflow or activity?
Answer
An activity for one side effect; a child workflow for its own lifecycle, waits, compensations or partitioning a huge job.
L4
What is continue-as-new?
Answer
Completing a run and atomically starting a new one with the same Workflow ID, empty history and current state.
L5
Patching or Worker Versioning?
Answer
Worker Versioning by default when you can run versioned deployments; patching for auto-upgrade workflows or when you cannot.
L6
What is a sticky queue?
Answer
A worker-specific queue that sends an execution's next workflow task back to the worker that has it cached.
L7
Why does the history shard count matter?
Answer
It bounds write concurrency and cannot change after the cluster is created.
Failure modes
Leaked booking after a timeout
The compensation was registered after the step, so a half-succeeded step was never undone.
Entity workflow terminated
It never continued-as-new and hit the history limit.
Running executions stall after a deploy
A command-changing edit shipped without a patch or versioning.
Misconceptions
Child workflows are for code organization.
Use them only for their own lifecycle, partitioning or deduplication.
Children carry over a parent's continue-as-new.
They do not; set the Parent Close Policy deliberately.
A patch can be removed once deployed.
Only after no open or replayable execution lacks its marker.
Interviewer traps
Fanning out 100,000 activities from one workflow.
Batch, or fan out through child workflows, to stay under the pending cap and history limits.
Scaling workers on CPU alone.
Scale on Schedule-To-Start latency and worker slot utilization.
Design scenario
Same prompt for every reader.
Requirements
No orphaned bookings, loyalty workflows that never hit history limits, weekly code changes without stalls, and predictable scaling.
Failure assumptions
- A provider times out after booking.
- A deploy adds a fraud check step.
- Traffic triples during holidays.
Constraints
- Self-hosted Temporal on PostgreSQL.
- Providers accept idempotency keys.
Prompt
Run a trip-booking service on Temporal with flight, hotel and car providers, plus per-customer loyalty workflows that live for years.
API
Which workflows, child workflows, signals and updates exist?
Data
How are compensations, idempotency keys and continue-as-new state structured?
Architecture
How do versioning, task queues, worker pools and history shards fit together?
Overview
Running durable workflows in production is mostly about five problems. Partial failure across services: model it as a saga whose compensations the engine guarantees to finish. Growth: split work into child workflows and keep each history small with continue-as-new. Change: deploy new code without breaking executions that started on old code, by patching or by pinning executions to the build they started on (Worker Versioning). Time and people: long waits on humans are just durable timers plus signals or updates. Capacity: workers are stateless pollers you scale horizontally, helped by sticky queues that keep replays rare, while the server's persistence layer (history shards on Cassandra, PostgreSQL or MySQL) is the part you size once and carefully. Activities stay idempotent through all of it.
Child workflows
A child workflow is a workflow started by another workflow. It has its own Workflow ID, history and limits, and the parent can await its result or let it run independently under a Parent Close Policy (terminate, request cancel, or abandon when the parent closes).
| Use a child workflow when | Use an activity when |
|---|---|
| The sub-process has its own lifecycle, waits, signals or compensations | It is one side effect: an API call, a query, a file write |
| You need to partition a huge job (for example, 1,000 children each handling 1,000 items) so no single history grows too large | You need any library or I/O without deterministic constraints |
| A separate Workflow ID gives useful deduplication (one child per customer) | Low overhead matters |
Temporal's docs are explicit that there is no reason to use child workflows just for code organization: start with activities and reach for a child only when you need one of those properties. Also note: a parent that continues-as-new does not carry its running children over.
Orchestrated saga in a workflow or choreographed over a bus?
Prefer
Orchestrated saga in a durable workflow
The workflow runs steps and compensations, and survives crashes.
- book_car failed non-retryably and compensations ran in reverse.
- cancel_flight retried after a 503 and finished.
- Nothing was left booked at the providers.
Alternative
Choreographed saga over a message bus
Each service reacts to events and emits its own compensation.
- No single place knows the saga's state.
- A lost compensation event is silent.
- Harder to debug, though fine for loosely coupled domains.
A saga with versioned code and continue-as-new
Diagram 1 condensed: compensate, version and checkpoint.
- 1
Register the compensation
Append the undo before calling the step. - 2
Run the step
An activity with retries and an idempotency key. - 3
Compensate on business failure
A non-retryable error runs undos in reverse order. - 4
Branch on a patch
Old histories take the old path; new runs record a marker. - 5
Continue as new
Before history grows too large, restart with the current state.
Sagas with compensation
A saga replaces one distributed transaction with a sequence of local steps, each paired with a compensation that semantically undoes it (refund, cancel, release). The existing saga lessons cover orchestration vs choreography and failure modes; here the point is what a durable engine adds: the compensation logic itself cannot be lost, because the workflow that runs it survives crashes, and each compensation is an activity with its own retry policy.
"""Two production patterns inside a durable workflow, simulated.
Part 1, saga with compensation (orchestrated saga):
- transient failures are retried by the activity retry policy (with backoff)
- a non-retryable business failure triggers compensations in reverse order
- the compensation is registered BEFORE calling the step, because a step can
succeed on the remote side and still look failed to us (timeout)
- compensations are idempotent and retried until they succeed
Part 2, continue-as-new for a long-lived entity workflow:
- every handled event grows the event history
- before a threshold, the workflow checkpoints its state and continues as a
new run (same Workflow ID, new Run ID, fresh history)
"""
import itertools
class Transient(Exception): pass
class NonRetryable(Exception): pass
attempts = {}
remote_bookings = set() # what really exists at the providers (a set: providers dedup on an idempotency key)
def provider(step):
"""The outside world, scripted. Returns normally or raises."""
n = attempts[step] = attempts.get(step, 0) + 1
if step == "book_hotel":
remote_bookings.add("hotel") # booking succeeds remotely...
if n == 1: raise Transient("timeout") # ...but our first call times out
if step == "book_flight": remote_bookings.add("flight")
if step == "book_car": raise NonRetryable("NoCarsAvailable")
if step == "cancel_flight" and n == 1: raise Transient("503 from airline")
if step.startswith("cancel_"): remote_bookings.discard(step[len("cancel_"):])
return f"{step}:ok"
def activity(step, max_attempts=5, initial=1.0, coeff=2.0):
"""Retry policy: exponential backoff, stop on non-retryable errors."""
delay = initial
for a in itertools.count(1):
try:
r = provider(step); print(f" {step:<14} attempt {a}: ok"); return r
except NonRetryable as e:
print(f" {step:<14} attempt {a}: NON-RETRYABLE {e}"); raise
except Transient as e:
if a >= max_attempts: raise
print(f" {step:<14} attempt {a}: transient ({e}), retry in {delay:.0f}s"); delay *= coeff
def trip_saga():
compensations = []
try:
for step, undo in [("book_flight", "cancel_flight"), ("book_hotel", "cancel_hotel"), ("book_car", "cancel_car")]:
compensations.append(undo) # register first: the step may half-succeed
activity(step)
return "booked"
except NonRetryable:
print(" -> compensating in reverse order")
for undo in reversed(compensations):
activity(undo, max_attempts=100) # compensations must eventually succeed
return "rolled back"
print("Part 1: trip booking saga")
print(" result:", trip_saga(), "| still booked at providers:", sorted(remote_bookings) or "nothing")
print("\nPart 2: entity workflow with continue-as-new (demo threshold: 20 events)")
EVENTS_PER_SIGNAL = 3 # e.g. signal received + workflow task scheduled/completed (simplified)
THRESHOLD = 20 # real Temporal warns at 10,240 events and hard-limits at 51,200 events or 50 MB
def account_run(run_id, state, deposits):
history = 2 # started + first workflow task
while deposits:
amt = deposits.pop(0); state["balance"] += amt; state["seq"] += 1; history += EVENTS_PER_SIGNAL
if history + EVENTS_PER_SIGNAL > THRESHOLD and deposits:
print(f" run {run_id}: {history} events, balance={state['balance']} seq={state['seq']} -> continue-as-new")
return state, True
print(f" run {run_id}: {history} events, balance={state['balance']} seq={state['seq']} -> idle, waiting for signals")
return state, False
deposits = [10] * 14
state, run = {"balance": 0, "seq": 0}, 1
while True:
state, cont = account_run(f"account-7/run-{run}", state, deposits)
if not cont: break
run += 1
print(f" same Workflow ID 'account-7', {run} runs, no history ever exceeded {THRESHOLD} events")Output:
Part 1: trip booking saga
book_flight attempt 1: ok
book_hotel attempt 1: transient (timeout), retry in 1s
book_hotel attempt 2: ok
book_car attempt 1: NON-RETRYABLE NoCarsAvailable
-> compensating in reverse order
cancel_car attempt 1: ok
cancel_hotel attempt 1: ok
cancel_flight attempt 1: transient (503 from airline), retry in 1s
cancel_flight attempt 2: ok
result: rolled back | still booked at providers: nothing
Part 2: entity workflow with continue-as-new (demo threshold: 20 events)
run account-7/run-1: 20 events, balance=60 seq=6 -> continue-as-new
run account-7/run-2: 20 events, balance=120 seq=12 -> continue-as-new
run account-7/run-3: 8 events, balance=140 seq=14 -> idle, waiting for signals
same Workflow ID 'account-7', 3 runs, no history ever exceeded 20 eventsThree details in Part 1 are the difference between a demo saga and a production one:
| Detail | Why | If you skip it |
|---|---|---|
| Register the compensation before calling the step | A step can succeed remotely and still time out locally (the hotel case) | A booking exists that nobody cancels |
| Compensations are idempotent and retried "forever" | They may run after a partial success, or twice | Refunds fail permanently or happen twice |
| Business failures are non-retryable errors | "No cars available" will not fix itself with backoff | Default unlimited retries hide the failure for hours |
| Steps carry idempotency keys (Workflow ID plus step name) | Activities are at-least-once | Double bookings and double charges |
Continue-as-new and history limits
Every execution's history is capped: Temporal warns at 10,240 events or 10 MB and terminates at 51,200 events or 50 MB. There is also a default cap of 2,000 pending activities, child workflows, signals or cancellation requests per execution (each separately). A long-lived entity workflow (a customer account, a device, a subscription) that handles events forever must therefore checkpoint and continue-as-new: start a fresh run with the same Workflow ID, a new Run ID, an empty history, and the current state passed as input. Part 2 of the code shows the loop. SDKs expose a "continue-as-new suggested" flag (for example, workflow.info().is_continue_as_new_suggested() in Python) so you do not hard-code thresholds; drain pending signal and update handlers before you continue.
Continuing-as-new periodically also helps with versioning: a new run picks up the current code, so an entity workflow never has to replay months of history on new code.
| History-growth tool | What it bounds | Trade-off |
|---|---|---|
| Continue-as-new | One run's history | State must be serializable input; children do not carry over |
| Child workflows | Fan-out size per history | Extra executions to observe |
| Bigger activities (batch inside one activity) | Number of events | Coarser checkpoints, more redo on failure |
| Pass references, not payloads | History bytes | An extra store to manage |
Diagram 1: a saga with versioned code and continue-as-new
Decisions
- 1
Step 1: order workflow starts with Workflow ID order-42
- nextStep 2: register compensation, then run step activity with retry policy
- 2
Step 2: register compensation, then run step activity with retry policy
- nextStep 3: activity outcome
- ?
Step 3: activity outcome
- nextStep 4: more steps?
- nextretry with backoff, same idempotency key
- nextFailure path: run compensations in reverse order, each retried until it succeeds
- ?
Step 4: more steps?
- nextStep 2: register compensation, then run step activity with retry policy
- nextStep 5: history near limit or continue-as-new suggested?
- ?
Step 5: history near limit or continue-as-new suggested?
- nextStep 6: continue-as-new with current state, same Workflow ID, new run
- nextStep 7: workflow completes
- 6
Step 6: continue-as-new with current state, same Workflow ID, new run
- nextStep 2: register compensation, then run step activity with retry policy
- 7
Step 7: workflow completes
- 8
retry with backoff, same idempotency key
- nextStep 2: register compensation, then run step activity with retry policy
- 9
Failure path: run compensations in reverse order, each retried until it succeeds
- nextStep 8: workflow ends as rolled back and emits an alert or event
- 10
Step 8: workflow ends as rolled back and emits an alert or event
Lesson map
Temporal Patterns in Production - Sagas, Child Workflows, Continue-As-New, Versioning & Worker Scaling
Sagas with compensations (register compensation first, non-retryable errors, runnable trip saga), child workflows and parent close policy, continue-as-new and history limits (runnable), versioning with patching vs Worker Versioning (runnable patched-marker and unsafe-change demo), human-in-the-loop with signals/updates and timers, idempotent activities, worker scaling and sticky queues, persistence and visibility stores and history shards; shipping-a-change decision chart.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Step 1: order workflow starts with Workflow ID order-42"] b["Step 2: register compensation, then run step activity with retry policy"] c["Step 3: activity outcome"] d["Step 4: more steps?"] e["Step 5: history near limit or continue-as-new suggested?"] f["Step 6: continue-as-new with current state, same Workflow ID, new run"] g["Step 7: workflow completes"] r["retry with backoff, same idempotency key"] x["Failure path: run compensations in reverse order, each retried until it succeeds"] y["Step 8: workflow ends as rolled back and emits an alert or event"] a -->|continues| b b -->|continues| c c -->|continues| d d -->|continues| b d -->|continues| e e -->|continues| f e -->|continues| g c -->|continues| r r -->|continues| b c -->|continues| x x -->|continues| y f -->|continues| b
Versioning: changing code under running workflows
The previous page showed why any change to the sequence of commands breaks replay of old histories. Two families of solutions exist:
// Changing workflow code while executions are in flight.
// v1: charge -> ship. v2 adds a fraud check before charge.
// Replaying an old (v1) history with v2 code breaks determinism unless the change
// is guarded by a patch marker (Temporal: patched()/GetVersion), or the old run is
// pinned to the old build (Worker Versioning). Simplified simulation.
type Hist = { cmds: string[]; markers: Set<string> };
type Code = (ctx: Ctx) => string[];
class NonDeterminism extends Error {}
class Ctx {
private i = 0; issued: string[] = [];
constructor(private h: Hist, private replaying: boolean) {}
patched(id: string): boolean {
// Not replaying: record a marker and take the new path.
// Replaying: take the new path only if the original run recorded the marker.
if (!this.replaying) { this.h.markers.add(id); return true; }
return this.h.markers.has(id);
}
activity(name: string) {
if (this.replaying && this.i < this.h.cmds.length && this.h.cmds[this.i] !== name)
throw new NonDeterminism(`history step ${this.i + 1} is ${this.h.cmds[this.i]}, code asked for ${name}`);
if (this.i >= this.h.cmds.length) this.h.cmds.push(name);
this.i++; this.issued.push(name);
}
}
const v1: Code = (c) => { c.activity("charge"); c.activity("ship"); return c.issued; };
const v2Unsafe: Code = (c) => { c.activity("fraud_check"); c.activity("charge"); c.activity("ship"); return c.issued; };
const v2Patched: Code = (c) => {
if (c.patched("add-fraud-check")) c.activity("fraud_check");
c.activity("charge"); c.activity("ship"); return c.issued;
};
// An order that started on v1 and is mid-flight: 'charge' is already in its history.
const inflightV1 = (): Hist => ({ cmds: ["charge"], markers: new Set() });
function tryReplay(label: string, code: Code, h: Hist) {
try { console.log(` ${label}: ok, commands = ${code(new Ctx(h, true)).join(" > ")}`); }
catch (e) { console.log(` ${label}: ${(e as Error).constructor.name}: ${(e as Error).message}`); }
}
console.log("1) Deploy v2 and replay an in-flight v1 order");
tryReplay("v2 unsafe ", v2Unsafe, inflightV1());
tryReplay("v2 patched", v2Patched, inflightV1());
console.log("2) A brand-new order on v2 patched code records the marker");
const fresh: Hist = { cmds: [], markers: new Set() };
console.log(` new run: ${v2Patched(new Ctx(fresh, false)).join(" > ")}, markers = [${[...fresh.markers]}]`);
tryReplay("replay of that new run", v2Patched, fresh);
console.log("3) Patch lifecycle: patched() -> deprecate the patch -> delete it");
console.log(" only after no open (or retained, replayable) execution still lacks the marker");
console.log("4) Worker Versioning alternative: route by build instead of branching in code");
const current = "build-v2";
for (const wf of [{ id: "order-1", startedOn: "build-v1", behavior: "Pinned" }, { id: "order-2", startedOn: "build-v1", behavior: "AutoUpgrade" }, { id: "order-3", startedOn: "build-v2", behavior: "Pinned" }]) {
const target = wf.behavior === "Pinned" ? wf.startedOn : current;
const note = wf.behavior === "AutoUpgrade" && wf.startedOn !== current ? " (must stay replay-safe: still needs patching)" : "";
console.log(` ${wf.id} started on ${wf.startedOn}, ${wf.behavior} -> next workflow task goes to ${target}${note}`);
}Output:
1) Deploy v2 and replay an in-flight v1 order
v2 unsafe : NonDeterminism: history step 1 is charge, code asked for fraud_check
v2 patched: ok, commands = charge > ship
2) A brand-new order on v2 patched code records the marker
new run: fraud_check > charge > ship, markers = [add-fraud-check]
replay of that new run: ok, commands = fraud_check > charge > ship
3) Patch lifecycle: patched() -> deprecate the patch -> delete it
only after no open (or retained, replayable) execution still lacks the marker
4) Worker Versioning alternative: route by build instead of branching in code
order-1 started on build-v1, Pinned -> next workflow task goes to build-v1
order-2 started on build-v1, AutoUpgrade -> next workflow task goes to build-v2 (must stay replay-safe: still needs patching)
order-3 started on build-v2, Pinned -> next workflow task goes to build-v2Expected1) Deploy v2 and replay an in-flight v1 order v2 unsafe : NonDeterminism: history step 1 is charge, code asked for fraud_check v2 patched: ok, commands = charge > ship 2) A brand-new order on v2 patched code records the marker new run: fraud_check > charge > ship, markers = [add-fraud-check] replay of that new run: ok, commands = fraud_check > charge > ship 3) Patch lifecycle: patched() -> deprecate the patch -> delete it only after no open (or retained, replayable) execution still lacks the marker 4) Worker Versioning alternative: route by build instead of branching in code order-1 started on build-v1, Pinned -> next workflow task goes to build-v1 order-2 started on build-v1, AutoUpgrade -> next workflow task goes to build-v2 (must stay replay-safe: still needs patching) order-3 started on build-v2, Pinned -> next workflow task goes to build-v2
Press Run. Snippets must be self-contained — no network, files, or native modules.
| Approach | How it works | Strengths | Costs |
|---|---|---|---|
Patching (workflow.patched() in Python, .NET and Ruby; GetVersion in Go; patched in TypeScript) | Branch in code; a marker in history records which path a run took | Works with any deployment style; fine-grained | Branches accumulate; three-step lifecycle (patch, deprecate, remove) per change |
| Worker Versioning (Worker Deployments, Build IDs) | Each execution is routed by build; Pinned workflows finish on the build they started on, Auto-Upgrade ones move to the current build | No branches for pinned workflows; ramp, verify and instant rollback | Must run several builds at once (rainbow deployments); auto-upgrade workflows still need patching |
| New workflow type or task queue | Start new executions on a new type; let old ones drain | Simplest mental model | Two code paths live until old executions end |
| Continue-as-new at safe points | Long-lived runs restart on new code | Bounds how long old code must be supported | Only at points where state can be checkpointed |
Temporal's docs now recommend Worker Versioning as the default for most teams that can run versioned deployments, with patching for auto-upgrade workflows and for teams that cannot. Whatever you choose, replay tests in CI (next page) are the safety net.
Long-running, human-in-the-loop flows
Approvals, KYC reviews, document signing and AI-agent "ask a human" steps all have the same shape: do some work, wait for a person or a deadline, continue. In a durable engine the wait is a timer racing a signal or update, it costs nothing while waiting, and it survives deploys. Practical rules:
- Model the human's decision as an update when the UI needs an immediate answer ("approved, next step is payment"), or a signal when fire-and-forget is fine.
- Always race the wait against a timer (escalate after 48 hours, auto-reject after 7 days).
- Expose queries for the UI ("waiting for manager since Tuesday").
- For multi-week waits, prefer Pinned versioning or periodic continue-as-new so the execution is not replaying ancient code.
Airflow 3.1 added Human-in-the-Loop operators for data pipelines (tasks wait in an awaiting_input state for a UI response); Step Functions uses the callback pattern with a task token (.waitForTaskToken), available on Standard workflows only.
Idempotent activities, in practice
| Technique | Example |
|---|---|
| Idempotency key derived from workflow identity | key = f"{workflow_id}:{activity_id}" passed to the payment API |
| Conditional writes | INSERT ... ON CONFLICT DO NOTHING, DynamoDB condition expressions |
| Check-then-act inside one transaction | "If shipment for order 42 exists, return it" |
| Heartbeat with progress details | A long file import heartbeats the last processed offset; the retry resumes from it |
| Natural deduplication at the receiver | Inbox table keyed by message ID |
Workers, sticky queues and scaling
Workers are stateless pollers apart from their workflow cache, so scaling is adding processes. Two mechanisms make that efficient:
- Sticky execution: after a worker processes a workflow task, the server routes that execution's next workflow tasks to a worker-specific sticky queue, so the cached state is reused instead of replaying history. If that worker does not pick the task up quickly (5 seconds by default), stickiness is dropped and any worker on the normal queue can take it (replaying). Sticky execution applies only to workflow tasks.
- Separate task queues for different resource profiles: CPU-heavy or GPU activities on their own queue and worker pool, workflow tasks on another.
Scale on Schedule-To-Start latency (tasks waiting for workers) and worker slot utilization, not on CPU alone. Too many workflows cached per worker raises memory; too few causes replay churn.
The persistence layer, at a high level
The Temporal Server has four independently scalable services: Frontend (stateless gateway: auth, rate limits, routing), History (owns workflow mutable state, timers and transfer queues, partitioned into history shards), Matching (hosts task queues and matches tasks to polling workers), and an internal Worker service. Workflow IDs hash onto history shards; each shard maps to a persistence partition with one concurrent writer, so the shard count bounds write concurrency and cannot be changed after the cluster is created. Persistence is Cassandra, PostgreSQL or MySQL (SQLite for development only), and visibility ("list running workflows") uses SQL advanced visibility or Elasticsearch. Temporal Cloud runs all of this for you.
Diagram 2: how to ship a workflow code change
Decisions
- 1
Step 1: change to workflow code is ready
- nextStep 2: does it change commands for histories already in flight?
- ?
Step 2: does it change commands for histories already in flight?
- nextShip normally, replay tests still run in CI
- nextStep 3: executions are short, minutes to hours?
- 3
Ship normally, replay tests still run in CI
- ?
Step 3: executions are short, minutes to hours?
- nextWorker Versioning with Pinned behavior, old versions drain
- nextStep 4: executions run for weeks or forever?
- 5
Worker Versioning with Pinned behavior, old versions drain
- nextStep 5: replay tests pass against recorded histories?
- ?
Step 4: executions run for weeks or forever?
- nextPatch the code path, then deprecate and remove the patch after old runs finish
- nextAutoUpgrade behavior plus patching
- 7
Patch the code path, then deprecate and remove the patch after old runs finish
- nextStep 5: replay tests pass against recorded histories?
- 8
AutoUpgrade behavior plus patching
- nextStep 5: replay tests pass against recorded histories?
- ?
Step 5: replay tests pass against recorded histories?
- nextShip normally, replay tests still run in CI
- nextFailure path: block the deploy, fix the patch or version routing
- 10
Failure path: block the deploy, fix the patch or version routing
What happens if you choose otherwise
- Choreographed saga over a message bus instead of an orchestrated one: no single place knows the saga's state, and a lost compensation event is silent. Fine for loosely coupled domains; harder to debug.
- Deploying workflow code changes without patching or versioning: running executions fail with non-determinism errors and stall until you roll back or reset them.
- One workflow per customer that never continues-as-new: it eventually hits the history limit and is terminated.
- Huge fan-out from one workflow (100,000 activities): you hit the pending-activity cap and bloat history. Batch, or fan out through child workflows.
Pitfalls
- Registering compensations after the step: leaks resources when a step half-succeeds.
- Forgetting that children do not survive a parent's continue-as-new: set the Parent Close Policy deliberately.
- Removing a patch while old executions remain: they fail on replay. Use replay tests against real histories.
- Picking a tiny history shard count on self-hosted Temporal: it is fixed for the life of the cluster.
Interview Q&A
How do you implement a saga in Temporal?
Answer
As workflow code: before each step, append its compensation to a list; run the step as an activity with a retry policy; on a non-retryable failure, run the compensations in reverse order as activities with generous retries. The engine guarantees the compensation code finishes even across crashes. Every activity uses an idempotency key.
What is continue-as-new and when do you need it?
Answer
It completes the current run and atomically starts a new run with the same Workflow ID, fresh history and the state you pass in. You need it for long-lived or high-event workflows to stay under history limits (warning at 10,240 events or 10 MB, hard limit 51,200 events or 50 MB) and to move long-lived executions onto new code.
How do you ship a change that adds a step to a running workflow?
Answer
Either guard it with a patch (patched("id")) so old histories take the old path and new runs record a marker, later deprecating and removing the patch; or use Worker Versioning so pinned executions finish on the old build while new ones start on the new build. Validate either way with replay tests.
What is a sticky queue?
Answer
A worker-specific task queue the server uses to send an execution's next workflow task back to the worker that has it cached, avoiding a full replay. If that worker does not respond quickly, the task falls back to the shared queue.
Why is the history shard count important?
Answer
Each shard is a unit of serialized writes, so the count bounds the cluster's write concurrency, and it cannot be changed after the cluster is created.
What does Pinned vs Auto-Upgrade mean in Worker Versioning?
Answer
Pinned executions finish on the build they started on; Auto-Upgrade executions move to the current build and still need patching.
How should a human approval wait be modeled?
Answer
A timer racing a signal or update, with queries for the UI, and Pinned versioning or continue-as-new for multi-week waits.
What does the Matching service do?
Answer
It hosts task queues and matches tasks to polling workers.
Check yourself
Pick a multi-step process you own. Write its steps and their compensations in order, mark which errors are non-retryable, and decide when a long-lived version of it should continue-as-new.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Sagas & Distributed Transactions — Orchestration, Choreography & Compensations, Orchestration vs Choreography — Central Coordinator vs Event Dance, Compensating Transactions — Idempotent Undo & Semantic Rollback, Saga State Machines — Timeouts, Retries & Deadlines, Saga Failure Modes — Poison Steps, Partial Failure & Reconciliation, Transactional Outbox, Inbox & Consumer Idempotency, Rolling, Blue-Green & Canary — Strategies, PDBs & Blast Radius, Progressive Delivery & Kill Switches — Canary, Ramp, Instant Rollback.