Data engineering
Part 4 of 6 · Workflow OrchestrationDurable Execution & the Temporal Model - Workflows vs Activities, Event History, Replay & Determinism
Workflows vs activities, event history and replay, determinism rules (SDK time/random, no I/O in workflow code) with a runnable replay engine showing a wall-clock non-determinism bug and the fix; timers, signals, queries and updates (runnable race demo); task queues and workers; retry policy defaults and the four activity timeouts; what exactly-once means (and does not) with idempotency keys; history and payload limits; workflow-vs-activity decision chart.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Workflow or activity: which may call an API?
Answer
The activity. Workflow code must not do I/O.
L2
What is the event history?
Answer
An append-only log of every command and result for one Workflow Execution.
L3
When does replay happen?
Answer
When no worker has the workflow cached: restarts, crashes, evictions, or queries on uncached workflows.
L4
What causes a non-determinism error?
Answer
Replay issues a different command than the history records, for example after branching on the wall clock.
L5
Signal, update or query?
Answer
Signal is async and recorded, update is sync and returns a result, query is a read that does not touch history.
L6
What are the four activity timeouts?
Answer
Schedule-To-Start, Start-To-Close, Schedule-To-Close and Heartbeat.
L7
Are activity side effects exactly-once?
Answer
No. They are at-least-once; use idempotency keys derived from the Workflow ID and activity ID.
Failure modes
Workflow stuck on a non-determinism error
Code branched on the wall clock or changed command order, so replay no longer matches history.
Double charge after a worker dies
The activity charged the card and died before reporting; the retry charged again without an idempotency key.
Dead worker noticed late
No Start-To-Close or Heartbeat timeout on a long activity, so only an unlimited Schedule-To-Close applies.
Misconceptions
Durable workflows run each side effect exactly once.
Workflow decisions are recorded once; activity effects are at-least-once.
Replay re-executes activities.
Finished activities return their recorded results during replay.
Queries are free.
A query on an uncached workflow replays its whole history.
Interviewer traps
Calling datetime.now() in workflow code.
Use workflow.now() or pass time in as input.
Leaving business errors retryable under the default policy.
Raise non-retryable error types so they fail fast instead of retrying forever.
Design scenario
Same prompt for every reader.
Requirements
No lost approvals, no double charges, a status page for support, and safe behavior when all workers are down for hours.
Failure assumptions
- Workers restart during deploys.
- An approval signal arrives while no worker runs.
- The payment worker dies after charging.
Constraints
- Payment API accepts idempotency keys.
- Payloads stay under 256 KB.
Prompt
Design an order-approval workflow that waits up to 7 days for a manager, then charges a card and ships.
API
Which signals, updates and queries does the workflow expose?
Data
What lands in history, and how is each activity's idempotency key built?
Architecture
Which task queues, workers, timeouts and retry policies do you set?
Overview
Durable execution makes ordinary code crash-proof. In Temporal (and Cadence, Azure Durable Functions, Restate, Inngest, each with its own variant), you split code into workflows, which orchestrate and must be deterministic, and activities, which do the real side effects and may fail. The server records every decision and result as an event history. When a worker dies, another worker re-runs the workflow function from the top and replays it against the history: activities that already finished return their recorded results instantly, timers that already fired are skipped, and execution continues from where it stopped. That is why workflow code may not read the clock, call random, or do I/O directly, and why activities are at-least-once and must be idempotent. "Exactly-once" in this world means the workflow's logic and its recorded decisions, never the side effects inside activities.
Workflows vs activities
| Workflow | Activity | |
|---|---|---|
| Role | Orchestration logic: order, branching, waiting, compensation | One unit of real work: call an API, write a DB, run a model |
| Must be deterministic? | Yes, it is replayed | No, it runs once per attempt and its result is recorded |
| Allowed | Calling activities, child workflows, timers, waiting on signals or conditions, SDK time and random APIs | Any I/O, any library, any duration (with heartbeats if long) |
| Not allowed | Direct network or disk I/O, system clock, unseeded random, threads, reading mutable globals | Assuming it runs only once |
| Retries by default | No (a Workflow Execution does not retry by default) | Yes: initial interval 1 s, backoff coefficient 2.0, maximum interval 100 x initial, unlimited attempts |
| Failure means | Usually a bug or a business decision | Usually transient; retried by policy |
| Runs where | On a worker, as a cached state machine, replayed when needed | On a worker, as a normal function call |
Wall clock or SDK time inside workflow code?
Prefer
SDK time (ctx.now / workflow.now)
Read time from history so replay sees the same value.
- Worker B replayed at 17:01 and finished with approved:ana.
- The branch matched the recorded auto_approve command.
- Timers and branches stay stable across workers.
Alternative
The system clock
Read datetime.now() and branch on it.
- Worker B at 17:01 asked for manual_review.
- History said auto_approve, so replay raised NonDeterminismError.
- The workflow is stuck until code is fixed or reset.
Crash, replay and continue
Diagram 1 condensed: history is the truth, code replays against it.
- 1
Record commands
Each activity, timer and child start is appended to history. - 2
Worker crashes
The cached workflow state is lost with the process. - 3
Replay on another worker
The code reruns from the top; recorded results return instantly. - 4
Continue
New commands are appended once replay catches up. - 5
Mismatch
A different command than history records raises a non-determinism error.
Event history and replay
A Workflow Execution is identified by a Workflow ID you choose (an order number, a user ID) and a system Run ID. Temporal guarantees at most one running execution per Workflow ID in a namespace, which gives you a natural deduplication key. Every command your code issues (schedule an activity, start a timer, start a child) and every event that comes back (activity completed, timer fired, signal received) is appended to the execution's history, which lives in the persistence store as an append-only log next to a compact mutable state row.
Workers cache running workflows in memory (and Temporal routes follow-up tasks back to the same worker, see the production page), so most of the time nothing replays. Replay happens when the cache is missing: a worker restart, a crash, a cache eviction, or a query on a non-cached workflow. The engine re-runs your function from the start, matching each command against the next recorded event. A mismatch is a non-determinism error.
"""Event-sourced replay, the core of durable execution (Temporal, Cadence,
Azure Durable Functions, Restate and Inngest all use a variant of it).
- Workflow code is a generator that yields COMMANDS (run activity, start timer).
- The engine appends EVENTS to an append-only history.
- A worker that crashes loses its memory; the next worker rebuilds state by
REPLAYING the code against the history: commands already in history get their
recorded results back instead of running again.
- Replay only works if the code issues the same commands in the same order.
Reading the wall clock inside workflow code breaks that. ctx.now() fixes it by
returning the time recorded in history (like workflow.now() in Temporal SDKs).
"""
class NonDeterminismError(Exception):
pass
class WorkerCrash(Exception):
pass
class Ctx:
def __init__(self, engine): self.e = engine
def activity(self, name, arg): return ("activity", name, arg)
def timer(self, seconds): return ("timer", f"{seconds}s", None)
def now(self): return self.e.task_time # deterministic: comes from history
ACTIVITIES = {
"credit_score": lambda a: 640,
"auto_approve": lambda a: f"approved:{a}",
"manual_review": lambda a: f"queued-for-human:{a}",
}
class Engine:
def __init__(self, history, wall_clock, crash_after_scheduling=None):
self.h, self.wall, self.crash = history, wall_clock, crash_after_scheduling
self.task_time = None
def run(self, wf, arg):
if not self.h:
self.h.append({"type": "WorkflowExecutionStarted", "input": arg})
cmd_events = [e for e in self.h if e["type"] in ("ActivityTaskScheduled", "TimerStarted")]
done = {e["ref"]: e for e in self.h if e["type"] in ("ActivityTaskCompleted", "TimerFired")}
task_times = [e["time"] for e in self.h if e["type"] == "WorkflowTaskStarted"]
gen, send, i = wf(Ctx(self), arg), None, 0
while True:
# each decision point is one workflow task; its time is recorded once, then replayed
if i < len(task_times):
self.task_time = task_times[i]
else:
self.task_time = self.wall
self.h.append({"type": "WorkflowTaskStarted", "time": self.wall})
try:
kind, name, carg = gen.send(send)
except StopIteration as stop:
self.h.append({"type": "WorkflowExecutionCompleted", "result": stop.value})
return stop.value
if i < len(cmd_events): # REPLAY: must match history
ev = cmd_events[i]
if ev["name"] != name:
raise NonDeterminismError(f"history has {ev['type']}({ev['name']}) at step {i+1}, code now asks for {kind}({name})")
else: # NEW command: record it first
ev = {"type": "ActivityTaskScheduled" if kind == "activity" else "TimerStarted", "name": name, "id": i}
self.h.append(ev)
if self.crash == name:
raise WorkerCrash(name)
if i not in done: # not finished yet: execute now
result = ACTIVITIES[name](carg) if kind == "activity" else "fired"
comp = {"type": "ActivityTaskCompleted" if kind == "activity" else "TimerFired", "ref": i, "result": result}
self.h.append(comp); done[i] = comp
send, i = done[i]["result"], i + 1
def hour(t): return int(t.split(":")[0])
def loan_bad(ctx, applicant):
score = yield ctx.activity("credit_score", applicant)
# BUG: wall clock inside workflow code; differs between first run and replay
now = WALL_CLOCK_FOR_DEMO # stands in for datetime.datetime.now()
if score > 700 or hour(now) < 17:
decision = yield ctx.activity("auto_approve", applicant)
else:
decision = yield ctx.activity("manual_review", applicant)
yield ctx.timer(86400) # durable timer: survives restarts
return decision
def loan_fixed(ctx, applicant):
score = yield ctx.activity("credit_score", applicant)
if score > 700 or hour(ctx.now()) < 17: # time recorded in history
decision = yield ctx.activity("auto_approve", applicant)
else:
decision = yield ctx.activity("manual_review", applicant)
yield ctx.timer(86400)
return decision
def show(h): return [e["type"] + (f"({e['name']})" if "name" in e else "") for e in h]
for label, wf in [("BUGGY (reads wall clock)", loan_bad), ("FIXED (uses ctx.now())", loan_fixed)]:
print(f"== {label}")
history = []
WALL_CLOCK_FOR_DEMO = "16:59"
try:
Engine(history, "16:59", crash_after_scheduling="auto_approve").run(wf, "ana")
except WorkerCrash as c:
print(f" worker A (16:59) crashed right after scheduling {c}; history so far:")
print(" ", show(history))
WALL_CLOCK_FOR_DEMO = "17:01" # worker B picks up two minutes later
try:
out = Engine(history, "17:01").run(wf, "ana")
print(f" worker B (17:01) replayed and finished -> {out}")
print(" ", show(history))
except NonDeterminismError as e:
print(f" worker B (17:01) NonDeterminismError: {e}")
print(" the workflow is stuck until the code is fixed (Temporal retries the workflow task)")Output:
== BUGGY (reads wall clock)
worker A (16:59) crashed right after scheduling auto_approve; history so far:
['WorkflowExecutionStarted', 'WorkflowTaskStarted', 'ActivityTaskScheduled(credit_score)', 'ActivityTaskCompleted', 'WorkflowTaskStarted', 'ActivityTaskScheduled(auto_approve)']
worker B (17:01) NonDeterminismError: history has ActivityTaskScheduled(auto_approve) at step 2, code now asks for activity(manual_review)
the workflow is stuck until the code is fixed (Temporal retries the workflow task)
== FIXED (uses ctx.now())
worker A (16:59) crashed right after scheduling auto_approve; history so far:
['WorkflowExecutionStarted', 'WorkflowTaskStarted', 'ActivityTaskScheduled(credit_score)', 'ActivityTaskCompleted', 'WorkflowTaskStarted', 'ActivityTaskScheduled(auto_approve)']
worker B (17:01) replayed and finished -> approved:ana
['WorkflowExecutionStarted', 'WorkflowTaskStarted', 'ActivityTaskScheduled(credit_score)', 'ActivityTaskCompleted', 'WorkflowTaskStarted', 'ActivityTaskScheduled(auto_approve)', 'ActivityTaskCompleted', 'WorkflowTaskStarted', 'TimerStarted(86400s)', 'TimerFired', 'WorkflowTaskStarted', 'WorkflowExecutionCompleted']The buggy workflow picked a branch from the wall clock. Worker A saw 16:59 and recorded "schedule auto_approve". Worker B replayed at 17:01, took the other branch, and asked for manual_review at a point where history says auto_approve. The engine cannot reconcile the two, so the workflow is stuck (Temporal fails the workflow task and keeps retrying it until you deploy fixed code or reset the execution). The fix reads time through the SDK, which returns the time recorded in history. The Temporal SDKs provide these replay-safe APIs: workflow.now() and workflow.random() in Python, workflow.Now(ctx) in Go, and in TypeScript the workflow sandbox replaces Date and Math.random with deterministic versions.
Diagram 1: crash, replay and continue
Decisions
- 1
Step 1: client starts workflow with Workflow ID order-42
- nextStep 2: server appends WorkflowExecutionStarted and schedules a workflow task
- 2
Step 2: server appends WorkflowExecutionStarted and schedules a workflow task
- nextStep 3: worker runs workflow code until it blocks, returns commands such as ScheduleActivity
- 3
Step 3: worker runs workflow code until it blocks, returns commands such as ScheduleActivity
- nextStep 4: server appends ActivityTaskScheduled and puts the activity on a task queue
- 4
Step 4: server appends ActivityTaskScheduled and puts the activity on a task queue
- nextStep 5: an activity worker runs the side effect and reports the result
- 5
Step 5: an activity worker runs the side effect and reports the result
- nextStep 6: server appends ActivityTaskCompleted and schedules the next workflow task
- 6
Step 6: server appends ActivityTaskCompleted and schedules the next workflow task
- nextStep 7: is the workflow still cached on a live worker?
- ?
Step 7: is the workflow still cached on a live worker?
- nextStep 8: continue from memory, issue next commands
- nextStep 8b: new worker replays code from the start against history
- 8
Step 8: continue from memory, issue next commands
- nextStep 10: workflow returns, server appends WorkflowExecutionCompleted
- 9
Step 8b: new worker replays code from the start against history
- nextStep 9: do replayed commands match history?
- ?
Step 9: do replayed commands match history?
- nextStep 8: continue from memory, issue next commands
- nextFailure path: non-determinism error, workflow task retries until code is fixed or execution is reset
- 11
Failure path: non-determinism error, workflow task retries until code is fixed or execution is reset
- 12
Step 10: workflow returns, server appends WorkflowExecutionCompleted
Lesson map
Durable Execution & the Temporal Model - Workflows vs Activities, Event History, Replay & Determinism
Workflows vs activities, event history and replay, determinism rules (SDK time/random, no I/O in workflow code) with a runnable replay engine showing a wall-clock non-determinism bug and the fix; timers, signals, queries and updates (runnable race demo); task queues and workers; retry policy defaults and the four activity timeouts; what exactly-once means (and does not) with idempotency keys; history and payload limits; workflow-vs-activity decision chart.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Step 1: client starts workflow with Workflow ID order-42"] b["Step 2: server appends WorkflowExecutionStarted and schedules a workflow task"] c["Step 3: worker runs workflow code until it blocks, returns commands such as ScheduleActivity"] d["Step 4: server appends ActivityTaskScheduled and puts the activity on a task queue"] e["Step 5: an activity worker runs the side effect and reports the result"] f["Step 6: server appends ActivityTaskCompleted and schedules the next workflow task"] g["Step 7: is the workflow still cached on a live worker?"] h["Step 8: continue from memory, issue next commands"] r["Step 8b: new worker replays code from the start against history"] m["Step 9: do replayed commands match history?"] x["Failure path: non-determinism error, workflow task retries until code is fixed or execution is reset"] z["Step 10: workflow returns, server appends WorkflowExecutionCompleted"] a -->|continues| b b -->|continues| c c -->|continues| d d -->|continues| e e -->|continues| f f -->|continues| g g -->|continues| h g -->|continues| r r -->|continues| m m -->|continues| h m -->|continues| x h -->|continues| z
What makes workflow code non-deterministic
| Source | Why it breaks replay | Replay-safe replacement |
|---|---|---|
System clock (datetime.now(), Date.now() outside the sandbox) | Different value on replay changes branches and timer durations | SDK time API (workflow.now()), or pass time in as an argument |
| Random numbers, UUIDs | New values on replay | SDK random and UUID helpers, or an activity |
| Network, database, file I/O | Results change; also a side effect repeated on every replay | Activities |
| Iterating unordered collections (in some languages) | Order differs between processes | Sort first |
| Threads, native async outside the SDK scheduler | Interleaving differs | SDK coroutines, workflow.wait_condition, child workflows |
| Reading mutable globals or config files | Value differs per process | Pass as input, or fetch in an activity |
| Changing code while executions are in flight | Old histories meet new commands | Patching or Worker Versioning (next page) |
Temporal's docs list which edits are safe without versioning: changing activity inputs, return values and timeouts, timer durations (mostly), and adding non-command calls; and which are not: adding, removing or reordering anything that produces a command (activities, timers, child workflows, signals to external workflows, completing or continuing-as-new).
Timers, signals, queries and updates
// Durable waits, messages and at-least-once activities in a replay engine.
// - A workflow waits for a Signal OR a durable Timer, whichever comes first.
// - A Query reads state by replaying history; it never appends events.
// - A Signal that arrives while no worker is alive is still recorded in history.
// - An activity whose worker dies after the side effect is retried after its
// Start-To-Close timeout: activities are at-least-once, so the side effect
// must be idempotent (here: an idempotency key the payment provider dedups on).
type Ev = { t: string; type: string; data?: string };
type Cmd = { kind: "race"; signal: string; timer: string } | { kind: "activity"; name: string; key: string };
interface Ctx { state: { status: string } }
function* expense(ctx: Ctx): Generator<Cmd, string, string> {
ctx.state.status = "waiting-for-approval";
const r = yield { kind: "race", signal: "approve", timer: "72h" };
if (r === "timer") { ctx.state.status = "auto-rejected"; return "rejected"; }
ctx.state.status = `paying (approved by ${r})`;
const receipt = yield { kind: "activity", name: "pay", key: "expense-981" };
ctx.state.status = "paid";
return receipt;
}
// Replay the code against history; return state plus the first command that has no result yet.
function replay(h: Ev[]): { state: Ctx["state"]; pending?: Cmd; result?: string } {
const ctx: Ctx = { state: { status: "new" } }; const g = expense(ctx); let send: string | undefined;
for (;;) {
const step = g.next(send as string);
if (step.done) return { state: ctx.state, result: step.value };
const c = step.value;
if (c.kind === "race") {
const sig = h.find((e) => e.type === "SignalReceived"); const fired = h.find((e) => e.type === "TimerFired");
if (!sig && !fired) return { state: ctx.state, pending: c };
send = sig ? sig.data! : "timer";
} else {
const done = h.find((e) => e.type === "ActivityTaskCompleted");
if (!done) return { state: ctx.state, pending: c };
send = done.data!;
}
}
}
function runScenario(useIdempotencyKey: boolean) {
const h: Ev[] = [{ t: "0h", type: "WorkflowExecutionStarted" }];
const ledger: string[] = []; const seenKeys = new Set<string>();
const provider = (key: string) => { // the external payment API
if (useIdempotencyKey && seenKeys.has(key)) return "rcpt-1 (deduplicated)";
seenKeys.add(key); ledger.push("charge $120"); return `rcpt-${ledger.length}`;
};
// worker 1 starts the workflow: the race command creates a durable timer
let r = replay(h);
if (r.pending?.kind === "race") h.push({ t: "0h", type: "TimerStarted", data: r.pending.timer });
console.log(` 0h query status -> "${replay(h).state.status}" (history length still ${h.length})`);
console.log(" 5h worker 1 crashes; no worker is running");
h.push({ t: "6h", type: "SignalReceived", data: "manager-lee" });
console.log(" 6h signal 'approve' accepted by the server and recorded while no worker is alive");
r = replay(h); // worker 2 replays from history
if (r.pending?.kind === "activity") {
h.push({ t: "7h", type: "TimerCanceled" }, { t: "7h", type: "ActivityTaskScheduled", data: r.pending.name });
provider(r.pending.key); // attempt 1 charges the card...
console.log(" 7h attempt 1 charged the card, then its worker died before reporting");
console.log(" 7h+ Start-To-Close timeout expires -> retry policy schedules attempt 2");
const receipt = provider(r.pending.key); // attempt 2
h.push({ t: "7h", type: "ActivityTaskStarted", data: "attempt=2" }, { t: "7h", type: "ActivityTaskCompleted", data: receipt });
}
r = replay(h); h.push({ t: "7h", type: "WorkflowExecutionCompleted", data: r.result });
console.log(` done: result=${r.result}, status="${r.state.status}", charges=${ledger.length}`);
console.log(` history: ${h.map((e) => e.type).join(", ")}`);
}
console.log("Without an idempotency key:"); runScenario(false);
console.log("With an idempotency key:"); runScenario(true);Output:
Without an idempotency key:
0h query status -> "waiting-for-approval" (history length still 2)
5h worker 1 crashes; no worker is running
6h signal 'approve' accepted by the server and recorded while no worker is alive
7h attempt 1 charged the card, then its worker died before reporting
7h+ Start-To-Close timeout expires -> retry policy schedules attempt 2
done: result=rcpt-2, status="paid", charges=2
history: WorkflowExecutionStarted, TimerStarted, SignalReceived, TimerCanceled, ActivityTaskScheduled, ActivityTaskStarted, ActivityTaskCompleted, WorkflowExecutionCompleted
With an idempotency key:
0h query status -> "waiting-for-approval" (history length still 2)
5h worker 1 crashes; no worker is running
6h signal 'approve' accepted by the server and recorded while no worker is alive
7h attempt 1 charged the card, then its worker died before reporting
7h+ Start-To-Close timeout expires -> retry policy schedules attempt 2
done: result=rcpt-1 (deduplicated), status="paid", charges=1
history: WorkflowExecutionStarted, TimerStarted, SignalReceived, TimerCanceled, ActivityTaskScheduled, ActivityTaskStarted, ActivityTaskCompleted, WorkflowExecutionCompletedExpectedWithout an idempotency key: 0h query status -> "waiting-for-approval" (history length still 2) 5h worker 1 crashes; no worker is running 6h signal 'approve' accepted by the server and recorded while no worker is alive 7h attempt 1 charged the card, then its worker died before reporting 7h+ Start-To-Close timeout expires -> retry policy schedules attempt 2 done: result=rcpt-2, status="paid", charges=2 history: WorkflowExecutionStarted, TimerStarted, SignalReceived, TimerCanceled, ActivityTaskScheduled, ActivityTaskStarted, ActivityTaskCompleted, WorkflowExecutionCompleted With an idempotency key: 0h query status -> "waiting-for-approval" (history length still 2) 5h worker 1 crashes; no worker is running 6h signal 'approve' accepted by the server and recorded while no worker is alive 7h attempt 1 charged the card, then its worker died before reporting 7h+ Start-To-Close timeout expires -> retry policy schedules attempt 2 done: result=rcpt-1 (deduplicated), status="paid", charges=1 history: WorkflowExecutionStarted, TimerStarted, SignalReceived, TimerCanceled, ActivityTaskScheduled, ActivityTaskStarted, ActivityTaskCompleted, WorkflowExecutionCompleted
Press Run. Snippets must be self-contained — no network, files, or native modules.
(The history above is simplified: real histories also contain workflow task events. Temporal also writes an activity's ActivityTaskStarted event only when the activity finishes, carrying the final attempt number, which is why intermediate retries do not bloat history.)
| Primitive | Direction | Writes to history? | Blocks the caller? | Use for |
|---|---|---|---|---|
Durable timer (workflow.sleep, await sleep) | Workflow to itself | Yes (started, fired) | n/a | Waits of seconds to months without holding a thread |
| Signal | Client to workflow, async | Yes | No; fire and forget | Approvals, external events, "cancel this order" |
| Query | Client reads workflow | No | Briefly; cannot change state or block on conditions | Status pages, debugging |
| Update | Client to workflow, sync | Yes, if accepted | Yes; returns a result or error, optional validator can reject | "Add item to cart and tell me the new total" |
| Signal-With-Start | Client | Yes | No | Start the workflow if not running, then deliver the signal |
Task queues and workers
Workers poll named task queues; the server never pushes to them. Workflow tasks and activity tasks can share a queue or use separate ones (for example, a GPU-only activity queue). Polling gives natural back-pressure: if workers are saturated, tasks simply wait in the queue, and the Schedule-To-Start latency metric shows it. Scaling is adding worker processes.
Retries and the four activity timeouts
| Timeout | Measures | Default | If it fires |
|---|---|---|---|
| Schedule-To-Start | Time an activity task waits in the queue for a worker | Unlimited | Not retried (another worker on the same queue would not help); fix capacity or routing |
| Start-To-Close | One attempt's maximum duration | None of its own; you must set it or Schedule-To-Close (if only Schedule-To-Close is set, it is used) | Attempt fails and the retry policy decides |
| Schedule-To-Close | The whole activity including all retries | Unlimited | Activity fails for good |
| Heartbeat | Maximum gap between heartbeats from a long-running attempt | None | Attempt fails fast instead of waiting for Start-To-Close; last heartbeat details let the next attempt resume |
Workflows have their own limits: Workflow Execution Timeout (whole chain including retries and continue-as-new), Workflow Run Timeout (one run), and Workflow Task Timeout (how long a worker may take to process one workflow task; 10 seconds by default). Set Start-To-Close on every activity, add Heartbeat for anything longer than a minute or so, and use non-retryable error types for business failures ("card declined") so they are not retried forever under the default unlimited policy.
What exactly-once means here (and what it does not)
| Claim | True? | Why |
|---|---|---|
| The workflow's code path runs to completion exactly once logically | Yes | Replays re-derive the same decisions; recorded results are reused |
| Each command is recorded once in history | Yes | The server deduplicates and appends atomically |
| An activity's side effect happens exactly once | No | A worker can perform the effect and die before reporting; the retry repeats it |
| A signal delivered once is processed once | Yes, within the workflow | It is an event in history; replay re-delivers it deterministically |
| Starting a workflow twice with the same Workflow ID creates two | No, while one is running | Workflow ID uniqueness plus the ID reuse and conflict policies |
So end-to-end exactly-once effects still need the classic tools: an idempotency key passed to the downstream API (derive it from Workflow ID plus activity ID), an upsert, or a transactional outbox on the receiving side. AWS documents the same split for Step Functions: Standard workflows have exactly-once workflow execution, but tasks re-run when you configure Retry, and asynchronous Express workflows are at-least-once.
Illustration: a Temporal Python workflow (not executed here)
import asyncio
from datetime import timedelta
from temporalio import workflow
from temporalio.common import RetryPolicy
with workflow.unsafe.imports_passed_through():
from activities import credit_score, auto_approve, manual_review, pay
@workflow.defn
class LoanWorkflow:
def __init__(self) -> None:
self.approved_by: str | None = None
@workflow.run
async def run(self, applicant: str) -> str:
opts = dict(start_to_close_timeout=timedelta(seconds=30),
retry_policy=RetryPolicy(maximum_attempts=5, non_retryable_error_types=["CardDeclined"]))
score = await workflow.execute_activity(credit_score, applicant, **opts)
if score > 700 or workflow.now().hour < 17: # replay-safe clock
return await workflow.execute_activity(auto_approve, applicant, **opts)
try: # durable wait: signal or 3 days
await workflow.wait_condition(lambda: self.approved_by is not None, timeout=timedelta(days=3))
except asyncio.TimeoutError:
return "rejected"
return await workflow.execute_activity(pay, f"loan-{workflow.info().workflow_id}", **opts)
@workflow.signal
def approve(self, manager: str) -> None:
self.approved_by = manager
@workflow.query
def status(self) -> str:
return "approved" if self.approved_by else "waiting"Diagram 2: workflow code or activity?
Decisions
- 1
Step 1: look at one line of logic
- nextStep 2: does it touch the outside world (network, disk, DB, clock, randomness)?
- ?
Step 2: does it touch the outside world (network, disk, DB, clock, randomness)?
- nextStep 3: is it the time or a random value?
- nextStep 4: is it a wait for time or for an external input?
- ?
Step 3: is it the time or a random value?
- nextUse the SDK API such as workflow.now or workflow.random, recorded or deterministic
- nextPut it in an activity with timeouts, retry policy and an idempotency key
- 4
Use the SDK API such as workflow.now or workflow.random, recorded or deterministic
- 5
Put it in an activity with timeouts, retry policy and an idempotency key
- ?
Step 4: is it a wait for time or for an external input?
- nextDurable timer, signal, update or wait condition in workflow code
- nextStep 5: is it pure, deterministic computation on workflow state?
- 7
Durable timer, signal, update or wait condition in workflow code
- ?
Step 5: is it pure, deterministic computation on workflow state?
- nextKeep it in workflow code
- nextFailure path: hidden nondeterminism such as unordered iteration or global state, refactor before shipping
- 9
Keep it in workflow code
- 10
Failure path: hidden nondeterminism such as unordered iteration or global state, refactor before shipping
What happens if you choose otherwise
- Doing I/O inside workflow code: it repeats on every replay and can return different results, causing non-determinism errors. Wrap it in an activity.
- One giant activity that does everything: you are back to task-level checkpoints, so a crash redoes all of it. Split at side-effect boundaries.
- No Start-To-Close timeout and no heartbeat on a long activity: a dead worker is detected only by Schedule-To-Close, which may be unlimited.
- Hand-rolled state machine in a database plus a cron poller: works, but you own timeouts, retries, visibility and races between pollers. That is the system durable execution packages.
Pitfalls
- Treating workflow code like request handlers: no
requests.get, no ORM calls, notime.time(). - Large payloads: every activity argument and result is stored in history; Temporal warns at 256 KB per payload and errors at 2 MB. Pass references.
- Unlimited retries on business errors: mark them non-retryable.
- Assuming queries are free on huge histories: a query on an uncached workflow replays its whole history.
Interview Q&A
Why must Temporal workflow code be deterministic?
Answer
Because the engine recovers state by replaying the code against the recorded event history. If replay issues different commands than the history contains, the engine cannot tell what really happened, so it raises a non-determinism error.
Are Temporal activities exactly-once?
Answer
No. They are at-least-once: a worker can finish the side effect and crash before the completion is recorded, and the retry runs it again. Use idempotency keys or idempotent writes. The workflow's decisions, on the other hand, are recorded exactly once.
What is the difference between Start-To-Close, Schedule-To-Close and Heartbeat timeouts?
Answer
Start-To-Close bounds one attempt; Schedule-To-Close bounds the whole activity across retries; Heartbeat bounds the silence between heartbeats so a dead worker on a long task is detected quickly. Schedule-To-Start bounds queue wait and signals missing worker capacity.
Signal vs Update vs Query?
Answer
A signal is an async write recorded in history with no response. An update is a sync, tracked write that returns a result or error and can be validated before acceptance. A query is a read that does not touch history and cannot block.
How can a workflow sleep for 30 days without a thread?
Answer
A durable timer is a history event plus a server-side timer. No worker holds anything while it waits; when it fires, a workflow task is scheduled and any worker replays or resumes the workflow.
What is Signal-With-Start for?
Answer
Starting a workflow if it is not running and delivering a signal in one call, so the first event is never lost.
Why does Temporal not retry Schedule-To-Start timeouts?
Answer
Another worker on the same queue would not help. The fix is more capacity or correct routing.
How should payloads be sized?
Answer
Small. Temporal warns at 256 KB per payload and errors at 2 MB, so pass references to stored data.
Check yourself
Take a workflow you know and list every line that reads time, randomness, configuration or the network. Move each into an SDK call or an activity, and write the idempotency key each activity would send.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Event Types — Notification, Event-Carried State Transfer & Event Sourcing, At-Least-Once vs Exactly-Once Delivery, Delivery Semantics — At-Least-Once, At-Most-Once & Exactly-Once, Retry Storms, Backoff & Jitter, Timeouts, Budgets & Deadline Propagation — End-to-End Latency Caps, Time, Clocks & Ordering in Distributed Systems - Physical Clocks, Lamport, Vector Clocks, HLC & TrueTime, Durable Objects Real-Time - WebSocket Hibernation, Alarms, Chat, Presence & Collaborative Editing.
Go Deeper
- Temporal: Workflow Definition and deterministic constraints
- Temporal: Event History
- Temporal: Activity timeouts (detecting failures)
- Temporal: Retry Policies
- Temporal: Signals, Queries and Updates
- Temporal: Workflow ID and Run ID
- Temporal: Self-hosted defaults and payload limits
- Azure Durable Functions: Orchestrator code constraints