Distributed systems
Part 3 of 6 · Durable ObjectsDurable Objects Concurrency - Single Thread, Input & Output Gates, blockConcurrencyWhile & the Races That Remain
Concurrency: single thread plus input and output gates, where interleaving still happens (fetch, timers, other objects), blockConcurrencyWhile cost, the external-call oversell race reproduced under wrangler dev (reserved 6 of 3) and fixed with claim-first and version check.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Where can requests interleave in a Durable Object?
Answer
Only at await points.
L2
What does the input gate guarantee?
Answer
No new events are delivered while the object waits on storage.
L3
What does the output gate guarantee?
Answer
No outgoing message leaves until the writes before it are durable.
L4
Which awaits are not protected?
Answer
fetch, R2, timers and calls to other objects.
L5
How did the reproduction reserve 6 of 3?
Answer
Each request checked stock, awaited an external call, then reserved, so all six saw stock before any wrote.
L6
Name two fixes.
Answer
Claim first with one synchronous conditional UPDATE, or compare-and-set on a version.
L7
When is blockConcurrencyWhile right?
Answer
Constructor-time initialization and migrations, or rare short sections you cannot restructure.
Failure modes
Check, await external, act oversells inventory
Requests interleave during the payment or HTTP call, and each acts on a stale check.
blockConcurrencyWhile around slow calls stalls the object
Every request queues behind the section, and a throw or 30 s timeout resets the object.
A claim that is never released leaks stock
A crash after claiming but before paying leaves units reserved without a TTL or alarm.
Misconceptions
Single-threaded code cannot race.
JavaScript interleaves at await, so non-storage awaits open the door to other requests.
Input gates protect every await.
They only cover storage operations.
blockConcurrencyWhile is the safe default for handlers.
It serializes everything behind it and should be rare outside initialization.
Interviewer traps
Holding a check across a payment call.
Claim the unit first, then pay with an idempotency key, then release on failure.
Copying Postgres SELECT FOR UPDATE habits.
There is no lock to hold. Make the critical part synchronous instead.
Design scenario
Same prompt for every reader.
Requirements
Never sell more than stock, never double-charge, and release stuck reservations after crashes.
Failure assumptions
- Payment calls take hundreds of milliseconds.
- Clients retry on timeout.
- The object can reset mid-request.
Constraints
- No global lock service.
- blockConcurrencyWhile only for initialization.
Prompt
Sell the last concert tickets through a Durable Object that must call a payment API before confirming.
API
What does reserve return, and which idempotency key reaches the payment API?
Data
Which columns support claim-first and version checks?
Architecture
Where do TTL reclamation alarms and payment retries run?
Overview
A Durable Object runs your code on one thread, like a browser tab. Synchronous code always runs to completion without interruption, so a check followed by a write with no await in between is atomic for free. The interesting part is await. JavaScript lets other requests run while one is waiting, which is where races come from. Cloudflare adds two rules. The input gate stops new events from being delivered while you wait on storage. The output gate holds your outgoing messages until your writes are durable. Together they make most storage code correct by default. They do not protect you while you wait on anything else: fetch(), R2, another object, a timer. A check, then an external call, then an act-on-the-check is still a race, and below we reproduce it on real Durable Objects.
How the code was checked: The real Durable Object TypeScript on this page type-checks with
tsc --strictagainst@cloudflare/workers-typesand was run locally underwrangler dev4.148.0 (workerd 2026-10-06) and the@cloudflare/vitest-plugintest pool. Blocks labeled simulation are sandbox models of the semantics, not Cloudflare code. Limits and prices are quoted from Cloudflare's docs as of 2026-10-07.
Why this matters
- Interviews: "Single-threaded means no races, right?" is a trap. You need to explain interleaving at await points, what the gates do, and two standard fixes (claim first, or compare-and-set on a version).
- Production: overselling inventory, double-booking a seat or double-charging a card in a DO almost always comes from awaiting an external call between the check and the write.
The rules, precisely
| Mechanism | What it does | What it does not do |
|---|---|---|
| Single thread | One piece of your JavaScript runs at a time per object | Stop interleaving at await |
| Input gate | While you await a storage operation, no other request, response or event is delivered to the object | Cover fetch(), R2, KV, other objects, timers: the gate opens and others may run |
| Output gate | Outgoing messages (replies, fetches) wait until pending writes are confirmed; if the write fails, they are discarded and the object resets | Make external side effects transactional |
| Write coalescing | Writes with no intervening await commit as one atomic batch | Span an external call |
blockConcurrencyWhile(fn) | No other events until fn finishes, even across external awaits; 30 s timeout; a throw resets the object | Scale: it is a global lock on the object |
allowConcurrency: true (KV reads) | Opts a read out of the input gate | Anything safe by default |
SQLite calls (sql.exec) are synchronous, so they never yield at all. A whole check-then-update in SQL is atomic without any gate.
How do you keep a check valid across an external call?
Prefer
Claim first, or compare-and-set
Make the decisive state change synchronous, then do the slow work.
- One conditional UPDATE claims the unit before the await.
- A version check rejects stale writers.
- Throughput stays high because nothing blocks the object.
Alternative
Check, await, then act
Other requests run during the await and act on the same stale check.
- Six requests reserved three units.
- Retries can double-charge without idempotency keys.
- Wrapping it in blockConcurrencyWhile stalls every request.
Where the gates help and where they stop
Diagram 1 condensed. The real reproduction and the simulation show both sides.
- 1
Run synchronous code
Nobody can interrupt it, so check-then-write with no await is atomic. - 2
Await storage
The input gate holds new events until storage returns. - 3
Await something else
fetch, R2, timers and other objects let other requests run. - 4
Restructure the method
Claim first or compare-and-set so the critical part stays synchronous.
Where interleaving happens (step-labeled)
Sequence
- 1
Request 1 → Object (one thread)
1. reserve(): read qty = 3 (sync SQL)
- 2
Request 2 → Object (one thread)
2. reserve() arrives, queued behind R1
- 3
Object (one thread) → External API
3. R1 awaits fetch to payment API (input gate OPENS)
- 4
Object (one thread) → Object (one thread)
4. R2 runs: reads qty = 3 too
- 5
Object (one thread) → External API
5. R2 awaits fetch as well
- 6
External API → Object (one thread)
6. R1 resumes, writes qty = 2
- 7
External API → Object (one thread)
7. R2 resumes, writes qty = 2 (stale)
- 8
Object (one thread)
8. FAILURE: two units sold, stock dropped by one
- 9
Object (one thread)
FIX: claim with one sync UPDATE ... WHERE qty > 0 before step 3
Lesson map
Durable Objects Concurrency - Single Thread, Input & Output Gates, blockConcurrencyWhile & the Races That Remain
Concurrency: single thread plus input and output gates, where interleaving still happens (fetch, timers, other objects), blockConcurrencyWhile cost, the external-call oversell race reproduced under wrangler dev (reserved 6 of 3) and fixed with claim-first and version check.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB r1["Request 1"] o["Object (one thread)"] r2["Request 2"] x["External API"] r1 -->|1. reserve(): read qty = 3 (sync SQL)| o r2 -->|2. reserve() arrives, queued behind R1| o o -->|3. R1 awaits fetch to payment API (input gate OPENS)| x o -->|5. R2 awaits fetch as well| x x -->|6. R1 resumes, writes qty = 2| o x -->|7. R2 resumes, writes qty = 2 (stale)| o
Real reproduction on Durable Objects
The Inventory object below has three versions of "reserve one unit". The external call is a 50 ms timer standing in for a payment API. Like fetch(), it is not storage I/O, so the input gate opens while it waits.
// Inventory: demonstrates the one race input gates do NOT stop - awaiting non-storage I/O.
import { DurableObject } from "cloudflare:workers";
import type { Env } from "./index";
// Stand-in for a payment / fraud API call. A timer, like fetch(), is not storage I/O,
// so the input gate opens and other requests may run while we wait.
const externalCall = (ms: number) => new Promise<void>((r) => setTimeout(r, ms));
export class Inventory extends DurableObject<Env> {
constructor(ctx: DurableObjectState, env: Env) {
super(ctx, env);
ctx.blockConcurrencyWhile(async () => {
ctx.storage.sql.exec("CREATE TABLE IF NOT EXISTS stock(id INTEGER PRIMARY KEY CHECK(id = 1), qty INTEGER NOT NULL, version INTEGER NOT NULL)");
});
}
reset(qty: number): void {
this.ctx.storage.sql.exec("INSERT INTO stock VALUES (1, ?, 0) ON CONFLICT(id) DO UPDATE SET qty = excluded.qty, version = 0", qty);
}
qty(): number { return this.ctx.storage.sql.exec("SELECT qty FROM stock").one().qty as number; }
// BUG: check, await external I/O, then act on the stale check.
async reserveRacy(): Promise<string> {
const qty = this.qty();
if (qty <= 0) return "sold-out";
await externalCall(50); // other reserveRacy calls interleave here
this.ctx.storage.sql.exec("UPDATE stock SET qty = ? WHERE id = 1", qty - 1); // writes a stale value
return "reserved";
}
// FIX 1: claim first (synchronous, atomic), then do the slow call, and compensate on failure.
async reserveClaimFirst(): Promise<string> {
const row = this.ctx.storage.sql
.exec("UPDATE stock SET qty = qty - 1, version = version + 1 WHERE id = 1 AND qty > 0 RETURNING qty").toArray();
if (row.length === 0) return "sold-out";
try {
await externalCall(50);
return "reserved";
} catch (e) {
this.ctx.storage.sql.exec("UPDATE stock SET qty = qty + 1 WHERE id = 1"); // compensation
throw e;
}
}
// FIX 2: optimistic check-and-set - remember the version, re-check it after the await.
async reserveCas(): Promise<string> {
for (let attempt = 0; attempt < 5; attempt++) {
const { qty, version } = this.ctx.storage.sql.exec("SELECT qty, version FROM stock").one() as { qty: number; version: number };
if (qty <= 0) return "sold-out";
await externalCall(50);
const updated = this.ctx.storage.sql
.exec("UPDATE stock SET qty = qty - 1, version = version + 1 WHERE id = 1 AND version = ? RETURNING qty", version).toArray();
if (updated.length === 1) return "reserved";
// someone else won the race; loop and re-read
}
return "conflict-retry-later";
}
}Real run: six buyers hit the same SKU object at once with 3 units in stock, under local wrangler dev:
Output (captured from the real run):
racy results={"reserved":6} final qty=2
claim results={"reserved":3,"sold-out":3} final qty=0
cas results={"reserved":3,"sold-out":3} final qty=0The racy version "reserved" six units and left the stock at 2: three oversold units, and a final count that is wrong too. Both fixes sell exactly three. The same assertion passes inside the Vitest plugin's workerd test pool (see the production page).
The same behavior, simulated
asyncio gives the same single thread and the same interleaving at await, so the semantics can be modeled outside Cloudflare (simulation, not Cloudflare code):
# SIMULATION (not Cloudflare code) of a Durable Object's single thread, input gate and the external-I/O race.
# asyncio gives us one thread with interleaving at every await, exactly like JavaScript in the object.
import asyncio
class Inventory:
def __init__(self, qty): self.qty, self.version = qty, 0
async def storage_read(self):
# Storage I/O: the input gate stays CLOSED, so no other request is delivered meanwhile.
return self.qty, self.version
async def external_call(self):
await asyncio.sleep(0.01) # fetch()/timer: the input gate OPENS, other requests run here
async def racy(self):
qty, _ = await self.storage_read()
if qty <= 0: return "sold-out"
await self.external_call() # 6 buyers all passed the check above
self.qty = qty - 1 # stale write: everyone writes 2
return "reserved"
async def claim_first(self):
if self.qty <= 0: return "sold-out" # check and decrement with no await between = atomic
self.qty -= 1; self.version += 1
await self.external_call()
return "reserved"
async def cas(self):
for _ in range(5):
qty, ver = await self.storage_read()
if qty <= 0: return "sold-out"
await self.external_call()
if self.version == ver: # nobody changed it while we waited
self.qty -= 1; self.version += 1; return "reserved"
return "conflict-retry-later"
async def run(mode):
inv = Inventory(3)
res = await asyncio.gather(*(getattr(inv, mode)() for _ in range(6)))
tally = {k: res.count(k) for k in sorted(set(res))}
print(f"{mode:<11} results={tally} final qty={inv.qty}")
for m in ["racy", "claim_first", "cas"]:
asyncio.run(run(m))Output (simulation):
racy results={'reserved': 6} final qty=2
claim_first results={'reserved': 3, 'sold-out': 3} final qty=0
cas results={'reserved': 3, 'sold-out': 3} final qty=0blockConcurrencyWhile: correct, and expensive
blockConcurrencyWhile makes the object behave like it holds a mutex across awaits. The docs recommend it for constructor-time setup (schema migrations, loading state), not for request handling. If each critical section takes 5 ms, that object tops out near 200 requests per second, and holding it across a slow fetch() is like holding a lock across I/O. Simulation of the ordering and the ceiling (not Cloudflare code):
// SIMULATION (not Cloudflare code): where interleaving happens, and what blockConcurrencyWhile costs.
const log: string[] = [];
const externalCall = (ms: number) => new Promise<void>((r) => setTimeout(r, ms)); // like fetch(): gate opens
async function handler(name: string) {
log.push(`${name}: start (sync code runs to completion, nobody can interrupt)`);
await externalCall(10); // other handlers may run during this await
log.push(`${name}: resumed after external call`);
}
// Tiny stand-in for blockConcurrencyWhile: one event at a time, even across awaits.
let chain: Promise<unknown> = Promise.resolve();
function blockConcurrencyWhile<T>(fn: () => Promise<T>): Promise<T> {
const run = chain.then(fn); chain = run.catch(() => undefined); return run;
}
async function main() {
await Promise.all([handler("req1"), handler("req2")]);
console.log("plain awaits:\n " + log.join("\n "));
log.length = 0;
await Promise.all(["req1", "req2"].map((n) => blockConcurrencyWhile(() => handler(n))));
console.log("inside blockConcurrencyWhile:\n " + log.join("\n "));
// Throughput ceiling when every request holds the object for its critical section.
for (const ms of [1, 5, 50, 300]) {
console.log(`critical section ${String(ms).padStart(3)} ms -> at most ${Math.floor(1000 / ms)} req/s for this object`);
}
}
main();
export {};Output (simulation):
plain awaits:
req1: start (sync code runs to completion, nobody can interrupt)
req2: start (sync code runs to completion, nobody can interrupt)
req1: resumed after external call
req2: resumed after external call
inside blockConcurrencyWhile:
req1: start (sync code runs to completion, nobody can interrupt)
req1: resumed after external call
req2: start (sync code runs to completion, nobody can interrupt)
req2: resumed after external call
critical section 1 ms -> at most 1000 req/s for this object
critical section 5 ms -> at most 200 req/s for this object
critical section 50 ms -> at most 20 req/s for this object
critical section 300 ms -> at most 3 req/s for this objectExpectedplain awaits: req1: start (sync code runs to completion, nobody can interrupt) req2: start (sync code runs to completion, nobody can interrupt) req1: resumed after external call req2: resumed after external call inside blockConcurrencyWhile: req1: start (sync code runs to completion, nobody can interrupt) req1: resumed after external call req2: start (sync code runs to completion, nobody can interrupt) req2: resumed after external call critical section 1 ms -> at most 1000 req/s for this object critical section 5 ms -> at most 200 req/s for this object critical section 50 ms -> at most 20 req/s for this object critical section 300 ms -> at most 3 req/s for this object
Press Run. Snippets must be self-contained — no network, files, or native modules.
Choosing a fix
Decisions
- 1
1. Does the operation await anything that is not storage?
- noSafe: gates plus sync SQL make it atomic
- yes2. Can you claim the resource first with one sync statement?
- 2
Safe: gates plus sync SQL make it atomic
- ?
2. Can you claim the resource first with one sync statement?
- yesClaim first, call out, compensate on failure
- no3. Is contention low?
- 4
Claim first, call out, compensate on failure
- nextMake the external call idempotent with a key
- ?
3. Is contention low?
- yesRead a version, call out, write only if the version is unchanged, retry
- no4. Is the external call fast and rare?
- 6
Read a version, call out, write only if the version is unchanged, retry
- nextMake the external call idempotent with a key
- ?
4. Is the external call fast and rare?
- yesblockConcurrencyWhile around it (accept the throughput cap)
- noSplit the work: queue the call, keep a pending state, finish in an alarm
- 8
blockConcurrencyWhile around it (accept the throughput cap)
- 9
Split the work: queue the call, keep a pending state, finish in an alarm
- nextMake the external call idempotent with a key
- 10
Make the external call idempotent with a key
| Fix | Correct under contention | Throughput | Complexity | Failure mode |
|---|---|---|---|---|
| Claim first plus compensation | Yes | High | Low | A crash after the claim leaves a unit held: add a TTL or reconcile in an alarm |
| Optimistic version check (CAS) | Yes | High at low contention, retry storms at high | Medium | Wasted external calls on conflict |
blockConcurrencyWhile | Yes | Capped at 1 / critical-section time | Low | All traffic to the object stalls behind a slow dependency |
| Do nothing ("it is single-threaded") | No | High | None | Oversell, double-charge |
Other concurrency facts worth knowing
- E-order: calls you make through one stub arrive in order. Calls through different stubs, or from different Workers, have no ordering guarantee.
- Constructor safety: wrap setup in
blockConcurrencyWhileso the first request never sees a half-built schema. - Errors inside
blockConcurrencyWhilereset the object. Catch inside if you want to keep running. - Alarms run like any other event: only one
alarm()runs at a time per object, but it can interleave with requests at await points. - Shared isolate memory: module-level variables can be shared by multiple objects in one isolate. Never keep per-object state in globals.
What happens if you choose otherwise
- Put a Redis lock (
SET NX PX) in front instead: you add a network hop and a lease that can expire mid-operation. You still need fencing. The object already serializes, so you only need to avoid awaiting between check and act. - Wrap every method in
blockConcurrencyWhile: correct, but a 300 ms partner API caps the object at about 3 requests per second. - Move the check into the external service: often best when the external system is the source of truth, like idempotent payment intents with an idempotency key.
Pitfalls
- Treating
await this.ctx.storage.get()andawait fetch()the same. The first keeps the input gate closed. The second opens it. - Holding a SQL cursor across an
await. Iteration after the await can see newer writes. - Forgetting that a failed write resets the object: in-memory fields set before the failure are gone.
- Retrying a non-idempotent external call after a reset. Pair external calls with idempotency keys.
Interview Q&A
If a Durable Object is single-threaded, how can it have a race?
Answer
Requests interleave at await. When the awaited thing is not storage, the input gate opens and another request can read the same state before the first one writes. Check, await external, then act equals a race.
What exactly do input gates and output gates guarantee?
Answer
Input gate: no new events are delivered while the object waits on storage, so storage read-modify-write is atomic. Output gate: no outgoing message leaves until the writes before it are durable, and if a write fails the messages are dropped and the object resets.
How would you sell the last ticket safely when you must call a payment API?
Answer
Claim the ticket first with one synchronous UPDATE ... WHERE remaining > 0 RETURNING. Then call the payment API with an idempotency key. On failure, release the claim. Add a TTL or an alarm to reclaim units stuck from crashes.
When is `blockConcurrencyWhile` the right call?
Answer
Constructor-time initialization and migrations, or rare short sections around an external call where you cannot restructure. Not routine request handling.
How does this compare with Postgres `SELECT ... FOR UPDATE`?
Answer
Same idea, a serialized critical section per entity. In Postgres the lock is held across the transaction, including any app-side external call you make while holding it, with deadlock risk. In a DO there is no lock to hold: you design the method so the critical part is synchronous.
Why did the racy reproduction end with quantity 2 after reserving 6?
Answer
Each request read the same stock before awaiting the external call, then reserved on stale data, so reservations and the final quantity disagree.
What happens if code inside blockConcurrencyWhile throws or runs past 30 seconds?
Answer
The object resets, and callers see errors before a fresh instance starts.
Does ordering on one stub prevent races between different callers?
Answer
No. E-order applies to calls on the same stub. Different callers still interleave at await points inside the object.
Check yourself
Find one handler in your code with a check, an external call and a write. Rewrite it as claim-first or compare-and-set, and add the idempotency key you would pass downstream.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Event Loop — Call Stack, Microtasks, Macrotasks & async/await, Reactor vs Proactor - How libuv, Nginx, Netty, Tokio & the Go Netpoller Work, Mutex vs RWLock, SlotWise Reservation Service Incident - HLD Debug & Fix Path, API Idempotency Keys, Lamport Clocks - Happens-Before, Logical Timestamps & Total Order.
Go Deeper
- Durable Objects: Easy, Fast, Correct - Choose three (input and output gates)
- Rules of Durable Objects: input/output gates, non-storage I/O races, blockConcurrencyWhile
- Durable Object State API (blockConcurrencyWhile, abort)
- SQLite storage API (transactionSync, cursor rules)
- Martin Kleppmann: How to do distributed locking