Distributed systems
Part 1 of 6 · Durable ObjectsCloudflare Durable Objects - Single-Instance Actors, Edge State & When to Use Them
Hub: a Durable Object is a single-threaded actor with its own SQLite, one live instance per ID worldwide; routing with getByName/idFromName/newUniqueId, stubs and RPC; DO vs KV, D1, R2, Queues, Redis, Postgres row locks, Orleans and Akka with a decision chart; runnable routing simulation; what happens if you pick the alternative.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is a Durable Object in one sentence?
Answer
A globally unique, single-threaded instance of a class, addressed by ID, with private transactional storage next to its code.
L2
How do callers on two continents reach the same object?
Answer
Both derive the same ID from the same name in the same namespace, and the platform routes both calls to where that ID lives.
L3
getByName vs newUniqueId?
Answer
Named IDs are deterministic and need a global check on first use. Unique IDs are random, faster on first use, and must be stored somewhere.
L4
Where is an object created and can it move?
Answer
Near the first caller or a location hint, and objects do not move after creation today.
L5
Why not Workers KV for a chat room?
Answer
KV is eventually consistent, so concurrent read-modify-write loses messages and other regions see stale lists.
L6
How does this compare with Orleans grains?
Answer
Same virtual actor idea. Orleans runs on your cluster with pluggable storage, Durable Objects run on Cloudflare with built-in SQLite and hibernating WebSockets.
L7
When is a Durable Object the wrong tool?
Answer
Global read-heavy config, large files, analytics across all entities, or one entity far above about 1,000 requests per second without a way to shard it.
Failure modes
One object for everything becomes a global bottleneck
A single object is one thread, so routing every request to it queues work and returns overloaded errors.
In-memory state vanishes on eviction or deploy
Fields not written to storage disappear when the object is evicted, moved or restarted.
KV used for ordered state loses updates
Two concurrent writers read the same value and one write silently disappears.
Misconceptions
Creating an ID or a stub creates the object.
The object comes to life on the first method call.
A Durable Object follows its user around the world.
Objects stay where they were created. Relocation is not available today.
Durable Objects replace every database.
They are per-entity owners. Cross-entity reporting still belongs in D1, Postgres or a warehouse fed by the objects.
Interviewer traps
Designing one global object for a hot counter or limiter.
Choose one object per room, user, tenant or document, and shard anything hotter than one object.
Treating in-memory fields as durable.
Persist state in SQLite as you go and rebuild caches in the constructor.
Design scenario
Same prompt for every reader.
Requirements
No double booking per venue, live updates to everyone viewing a venue, and reporting across all venues.
Failure assumptions
- Deploys restart every object and drop WebSockets.
- A few venues get launch-day spikes.
- Viewers are spread across continents.
Constraints
- No global lock service.
- Reporting must not query every object synchronously.
Prompt
Design a multi-tenant booking service where each venue's seat map must never double-book, with live seat updates for viewers.
API
Which name maps a request to its venue object, and which RPC methods does it expose?
Data
What lives in each object's SQLite, and what flows to a shared reporting store?
Architecture
Where do the router Worker, venue objects and reporting database sit?
Overview
A Durable Object is a tiny, single-threaded server with its own private SQLite database, and there is exactly one of it per ID in the whole world. You never start or stop it. You name it (getByName("room:42")), call a method on it, and Cloudflare routes the call to the one live instance, creating it near the first caller if it does not exist yet. Two callers in Tokyo and Chicago who use the same name reach the same object, so the object becomes the natural place to serialize everything that touches that one thing: a chat room, a document, a tenant's rate limit, a seat map.
That is the actor model, run as a managed service. The engineering question is always the same: what is the atom of coordination? Pick it well (one object per room, per user, per document) and you get strong consistency without locks and near-unlimited horizontal scale. Pick it badly (one object for everything) and you have rebuilt a single-threaded global bottleneck.
How the code was checked: The real Durable Object TypeScript on this page type-checks with
tsc --strictagainst@cloudflare/workers-typesand was run locally underwrangler dev4.148.0 (workerd 2026-10-06) and the@cloudflare/vitest-plugintest pool. Blocks labeled simulation are sandbox models of the semantics, not Cloudflare code. Limits and prices are quoted from Cloudflare's docs as of 2026-10-07.
Why this matters in interviews and in production
- Interviews: "Design a chat app / collaborative editor / rate limiter / booking system" all hit the same wall: where does the shared, ordered state live, and how do you avoid races without a global lock? Durable Objects are a concrete, modern answer you can compare with Redis, Postgres row locks, Raft groups and Orleans grains.
- Production: they remove a whole tier (sticky WebSocket servers, a Redis cluster for coordination, a lock service) but impose hard rules: one object handles roughly 500 to 1,000 simple requests per second, in-memory state disappears on eviction, deploys restart every object, and the object lives in one location.
The mental model: one actor per ID
| Property | Stateless Worker | Durable Object |
|---|---|---|
| Instances | Many, anywhere, per request | Exactly one live instance per ID, worldwide |
| State between requests | None you can rely on | In-memory fields (until eviction) plus private SQLite storage |
| Concurrency | Many requests in parallel on many machines | One thread; requests interleave only at await points |
| Location | Nearest data center to the user | Created near the first caller (or a location hint), then stays there |
| Addressed by | URL / route | DurableObjectId (64 hex digits) from a name, a random ID or a stored string |
| Lifecycle | Per request | Created lazily on first call, hibernates or is evicted when idle, restarted on deploy |
The docs describe each object as an actor: it receives messages (HTTP, RPC, WebSocket frames, alarms), runs them one at a time in its own context with its own storage, and sends messages out. Unlike Erlang or Akka you never spawn or supervise it, and unlike a database row it runs your code next to the data.
Where does ordered per-entity state live?
Prefer
One Durable Object per entity
Every update for one room, document or tenant reaches one single-threaded owner with its own SQLite.
- No lock anywhere, yet 200 concurrent increments lose nothing.
- Live WebSockets live next to the state.
- Scale out by adding objects, not by growing one.
Alternative
Shared store plus your own coordination
KV, Redis or Postgres can hold the data, but you add locks, sticky servers or pub/sub to order updates.
- KV loses concurrent read-modify-write updates.
- Postgres hot rows queue on locks.
- Redis means running and failing over a cluster.
From a name to the one live instance
The hub map. Sibling pages cover storage, concurrency, real-time, scaling and production.
- 1
Name the entity
getByName or idFromName turns room:42 into the same ID everywhere. - 2
Route to the one instance
The platform finds where the ID lives, or creates it near the first caller. - 3
Run one request at a time
The object serializes work and writes to its own SQLite. - 4
Shard what is too hot
One object is one thread. A global singleton returns overloaded errors.
Durable Objects vs the alternatives
| Option | Consistency | Where your code runs | Good at | Weak at | If you pick it instead of a DO |
|---|---|---|---|---|---|
| Durable Objects | Strong and serialized per object | Inside the object, next to its data | Per-entity coordination, WebSockets, per-entity schedules | Cross-entity queries, one hot object | (baseline) |
| Workers KV | Eventually consistent; the docs say changes can take 60 s or more to show up in other locations | In a Worker, data fetched over the network | Read-heavy config, sessions, flags | Counters, locks, read-modify-write | Lost updates and stale reads on anything that must be ordered |
| D1 | Strong within one database | Worker talks to the DB over the network | One relational database, migrations, HTTP API, ad hoc SQL | Per-entity isolation at huge scale; a D1 database maxes out at 10 GB | One shared DB serializes everyone; you add row locks and retries |
| R2 | Strong read-after-write per object | Worker | Blobs, media, backups | Small hot mutable state | Fine for files, wrong for coordination |
| Queues | At-least-once delivery | Consumer Worker | Async jobs, retries, batching | Interactive, ordered, stateful logic | Good companion, not a replacement |
| Redis (self-run or managed) | Single primary, in-memory, optional persistence | App servers call over the network | Fast counters, Lua scripts, pub/sub | Durability by default, per-tenant isolation, you operate it | You run and scale a cluster, and failover can lose acknowledged writes |
Postgres row locks (SELECT ... FOR UPDATE) | Serializable per row, ACID | App servers | Cross-entity transactions and reporting | Lock contention, connection limits, one region | Hot rows queue on locks, deadlocks need retries |
| Orleans grains / Akka actors | Single activation per grain (Orleans) or per entity with cluster sharding (Akka) | Your cluster | Rich actor frameworks, any cloud | You run the cluster, membership and storage | Same model, but you own capacity, placement and upgrades |
Rule of thumb: use a Durable Object when the state has a natural owner (one room, one doc, one tenant) and you need ordering or live connections. Put cross-entity reporting in D1 or Postgres fed by the objects, and keep blobs in R2.
How a request reaches an object (step-labeled)
Decisions
- 1
1. Client request
- next2. Nearest Worker (stateless router)
- 2
2. Nearest Worker (stateless router)
- next3. Derive ID: getByName / idFromName / newUniqueId / idFromString
- 3
3. Derive ID: getByName / idFromName / newUniqueId / idFromString
- next4. get() returns a stub at once (no network yet)
- 4
4. get() returns a stub at once (no network yet)
- next5. First call: where does this ID live?
- ?
5. First call: where does this ID live?
- exists6a. Route to the host data center
- new named ID6b. Global uniqueness check, create near caller or hint
- 6
6a. Route to the host data center
- next7. Constructor if cold, then blockConcurrencyWhile setup
- 7
6b. Global uniqueness check, create near caller or hint
- next6a. Route to the host data center
- 8
7. Constructor if cold, then blockConcurrencyWhile setup
- next8. RPC method runs single-threaded, reads and writes SQLite
- 9
8. RPC method runs single-threaded, reads and writes SQLite
- next9. Output gate holds the reply until the write is confirmed
- F1. Too many queued requestsOverloaded error: do not retry, shard instead
- 10
9. Output gate holds the reply until the write is confirmed
- next10. Reply to Worker, Worker replies to client
- F2. Write fails or object resetsReply replaced by error: recreate stub, retry only if idempotent
- 11
10. Reply to Worker, Worker replies to client
- 12
Overloaded error: do not retry, shard instead
- 13
Reply replaced by error: recreate stub, retry only if idempotent
Lesson map
Cloudflare Durable Objects - Single-Instance Actors, Edge State & When to Use Them
Hub: a Durable Object is a single-threaded actor with its own SQLite, one live instance per ID worldwide; routing with getByName/idFromName/newUniqueId, stubs and RPC; DO vs KV, D1, R2, Queues, Redis, Postgres row locks, Orleans and Akka with a decision chart; runnable routing simulation; what happens if you pick the alternative.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c["1. Client request"] w["2. Nearest Worker (stateless router)"] id["3. Derive ID: getByName / idFromName / newUniqueId / idFromString"] s["4. get() returns a stub at once (no network yet)"] r["5. First call: where does this ID live?"] h["6a. Route to the host data center"] u["6b. Global uniqueness check, create near caller or hint"] k["7. Constructor if cold, then blockConcurrencyWhile setup"] m["8. RPC method runs single-threaded, reads and writes SQLite"] g["9. Output gate holds the reply until the write is confirmed"] ok["10. Reply to Worker, Worker replies to client"] ov["Overloaded error: do not retry, shard instead"] er["Reply replaced by error: recreate stub, retry only if idempotent"] c -->|continues| w w -->|continues| id id -->|continues| s s -->|continues| r r -->|exists| h r -->|new named ID| u u -->|continues| h h -->|continues| k k -->|continues| m m -->|continues| g g -->|continues| ok m -->|F1. Too many queued requests| ov g -->|F2. Write fails or object resets| er
Step notes:
- Step 3:
getByName(name)andidFromName(name)are deterministic: the same string in the same namespace always yields the same ID.newUniqueId()makes a random ID you must store somewhere (a cookie, KV, another object). The docs notenewUniqueIdhas lower first-use latency because a named ID needs a worldwide check the first time it is used, to make sure nobody on the other side of the planet created it at the same moment. - Step 4: creating an ID or a stub does not create the object. The object comes to life on the first method call.
- Step 6b: an object is created near the first caller unless you pass a
locationHinton that firstget(). The docs say objects do not currently move after creation (relocation is planned). - Step 9: with SQLite-backed objects, Cloudflare's write path forwards each commit to five follower machines and confirms it once three acknowledge (from the Cloudflare SQLite-in-DO engineering post). Only then does the output gate release your reply.
RPC, not hand-written HTTP
With a compatibility date of 2024-04-03 or later you call public methods directly: await env.COUNTER.getByName("views").increment(5). The stub is typed from the class. Calls made on the same stub arrive in the order you made them (E-order, implemented by Cap'n Proto RPC). There is no ordering between different stubs. If a stub throws, it may be broken: make a fresh one. WebSocket upgrades still go through the object's fetch() handler.
Each RPC method call on a stub is billed as one request. If a method returns an object that extends RpcTarget, calls on that returned stub ride the same session and are not billed separately.
Real code: the front Worker and a counter object
This is the real router used throughout this cluster. It type-checks with tsc --strict against @cloudflare/workers-types and runs under wrangler dev (local workerd).
// Front Worker: stateless router. Picks WHICH object handles a request, then calls it over RPC.
import { ChatRoom } from "./chat";
import { Counter, shardedIncrement, shardedRead } from "./counter";
import { RateLimiter } from "./ratelimit";
import { LeaseLock } from "./lease";
import { Inventory } from "./inventory";
export { ChatRoom, Counter, RateLimiter, LeaseLock, Inventory };
export interface Env {
CHAT_ROOM: DurableObjectNamespace<ChatRoom>;
COUNTER: DurableObjectNamespace<Counter>;
RATE_LIMITER: DurableObjectNamespace<RateLimiter>;
LEASE: DurableObjectNamespace<LeaseLock>;
INVENTORY: DurableObjectNamespace<Inventory>;
}
const json = (v: unknown, status = 200) =>
new Response(JSON.stringify(v), { status, headers: { "content-type": "application/json" } });
export default {
async fetch(request: Request, env: Env): Promise<Response> {
const url = new URL(request.url);
const [, kind, name = "default", action = ""] = url.pathname.split("/");
const q = url.searchParams;
switch (kind) {
case "chat": {
// Validate in the Worker so junk requests are not billed against the object.
if (request.headers.get("Upgrade") !== "websocket") return json({ error: "expected websocket" }, 426);
// Same name => same object, worldwide. locationHint only matters on the very first get.
const stub = env.CHAT_ROOM.getByName(`room:${name}`, { locationHint: "enam" });
return stub.fetch(request); // WebSocket upgrades still go through fetch()
}
case "counter": {
const stub = env.COUNTER.getByName(name);
return json({ value: action === "inc" ? await stub.increment(Number(q.get("by") ?? 1)) : await stub.get() });
}
case "sharded": {
const shards = Number(q.get("shards") ?? 8);
if (action === "inc") return json({ shardValue: await shardedIncrement(env, name, shards) });
return json(await shardedRead(env, name, shards));
}
case "rl": {
const d = await env.RATE_LIMITER.getByName(`key:${name}`).take(5, 1); // burst 5, 1 token/sec
return json(d, d.allowed ? 200 : 429);
}
case "lease": {
const stub = env.LEASE.getByName(`lock:${name}`);
if (action === "acquire") return json(await stub.acquire(q.get("owner") ?? "?", Number(q.get("ttl") ?? 30000)));
if (action === "release") return json({ released: await stub.release(q.get("owner") ?? "?", Number(q.get("token"))) });
return json({ error: "unknown action" }, 400);
}
case "inv": {
const stub = env.INVENTORY.getByName(`sku:${name}`);
if (action === "reset") { await stub.reset(Number(q.get("qty") ?? 3)); return json({ qty: await stub.qty() }); }
if (action === "racy") return json({ result: await stub.reserveRacy() });
if (action === "claim") return json({ result: await stub.reserveClaimFirst() });
if (action === "cas") return json({ result: await stub.reserveCas() });
return json({ qty: await stub.qty() });
}
}
return json({ error: "not found" }, 404);
},
} satisfies ExportedHandler<Env>;// Counter: one object per counter name. Sharded counters spread one hot name across N objects.
import { DurableObject } from "cloudflare:workers";
import type { Env } from "./index";
export class Counter extends DurableObject<Env> {
constructor(ctx: DurableObjectState, env: Env) {
super(ctx, env);
ctx.blockConcurrencyWhile(async () => {
ctx.storage.sql.exec("CREATE TABLE IF NOT EXISTS kv(k TEXT PRIMARY KEY, v INTEGER NOT NULL)");
});
}
// RPC method: callers do `await stub.increment(5)`; no Request/Response parsing.
increment(by = 1): number {
// UPSERT + RETURNING is one synchronous statement: atomic, no await, no interleaving.
return this.ctx.storage.sql
.exec("INSERT INTO kv(k, v) VALUES ('n', ?) ON CONFLICT(k) DO UPDATE SET v = v + excluded.v RETURNING v", by)
.one().v as number;
}
get(): number {
const rows = this.ctx.storage.sql.exec("SELECT v FROM kv WHERE k = 'n'").toArray();
return rows.length ? (rows[0].v as number) : 0;
}
}
// Worker-side helpers for a sharded counter: writes pick a random shard, reads fan in all shards.
export async function shardedIncrement(env: Env, name: string, shards: number): Promise<number> {
const shard = Math.floor(Math.random() * shards);
return env.COUNTER.getByName(`${name}#${shard}`).increment(1);
}
export async function shardedRead(env: Env, name: string, shards: number): Promise<{ total: number; perShard: number[] }> {
const perShard = await Promise.all(
Array.from({ length: shards }, (_, i) => env.COUNTER.getByName(`${name}#${i}`).get()),
);
return { total: perShard.reduce((a, b) => a + b, 0), perShard };
}Real run (local wrangler dev, wrangler 4.148.0, workerd 2026-10-06). Three sequential increments, then 200 concurrent increments against the same name, then 400 increments spread over 8 shards:
Output (captured from the real run):
inc -> { value: 1 }
inc -> { value: 2 }
inc -> { value: 3 }
after 200 concurrent incs -> { value: 203 }
sharded total -> 400 perShard -> [43,52,64,41,40,49,54,57]No update was lost under 200 concurrent requests even though there is no lock anywhere in the code: the object is single-threaded and the UPSERT runs synchronously.
Simulation: one name, one instance, worldwide
This Python simulation (not Cloudflare code) models the routing rules above: deterministic IDs per namespace, a directory recording where each ID lives, and creation near the first caller.
# SIMULATION (not Cloudflare code): how a name becomes exactly one live object, worldwide.
# Models: idFromName = deterministic hash, a global directory that records where an ID lives,
# and "create near the first caller". Latencies are illustrative example values, not measurements.
import hashlib
def id_from_name(namespace: str, name: str) -> str:
# Real IDs are 64 hex digits derived from the name within one namespace; we mimic the shape.
return hashlib.sha256(f"{namespace}/{name}".encode()).hexdigest()
directory = {} # id -> colo hosting the single live instance
instances = {} # id -> in-memory state of that instance
def call(caller_colo: str, namespace: str, name: str, op: str):
oid = id_from_name(namespace, name)
if oid not in directory:
# first use of a NAMED id: the system must make sure nobody else created it elsewhere
directory[oid] = caller_colo # created near the first caller
instances[oid] = {"messages": []}
route = f"created in {caller_colo} (first use, global uniqueness check)"
else:
route = f"routed {caller_colo} -> {directory[oid]}"
instances[oid]["messages"].append(op) # all ops land on the SAME instance
return oid[:12], route, len(instances[oid]["messages"])
for colo, op in [("ORD", "alice: hi"), ("FRA", "bob: hallo"), ("SIN", "chen: hello"), ("ORD", "alice: welcome")]:
print(colo, "->", *call(colo, "ChatRoom", "room:42", op))
print("other room ->", *call("FRA", "ChatRoom", "room:43", "dana: new room"))
print("same name, other class ->", id_from_name("Counter", "room:42")[:12], "(different namespace, different object)")
print("live instances:", len(instances), "| messages in room:42:", instances[id_from_name("ChatRoom", "room:42")]["messages"])Output (simulation):
ORD -> 8f48c4f635be created in ORD (first use, global uniqueness check) 1
FRA -> 8f48c4f635be routed FRA -> ORD 2
SIN -> 8f48c4f635be routed SIN -> ORD 3
ORD -> 8f48c4f635be routed ORD -> ORD 4
other room -> 0e05675e891b created in FRA (first use, global uniqueness check) 1
same name, other class -> 8873a7984c55 (different namespace, different object)
live instances: 2 | messages in room:42: ['alice: hi', 'bob: hallo', 'chen: hello', 'alice: welcome']Decision chart
Decisions
- 1
Start: what does the request need?
- next1. Must many clients see one ordered history of the same thing?
- ?
1. Must many clients see one ordered history of the same thing?
- no2. Mostly reads, staleness OK?
- yes4. Does the state split by user, room, tenant or doc?
- ?
2. Mostly reads, staleness OK?
- yesWorkers KV
- no3. Large files?
- 4
Workers KV
- ?
3. Large files?
- yesR2
- noD1 or Postgres
- 6
R2
- 7
D1 or Postgres
- ?
4. Does the state split by user, room, tenant or doc?
- yesOne Durable Object per entity
- no5. Over about 500 req/s on that one thing?
- 9
One Durable Object per entity
- next6. Need reports across all entities?
- ?
5. Over about 500 req/s on that one thing?
- noA single Durable Object is fine
- yesShard: many objects plus a fan-in or rollup
- 11
A single Durable Object is fine
- 12
Shard: many objects plus a fan-in or rollup
- ?
6. Need reports across all entities?
- yesObjects for the hot path, stream to D1 or Postgres
- 14
Objects for the hot path, stream to D1 or Postgres
The same logic as code, run over eight example workloads (decision helper, not Cloudflare code):
// SIMULATION / decision helper (not Cloudflare code): pick a store from the requirements.
type Need = {
coordination: boolean; // many clients must see one serialized order of updates
perEntityState: boolean; // state partitions cleanly by user / room / tenant / doc
readHeavyGlobal: boolean; // mostly reads, tolerate ~60 s staleness across locations
bigBlobs: boolean; // files / media / large objects
crossEntitySql: boolean; // ad hoc queries and joins ACROSS all entities
asyncWork: boolean; // fire-and-forget jobs with retries
realtime: boolean; // WebSockets / presence
};
function choose(n: Need): string {
if (n.bigBlobs) return "R2 (blob store), keep metadata elsewhere";
if (n.asyncWork && !n.coordination) return "Queues (at-least-once jobs)";
if (n.coordination || n.realtime) {
if (n.crossEntitySql) return "Durable Objects for the hot path + D1/Postgres for reporting";
return n.perEntityState ? "Durable Objects (one object per entity)" : "Durable Objects, but find a partition key first";
}
if (n.readHeavyGlobal) return "Workers KV (eventually consistent edge cache)";
if (n.crossEntitySql) return "D1 or Postgres (one shared SQL database)";
return "plain stateless Worker";
}
const base: Need = { coordination: false, perEntityState: false, readHeavyGlobal: false, bigBlobs: false, crossEntitySql: false, asyncWork: false, realtime: false };
const cases: [string, Partial<Need>][] = [
["chat room with presence", { coordination: true, realtime: true, perEntityState: true }],
["feature-flag config read on every request", { readHeavyGlobal: true }],
["seat booking per venue", { coordination: true, perEntityState: true }],
["user avatar uploads", { bigBlobs: true }],
["send welcome emails", { asyncWork: true }],
["finance dashboard across all tenants", { crossEntitySql: true }],
["per-tenant rate limit + monthly report", { coordination: true, perEntityState: true, crossEntitySql: true }],
["one global leaderboard, 50k writes/s", { coordination: true }],
];
for (const [name, n] of cases) console.log(name.padEnd(42), "->", choose({ ...base, ...n }));Output (simulation):
chat room with presence -> Durable Objects (one object per entity)
feature-flag config read on every request -> Workers KV (eventually consistent edge cache)
seat booking per venue -> Durable Objects (one object per entity)
user avatar uploads -> R2 (blob store), keep metadata elsewhere
send welcome emails -> Queues (at-least-once jobs)
finance dashboard across all tenants -> D1 or Postgres (one shared SQL database)
per-tenant rate limit + monthly report -> Durable Objects for the hot path + D1/Postgres for reporting
one global leaderboard, 50k writes/s -> Durable Objects, but find a partition key firstExpectedchat room with presence -> Durable Objects (one object per entity) feature-flag config read on every request -> Workers KV (eventually consistent edge cache) seat booking per venue -> Durable Objects (one object per entity) user avatar uploads -> R2 (blob store), keep metadata elsewhere send welcome emails -> Queues (at-least-once jobs) finance dashboard across all tenants -> D1 or Postgres (one shared SQL database) per-tenant rate limit + monthly report -> Durable Objects for the hot path + D1/Postgres for reporting one global leaderboard, 50k writes/s -> Durable Objects, but find a partition key first
Press Run. Snippets must be self-contained — no network, files, or native modules.
What happens if you choose otherwise
- Put a chat room's state in Workers KV: two people post at once, both read the same list, both write it back, and one message disappears. Other regions may not see the new list for a minute.
- Put every room in one Postgres table with row locks: correct, but every message is a network round trip plus a lock. Live connections still need sticky WebSocket servers plus Redis pub/sub between them.
- Use one global Durable Object for everything: correct and simple until about 500 to 1,000 requests per second, then you get "Durable Object is overloaded" errors. The docs list this as an anti-pattern.
- Run Orleans or Akka yourself: the same programming model, but you now own cluster membership, rebalancing, persistence and upgrades.
Pitfalls
- Global variables are shared. Several objects of the same class can share one isolate, so a module-level variable leaks between objects. Use instance fields.
- The object does not know who it is unless you tell it.
ctx.id.nameis set when the caller usedidFromNameorgetByName(names up to 1,024 bytes), but not viaidFromStringor fornewUniqueIdobjects. Pass identity in aninit()call or store it. - In-memory fields vanish on hibernation, eviction, deploys and crashes. Persist anything you need.
- Not every data center hosts objects. Hints for some regions (South America, Africa, Middle East) land in a nearby supported region.
- Always
awaitRPC calls. Unawaited calls swallow errors and lose return values.
Real-world anchors
- Cloudflare's storage guide says D1 and Queues are built on Durable Objects.
- The open-source workers-chat-demo is the canonical chat room per object design.
where.durableobjects.liveshows which locations currently host objects.- The Orleans "virtual actor" paper and docs describe the same idea: actors that always exist logically and are activated on demand.
The cluster map
- Durable Objects Storage - SQLite vs Legacy KV, Transactions, Write Coalescing & Point-in-Time Recovery: SQLite vs the legacy KV backend, transactions, write coalescing, point-in-time recovery and limits.
- Durable Objects Concurrency - Single Thread, Input & Output Gates, blockConcurrencyWhile & the Races That Remain: the single thread, input and output gates, blockConcurrencyWhile, and the race you can still write, with a real reproduction.
- Durable Objects Real-Time - WebSocket Hibernation, Alarms, Chat, Presence & Collaborative Editing: WebSocket Hibernation, alarms, chat, presence, collaborative editing, fan-out and backpressure.
- Durable Objects Scaling - Sharding by ID, Hot Objects, Location Hints, Limits & Cost: sharding by ID, hot objects, location hints, jurisdictions and cost, plus a rate limiter, a lease lock and a sharded counter.
- Durable Objects in Production - Resets, Deploys & Class Migrations, Observability, Testing & Interview Q&A: resets, deploys and class migrations, observability, testing with the Vitest plugin, and interview Q&A.
Interview Q&A
What is a Durable Object in one sentence?
Answer
A globally unique, single-threaded instance of a class, addressed by ID, with private transactional storage colocated with its code. It is an actor run as a managed service.
How do two Workers on different continents end up talking to the same object?
Answer
Both derive the same ID from the same name in the same namespace. The platform tracks where that ID lives and routes both calls there. The first use of a named ID includes a global check so only one instance is ever created.
Why not just use Redis for a chat room?
Answer
Redis gives you fast shared data, but you still need WebSocket servers, sticky routing, pub/sub between servers and someone to run Redis. A Durable Object per room holds the sockets, the ordering and the history in one place. Redis wins if you need one shared keyspace across many entities or you are not on Cloudflare.
When is a Durable Object the wrong tool?
Answer
Global read-heavy config (use KV), large files (R2), analytics across all entities (D1 or a warehouse), and any single entity that needs far more than about 1,000 requests per second without a way to shard it.
How does this compare with Orleans grains?
Answer
Same virtual actor idea: one activation per identity, single-threaded turns, activated on demand. Orleans runs on your cluster with pluggable storage. Durable Objects run on Cloudflare with built-in SQLite storage, global routing and hibernating WebSockets, but only JavaScript or WASM code.
What do you lose by choosing Durable Objects?
Answer
Cross-entity transactions and queries, control over placement beyond hints, in-memory state durability, and portability. You also have to design for restarts on every deploy.
getByName, idFromName or newUniqueId: which do you use?
Answer
Use getByName or idFromName when callers can derive the ID from a natural name. Use newUniqueId when you mint a fresh entity, store the ID somewhere, and want lower first-use latency.
Does creating a stub start the object?
Answer
No. Creating an ID or a stub is local. The object comes to life on the first method call.
What ordering do calls on one stub get?
Answer
Calls made on the same stub arrive in the order you made them (E-order). There is no ordering between different stubs.
Check yourself
Pick a system you know with shared mutable state. Name its atom of coordination, write the object name you would route on, and estimate the peak requests per second one object would see.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Scaling WebSockets — Sticky Sessions, Fan-out & Backpressure, Distributed Locks — Correctness, Leases & Fencing Tokens, Raft Consensus — Leader Election, Log Replication & Safety, Time, Clocks & Ordering in Distributed Systems - Physical Clocks, Lamport, Vector Clocks, HLC & TrueTime, SlotWise Reservation Service Incident - HLD Debug & Fix Path, Rate Limiter for an LLM Gateway - LLD Spec & Concepts, Browser Engines & PWAs — Service Workers, Rendering & Process Model.
Go Deeper
- Cloudflare Durable Objects documentation
- What are Durable Objects?
- Rules of Durable Objects
- Durable Object Namespace API (getByName, idFromName, newUniqueId)
- Durable Object Stub (E-order)
- Workers RPC
- Choose a data or storage product
- Introducing Workers Durable Objects (2020 launch post)
- Cloudflare workers-chat-demo
- Microsoft Orleans overview
- Actor model (Wikipedia)