Distributed systems
Part 6 of 6 · Durable ObjectsDurable Objects in Production - Resets, Deploys & Class Migrations, Observability, Testing & Interview Q&A
Production: failure catalog (deploys, eviction, write failure resets, blockConcurrencyWhile throws, overloaded vs retryable errors), version skew, declarative exports class lifecycle (create, rename three-deploy alias, transfer, delete), at-least-once alarms, observability, testing with @cloudflare/vitest-plugin (4 real passing tests), interview Q&A.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What happens to objects on deploy?
Answer
They restart, WebSockets disconnect, and old and new versions briefly coexist.
L2
Are there shutdown hooks?
Answer
No. Persist progress incrementally instead.
L3
Which errors should a caller retry?
Answer
Only retryable errors on idempotent calls, with backoff and a fresh stub.
L4
Why can an alarm double-charge?
Answer
Alarms are at least once and retried on throw, so a charge before a later failure repeats.
L5
How do you rename a class safely?
Answer
Deploy an alias, then the renamed tombstone with a live entry, then remove the alias.
L6
How do you test eviction recovery?
Answer
Write state, call evictDurableObject, then assert durable state survived.
L7
How do you debug one object among millions?
Answer
Filter metrics and logs by its ID or name and inspect its tables in Data Studio.
Failure modes
A rename orphans data
Changing the class name without a renamed exports entry and alias strands every object's storage.
An alarm charges a customer three times
Retries after a later failure repeat the side effect unless a dedupe key is stored first.
A renamed field breaks old objects during rollout
A new Worker sends the new shape to an object still on old code and gets 400 bad message.
Misconceptions
Deploys wait for objects to finish.
Objects restart, and storage access from the old instance throws.
Alarms run exactly once.
They are at least once and retried up to 6 times on throw.
Overloaded errors are transient and safe to retry.
Retrying deepens the overload. Shed load or shard.
Interviewer traps
Removing a field from an RPC contract in one deploy.
Change contracts additively and remove old fields after every caller and object moved.
Deleting a namespace to clean up.
A deleted state permanently removes the namespace and all its data.
Design scenario
Same prompt for every reader.
Requirements
No double charges, no lost charges during long payment outages, and a rename without data loss.
Failure assumptions
- Payment provider outages last minutes.
- Deploys happen several times a day.
- Old and new versions coexist during rollout.
Constraints
- No shutdown hooks.
- Alarm retries stop after 6 attempts.
Prompt
Ship a billing Durable Object that charges customers from an alarm and must survive deploys, retries and a class rename.
API
Which RPC fields can change, and how do old objects read new requests?
Data
Where is the dedupe key stored and when is it written?
Architecture
How do alarms reschedule during long outages, and what do you monitor?
Overview
Durable Objects are designed to restart. Deploys restart every object and drop every WebSocket. The runtime evicts idle objects, moves them between machines, and resets an object whose write failed, whose blockConcurrencyWhile threw or timed out, or which exceeded its limits. There are no shutdown hooks. So production readiness means: persist as you go, make every handler and alarm safe to run twice, keep RPC contracts compatible across versions that briefly coexist, change classes through the declarative exports config (create, rename, transfer, delete), watch the right metrics, and test against the real runtime with the Workers Vitest plugin.
How the code was checked: The real Durable Object TypeScript on this page type-checks with
tsc --strictagainst@cloudflare/workers-typesand was run locally underwrangler dev4.148.0 (workerd 2026-10-06) and the@cloudflare/vitest-plugintest pool. Blocks labeled simulation are sandbox models of the semantics, not Cloudflare code. Limits and prices are quoted from Cloudflare's docs as of 2026-10-07.
Why this matters
- Interviews: the senior signal is knowing how the system fails: "What happens to in-flight requests on deploy? What if the alarm runs twice? What if the object resets mid-transfer?"
- Production: most Durable Object incidents are self-inflicted: a rename that orphaned data, a non-idempotent alarm, a contract change that broke old objects during rollout, or a global singleton that started returning overloaded errors.
Failure catalog
| Event | What the runtime does | What you see | What to do |
|---|---|---|---|
| Deploy (code update) | Rolls out globally, eventually consistent; restarts objects | WebSockets disconnect; for seconds to minutes a new Worker may call an old object version | Compatible RPC contracts, client reconnect with backoff |
| Idle | Hibernates after about 10 s if eligible, otherwise evicts after 70 to 140 s | Constructor runs again on next event; memory is empty | Persist state, cheap constructor |
| Failed storage write | Resets the object, discards held outgoing messages | Callers get errors, nothing partial is visible | Retry only idempotent calls |
Throw or 30 s timeout inside blockConcurrencyWhile | Resets the object | Errors, then a fresh instance | Catch inside if you want to continue |
| Uncaught exception | May leave unknown state; runtime may terminate the instance | Error with .remote = true | try/catch boundaries |
| Overload | Queues, then rejects | .overloaded = true | Never retry; shard or shed |
| Transient infra error | Fails the call (network blip, internal error) | .retryable = true | Retry with exponential backoff if idempotent; recreate the stub |
| Runtime update or host move | Restart; in-flight requests get up to 30 s if they do not touch storage | Storage access from the old instance throws | Same as deploy |
ctx.abort() | Immediate reset; an interrupted alarm retries unless { retryAlarm: false } | Log entry with your message | Use for PITR restores and poison states |
The known-issues page adds a subtle one: global uniqueness is enforced when an event starts and when storage is accessed. An event that never touches storage may keep running on an instance that has already been replaced elsewhere.
How do you treat restarts?
Prefer
Design every event as possibly the first
Persist as you go, dedupe side effects, and keep contracts compatible across versions.
- The idempotent alarm charged once across three attempts.
- Additive contracts work in both rollout directions.
- Vitest proves state survives eviction.
Alternative
Assume the object stays up
In-memory progress, one-shot alarms and breaking contracts fail at the first deploy.
- The naive alarm charged three times.
- A renamed field returned 400 from old objects.
- A rename without exports strands data.
An object's life and its failure paths
Diagram 1 condensed. The failure catalog lists every reset and what to do.
- 1
Serve and persist
Write progress to storage as you go. - 2
Reset on deploy or failure
Deploys, failed writes and limit breaches restart the object. - 3
Resume from storage
The constructor runs again and memory starts empty. - 4
Retry only what is safe
Idempotent calls on retryable errors, never overloaded ones.
Lifecycle and failure paths (step-labeled)
States
- 1
Start → Inactive
Start → Inactive
- 2
Inactive → Active
Inactive → Active
1. first request or alarm runs constructor
- 3
Active → IdleHibernatable
Active → IdleHibernatable
2. no events, nothing pending
- 4
Active → IdleNonHibernatable
Active → IdleNonHibernatable
3. pending timers, I/O or standard WebSockets
- 5
IdleHibernatable → Hibernated
IdleHibernatable → Hibernated
4. about 10 s idle, sockets stay open
- 6
Hibernated → Active
Hibernated → Active
5. new event, constructor runs again
- 7
IdleNonHibernatable → Inactive
IdleNonHibernatable → Inactive
6. evicted after 70 to 140 s idle
- 8
Active → Inactive
Active → Inactive
F1. deploy, write failure, abort or exception reset
- 9
Hibernated → Inactive
Hibernated → Inactive
F2. runtime moves the object
Lesson map
Durable Objects in Production - Resets, Deploys & Class Migrations, Observability, Testing & Interview Q&A
Production: failure catalog (deploys, eviction, write failure resets, blockConcurrencyWhile throws, overloaded vs retryable errors), version skew, declarative exports class lifecycle (create, rename three-deploy alias, transfer, delete), at-least-once alarms, observability, testing with @cloudflare/vitest-plugin (4 real passing tests), interview Q&A.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB inactive["Inactive"] active["Active"] idlehibernatable["IdleHibernatable"] idlenonhibernatable["IdleNonHibernatable"] hibernated["Hibernated"] inactive -->|1. first request or alarm runs constructor| active active -->|2. no events, nothing pending| idlehibernatable active -->|3. pending timers, I/O or standard WebSockets| idlenonhibernatable idlehibernatable -->|4. about 10 s idle, sockets stay open| hibernated hibernated -->|5. new event, constructor runs again| active idlenonhibernatable -->|6. evicted after 70 to 140 s idle| inactive active -->|F1. deploy, write failure, abort or exception reset| inactive hibernated -->|F2. runtime moves the object| inactive
Deploys, version skew and class migrations
Version skew. During a rollout, a request can hit the new Worker which then calls an object still running the old code (and the reverse). With gradual deployments, the window lasts as long as both versions are live. Rule: change contracts additively, read tolerantly, remove old fields only after every caller and object has moved. Simulation (not Cloudflare code):
// SIMULATION (not Cloudflare code): during a rollout a NEW Worker can call an OLD object version
// (and vice versa) for seconds to minutes. RPC payloads must be forward and backward compatible.
type V1Msg = { user: string; text: string };
type V2Msg = { user: string; text: string; replyTo?: number }; // additive, optional => safe
type V2Bad = { author: string; body: string }; // renamed fields => breaks old code
function oldObjectHandle(msg: any): string { // version 1 still running somewhere
if (typeof msg.user !== "string" || typeof msg.text !== "string") throw new Error("400 bad message");
return `stored "${msg.text}" from ${msg.user}`;
}
function newObjectHandle(msg: Partial<V2Msg> & Partial<V2Bad>): string { // tolerant reader in version 2
const user = msg.user ?? msg.author, text = msg.text ?? msg.body;
if (!user || !text) throw new Error("400 bad message");
return `stored "${text}" from ${user}${msg.replyTo ? ` (reply to #${msg.replyTo})` : ""}`;
}
const cases: [string, unknown, (m: any) => string][] = [
["old Worker -> new object", { user: "a", text: "hi" } as V1Msg, newObjectHandle],
["new Worker (additive) -> old object", { user: "b", text: "yo", replyTo: 7 } as V2Msg, oldObjectHandle],
["new Worker (renamed) -> old object", { author: "c", body: "hey" } as V2Bad, oldObjectHandle],
["new Worker (renamed) -> new object", { author: "c", body: "hey" } as V2Bad, newObjectHandle],
];
for (const [name, msg, h] of cases) {
try { console.log(name.padEnd(38), "OK ", h(msg)); } catch (e) { console.log(name.padEnd(38), "FAIL", (e as Error).message); }
}
export {};Output (simulation):
old Worker -> new object OK stored "hi" from a
new Worker (additive) -> old object OK stored "yo" from b
new Worker (renamed) -> old object FAIL 400 bad message
new Worker (renamed) -> new object OK stored "hey" from cExpectedold Worker -> new object OK stored "hi" from a new Worker (additive) -> old object OK stored "yo" from b new Worker (renamed) -> old object FAIL 400 bad message new Worker (renamed) -> new object OK stored "hey" from c
Press Run. Snippets must be self-contained — no network, files, or native modules.
Class lifecycle with exports. Since June 2026 the declarative exports field in the Wrangler config replaces the imperative migrations array (both still work, but one Worker uses only one). Our lab's config:
{
// Real Worker + Durable Objects used by the Study lesson. Runs locally with `wrangler dev`.
"name": "do-study-lab",
"main": "src/index.ts",
"compatibility_date": "2026-10-01",
"durable_objects": {
"bindings": [
{ "name": "CHAT_ROOM", "class_name": "ChatRoom" },
{ "name": "COUNTER", "class_name": "Counter" },
{ "name": "RATE_LIMITER", "class_name": "RateLimiter" },
{ "name": "LEASE", "class_name": "LeaseLock" },
{ "name": "INVENTORY", "class_name": "Inventory" }
]
},
// Declarative class lifecycle (replaces the legacy "migrations" array).
"exports": {
"ChatRoom": { "type": "durable-object", "storage": "sqlite" },
"Counter": { "type": "durable-object", "storage": "sqlite" },
"RateLimiter": { "type": "durable-object", "storage": "sqlite" },
"LeaseLock": { "type": "durable-object", "storage": "sqlite" },
"Inventory": { "type": "durable-object", "storage": "sqlite" }
},
"observability": { "enabled": true }
}| Operation | exports entry | Danger |
|---|---|---|
| Create | { "type": "durable-object", "storage": "sqlite" } | New namespaces are always SQLite |
| Delete | "state": "deleted" | Permanently deletes the namespace and all its data; rejected if code still exports the class or another Worker binds to it |
| Rename | Old name "state": "renamed", "renamed_to": "New" plus a live entry for New | Not atomic with the code rollout: use the three-deploy pattern (alias export { New as Old }, apply the rename, remove the alias) |
| Transfer to another Worker | Target: "expecting-transfer" with transfer_from; source: "transferred" with transferred_to | A four-deploy sequence; bind on the target only after the source deploy lands |
Lifecycle changes apply only through wrangler deploy. If the config has exports entries, wrangler versions upload fails fast. Data-schema changes are a different thing: run them inside each object in the constructor, as shown on the storage page.
Alarms that run twice
Alarms are at least once: a throw is retried with backoff from 2 s, up to 6 times, and in rare cases an alarm can fire more than once even without a throw. Simulation of why side effects need a dedupe key, and what happens in a long outage (not Cloudflare code):
# SIMULATION (not Cloudflare code) of alarm delivery: at-least-once, retried on uncaught exceptions
# with exponential backoff starting at 2 s, up to 6 retries (numbers from the Alarms docs).
def deliver(handler, fail_first_n, max_retries=6):
attempts, delay, t = 0, 2.0, 0.0
while True:
attempts += 1
ok = handler(attempt=attempts, will_fail=attempts <= fail_first_n)
if ok: return f"succeeded on attempt {attempts} at t={t:.0f}s"
if attempts > max_retries: return f"gave up after {attempts} attempts (t={t:.0f}s); alarm is gone until someone calls setAlarm again"
t += delay; delay *= 2
charges, seen = [], set()
def naive(attempt, will_fail):
charges.append(f"charge#{attempt}") # side effect happens BEFORE the failure below
return not will_fail
def idempotent(attempt, will_fail):
if "invoice-42" not in seen: # dedupe key stored in the object's SQLite in real code
seen.add("invoice-42"); charges.append("charge")
return not will_fail
print("naive :", deliver(naive, fail_first_n=2), "| charges:", charges); charges.clear()
print("idempotent:", deliver(idempotent, fail_first_n=2), "| charges:", charges); charges.clear()
print("long outage:", deliver(idempotent, fail_first_n=99))
print("fix for long outages: catch inside alarm(), record the error, call setAlarm(now + backoff) yourself")Output (simulation):
naive : succeeded on attempt 3 at t=6s | charges: ['charge#1', 'charge#2', 'charge#3']
idempotent: succeeded on attempt 3 at t=6s | charges: ['charge']
long outage: gave up after 7 attempts (t=126s); alarm is gone until someone calls setAlarm again
fix for long outages: catch inside alarm(), record the error, call setAlarm(now + backoff) yourselfThe docs' advice for long outages: catch inside alarm() and schedule your own next attempt, so you are not limited to 6 automatic retries.
Calling objects safely from a Worker
Errors from a stub carry flags. Recreate the stub after any error (a broken stub keeps failing), retry only .retryable errors on idempotent calls with exponential backoff, and never retry .overloaded ones. Real, type-checked helper:
// Worker-side call wrapper: recreate the stub after errors, retry only retryable + idempotent calls.
type DOError = Error & { retryable?: boolean; overloaded?: boolean; remote?: boolean };
export async function callWithRetry<T>(makeStub: () => T, op: (s: T) => Promise<unknown>, idempotent: boolean, max = 4) {
for (let attempt = 0; ; attempt++) {
const stub = makeStub(); // a stub that threw may be "broken": always make a fresh one
try {
return await op(stub);
} catch (err) {
const e = err as DOError;
if (e.overloaded) throw e; // never retry overload: it makes the hot object hotter
if (!e.retryable || !idempotent || attempt >= max) throw e;
const backoff = Math.min(2000, 100 * 2 ** attempt) * (0.5 + Math.random()); // exponential + jitter
await new Promise((r) => setTimeout(r, backoff));
}
}
}Observability
- Metrics: namespace and per-object charts in the dashboard (filter by ID or name), backed by GraphQL datasets
durableObjectsInvocationsAdaptiveGroups,durableObjectsPeriodicGroups,durableObjectsStorageGroupsanddurableObjectsSubrequestsAdaptiveGroups. - Memory: the Memory usage chart is per isolate (P50 to P999), and an isolate can host several objects. A rising trend suggests a leak.
- Logs: enable
"observability": { "enabled": true }. Request logs carry$workers.durableObjectId.wrangler tailstreams live logs, but WebSocket request logs only appear when the socket closes, and tail should not be attached to heavy traffic. - Data Studio: view and edit a SQLite-backed object's tables from the dashboard.
- Log what you will need: object ID or name, method, rows read and written (
cursor.rowsReadandrowsWritten), andretryableoroverloadedflags on errors.
Testing with the Workers Vitest plugin
The docs now use @cloudflare/vitest-plugin (cloudflareTest() in vitest.config.ts, Vitest 4.1 or later), which runs tests inside workerd with real bindings. Helpers from cloudflare:test include runInDurableObject (reach into an instance and its storage), runDurableObjectAlarm (fire a pending alarm now), evictDurableObject and evictAllDurableObjects (tear down in-memory state while keeping storage), and listDurableObjectIds. Older projects used the @cloudflare/vitest-pool-workers package. Miniflare is the local simulator underneath wrangler dev and these tests.
import { cloudflareTest } from "@cloudflare/vitest-plugin";
import { defineConfig } from "vitest/config";
export default defineConfig({
plugins: [cloudflareTest({ wrangler: { configPath: "./wrangler.jsonc" } })],
});// Runs inside workerd via @cloudflare/vitest-plugin: real Durable Objects, real SQLite, real alarms.
import { env } from "cloudflare:workers";
import { evictDurableObject, runDurableObjectAlarm, runInDurableObject } from "cloudflare:test";
import { describe, expect, it } from "vitest";
import type { LeaseLock } from "../src/lease";
describe("Durable Objects lab", () => {
it("same name -> same object; 100 concurrent RPCs lose no updates", async () => {
const stub = env.COUNTER.getByName("t-counter");
await Promise.all(Array.from({ length: 100 }, () => stub.increment(1)));
expect(await env.COUNTER.getByName("t-counter").get()).toBe(100);
});
it("eviction wipes memory but the persisted bucket survives (no free refill)", async () => {
const stub = env.RATE_LIMITER.getByName("key:evict");
for (let i = 0; i < 5; i++) await stub.take(5, 0.001); // drain the burst; refill is ~0
await evictDurableObject(stub); // constructor runs again on next call
const d = await stub.take(5, 0.001);
expect(d.allowed).toBe(false);
});
it("racy reserve oversells, claim-first does not", async () => {
const racy = env.INVENTORY.getByName("sku:t-racy");
await racy.reset(3);
const r1 = await Promise.all(Array.from({ length: 6 }, () => racy.reserveRacy()));
expect(r1.filter((x) => x === "reserved").length).toBeGreaterThan(3); // the bug, reproduced
const safe = env.INVENTORY.getByName("sku:t-claim");
await safe.reset(3);
const r2 = await Promise.all(Array.from({ length: 6 }, () => safe.reserveClaimFirst()));
expect(r2.filter((x) => x === "reserved").length).toBe(3);
expect(await safe.qty()).toBe(0);
});
it("lease alarm clears an expired holder; fencing token keeps increasing", async () => {
const stub = env.LEASE.getByName("lock:t");
const a = await stub.acquire("a", 60_000);
// Simulate the TTL passing: rewrite expires_at instead of sleeping for a minute.
await runInDurableObject(stub, (_inst: LeaseLock, state) => {
state.storage.sql.exec("UPDATE lease SET expires_at = ? WHERE id = 1", Date.now() - 1);
});
expect(await runDurableObjectAlarm(stub)).toBe(true); // fire the pending alarm now
await runInDurableObject(stub, (_inst: LeaseLock, state) => {
const row = state.storage.sql.exec("SELECT owner, token FROM lease").one();
expect(row.owner).toBeNull();
expect(row.token).toBe(a.token);
});
const b = await stub.acquire("b", 1000);
expect(b.token).toBe((a.token ?? 0) + 1);
});
});Real run (Vitest 4.1.11 with @cloudflare/vitest-plugin, all inside workerd):
Output (captured from the real run):
✓ test/do.test.ts > Durable Objects lab > same name -> same object; 100 concurrent RPCs lose no updates 86ms
✓ test/do.test.ts > Durable Objects lab > eviction wipes memory but the persisted bucket survives (no free refill) 22ms
✓ test/do.test.ts > Durable Objects lab > racy reserve oversells, claim-first does not 155ms
✓ test/do.test.ts > Durable Objects lab > lease alarm clears an expired holder; fencing token keeps increasing 21ms
Test Files 1 passed (1)
Tests 4 passed (4)What to test, in order of value: concurrency invariants (fire N calls with Promise.all), eviction recovery (state survives evictDurableObject), alarms (fire with runDurableObjectAlarm, run twice, assert idempotence), and contract compatibility between versions.
Operational checklist
Decisions
- 1
1. New change ready
- next2. Does it change an RPC contract?
- ?
2. Does it change an RPC contract?
- yesMake it additive, tolerant reader, deploy objects and Workers in a safe order
- no3. Does it change a class name, delete or move a class?
- 3
Make it additive, tolerant reader, deploy objects and Workers in a safe order
- next3. Does it change a class name, delete or move a class?
- ?
3. Does it change a class name, delete or move a class?
- yesUse exports tombstones, follow the multi-deploy pattern, back up data first
- no4. Does it change SQLite schema?
- 5
Use exports tombstones, follow the multi-deploy pattern, back up data first
- next4. Does it change SQLite schema?
- ?
4. Does it change SQLite schema?
- yesMigration in constructor via blockConcurrencyWhile, expand then contract
- no5. Run Vitest plugin suite: concurrency, eviction, alarms
- 7
Migration in constructor via blockConcurrencyWhile, expand then contract
- next5. Run Vitest plugin suite: concurrency, eviction, alarms
- 8
5. Run Vitest plugin suite: concurrency, eviction, alarms
- next6. Deploy, expect WebSocket reconnects, watch overload and error metrics
- 9
6. Deploy, expect WebSocket reconnects, watch overload and error metrics
- F1. errors spikeRoll back code; for data damage use per-object PITR
- 10
Roll back code; for data damage use per-object PITR
Interview Q&A
What happens to a Durable Object when you deploy?
Answer
The new code rolls out globally and objects restart. WebSockets are disconnected. In-flight requests that do not touch storage may finish, but storage access from the old instance throws. For a short window, new Workers may call objects still on the old version, so contracts must be compatible both ways.
How do you rename a Durable Object class without losing data?
Answer
Use an exports renamed tombstone pointing at a live entry for the new name. To avoid errors during rollout, deploy an alias first (export { NewName as OldName }), then the rename, then remove the alias.
Your alarm charges customers and sometimes double-charges. Why?
Answer
Alarms are at least once and retried on throw. If the charge happens before a later failure, the retry charges again. Store a dedupe key per charge in the object's SQLite and check it first, or use the payment provider's idempotency key.
How do you test a Durable Object's recovery after eviction?
Answer
With the Vitest plugin: write state, call evictDurableObject(stub), then assert durable state survived and in-memory caches reset. We did exactly this for the rate limiter.
The object returns "Durable Object is overloaded". Should the caller retry?
Answer
No. The error has .overloaded = true, and retrying deepens the overload. Shed load and shard the hot entity.
Why are there no shutdown hooks?
Answer
Cloudflare cannot guarantee they would run in every failure case, and code would come to depend on them. Write progress to storage incrementally instead. Storage writes are fast, and the output gate keeps them honest.
How do you debug one misbehaving object among millions?
Answer
Filter dashboard metrics and logs by its ID or name ($workers.durableObjectId), inspect its tables in Data Studio, and reproduce locally by addressing the same name in a Vitest test.
What should an alarm do during a long outage?
Answer
Catch errors inside alarm(), record them, and call setAlarm with your own backoff, so you are not limited to 6 automatic retries.
Which tests give the most value for a Durable Object?
Answer
Concurrency invariants with Promise.all, eviction recovery, alarms fired twice for idempotence, and contract compatibility between versions.
Check yourself
Pick one side effect your service performs from a timer or retry loop. Write the dedupe key you would store before it, and the test that proves it runs once when fired twice.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Testing Strategies — Pyramid, Contracts, Property-Based & Flakes, Load, Chaos & Production Validation — Shadow Traffic, Canaries & Game Days, Observability Triad — Metrics, Logs & Distributed Tracing, Zero-Downtime Database Migrations — Expand/Contract, Dual Write & Online DDL, API Idempotency Keys, Failover & Split Brain — Detection, Promotion, Fencing & Lost Writes, Disaster Recovery & Multi-Region - RTO/RPO, Backups, Pilot Light to Active-Active.
Go Deeper
- Durable Object class exports (create, rename, transfer, delete)
- Known issues: global uniqueness and code updates
- Error handling: retryable and overloaded
- Troubleshooting
- Metrics and analytics
- Testing Durable Objects with the Vitest plugin
- Workers Vitest integration
- Gradual deployments (Durable Objects section)
- @cloudflare/vitest-plugin on npm
- Cloudflare Actors framework