Dual Write / Dual Read — Migrating Data Paths Safely
When rows must live in a new table or store, dual-write keeps both paths updated and dual-read flips traffic. The order is write-both/read-old, shadow compare, read-new, then write-new-only. Partial failure, ordering, and idempotency dominate. An outbox makes the second write durable intent.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
The new store is not the row's home yet
Prefer
Dual-write, then shadow, then flip
The app writes both paths. Users still read the old source of truth until a flag says otherwise. Rollback is that flag.
- You see production write load on the new path before users depend on it.
- No stream platform required. The cost is double writes and a real error policy.
Alternative
CDC cutover
One write path. A stream applies changes to the new store. Less application branching, more pipeline lag and ops.
- Prefer it when the org already runs log capture well.
- Still needs checksums. A silent apply bug is not fixed by removing the second RPC.
Overview
When data must live in a new table, a new shard, or a new store, dual-write keeps both paths updated and dual-read lets you flip traffic gradually. A column rename on the same table does not need this machinery. Expand/contract is enough there. Dual-write is for a path that moves.
The progression:
- Write-both / read-old. The new path stays warm. Users still see the old source of truth.
- Shadow read. Compare new against old on a sample, or asynchronously. Do not serve the new value yet.
- Read-new. A feature flag. Rollback is a flag flip, not a restore.
- Write-new-only. Only after backfill and reconcile say the old path is idle and caught up.
Ordering, partial failure, and idempotency dominate the bugs. You rarely get one transaction across two stores. Design for a chosen failure policy plus reconcile.
Flow
- 1
1. Baseline: write old, read old
- next2. Expand: write both, read old
- 2
2. Expand: write both, read old
- next3. Shadow compare, still serve old
- 3
3. Shadow compare, still serve old
- next4. Flag: read new, keep dual-write
- 4
4. Flag: read new, keep dual-write
- next5. Stop old writes after reconcile
- 5
5. Stop old writes after reconcile
- next6. Contract: drop the old path
- 6
6. Contract: drop the old path
Lesson map
Dual Write / Dual Read — Migrating Data Paths Safely
When rows must live in a new table or store, dual-write keeps both paths updated and dual-read flips traffic. The order is write-both/read-old, shadow compare, read-new, then write-new-only. Partial failure, ordering, and idempotency dominate. An outbox makes the second write durable intent.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Baseline: write old, read old"] b["2. Expand: write both, read old"] c["3. Shadow compare, still serve old"] d["4. Flag: read new, keep dual-write"] a -->|1. Baseline: write old, read old| b b -->|2. Expand: write both, read old| c c -->|3. Shadow compare, still serve old| d
Phases, and the two ways back
Rollback from read-new returns to read-old while dual-write is still on. A shadow mismatch aborts the flip. It does not serve the new row.
- 1
Write old, read old
Baseline. No new path in the request. - 2
Write both, read old
New mutations hit both stores. Historical rows are still a backfill problem. - 3
Shadow the new read
Compare. Increment a mismatch metric. The response still comes from the old store. - 4
Flag the read to new
Users depend on the new path. Keep dual-write so a flag flip back is not a stale read. - 5
Write new only, then contract
Stop old writes only after reconcile. Dropping the old path is the irreversible zone on the cutover page.
A mismatch during shadow sends you back to repair, still serving old. A bad canary sends the read flag back to old and keeps dual-write on, so the old store does not miss updates while you investigate. That flag story is rollback and cutover.
Consistency hazards
- Partial write. Old succeeds and new fails, or the reverse. Without a repair path the stores diverge.
- Ordering. Two concurrent updates to the same key can apply in different orders on each store. A version column or last-write token has to be part of the design, not a follow-up.
- Idempotency. Retries must not double-apply. Use an idempotency key or a version. The HTTP key machine itself is the idempotency keys lesson. Here the key only has to make the second store's upsert a no-op.
- Read-your-writes. The user writes, then reads a replica or the other store, and sees a miss or a stale value. During read-old, the read must hit the store you just wrote as source of truth.
- No cross-store transaction. Two commits are not atomic. Pick a policy and a reconcile job.
Fail closed: if either write fails, fail the request. Stronger, higher error rate.
Succeed and enqueue: commit the old store (or an outbox row) and repair the new store asynchronously. Softer, and you now own a lag SLO.
Why dual-write beats a naive cutover: you validate the new path under production writes before users depend on it. Why a CDC cutover can beat dual-write: one write path, a reusable stream, less application branching — paid for with pipeline ops and lag. Prefer CDC when the organization already runs it well. Prefer dual-write when the move is small or CDC is not available.
Logical replication is the database-native cousin of that stream for Postgres. It is an alternative path when you are standing up a second cluster, not a second lesson in slots and decoding.
Fail-closed dual-write
The sketch is in memory so you can run it. The policy is the point: both upserts succeed, or the call reports failure. A retry with the same idempotency key is a no-op success. The demo does not roll back the old store when the new store fails — production must either do that or enqueue a repair. The metric is the part you are not allowed to skip.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Shadow read
Shadow mode compares and still returns the old balance. Read-new returns the new row, with a fallback to old during the window where you might roll back and a key exists on only one side. The fallback is not permission to skip reconcile.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
What is the order of dual-write versus backfill?
Answer
Expand the schema, start dual-write for new mutations, backfill historical rows, shadow-read, then flip reads. A backfill alone races with live writes: a row updated after you copied it stays stale on the target. Dual-write covers the moving head. The backfill covers history. The next page is that job.
One of the two writes fails. What do you do?
Answer
Whatever you agreed before coding. Fail the request, or enqueue a repair. Never ignore the miss. Chart dual_write_partial. A succeed-on-old policy without a worker is a silent divergence.
How does an outbox help dual-write?
Answer
The same database transaction as the primary write records intent. An async worker applies that intent to the second store with retries and an idempotency key. You do not lose the second write because the process crashed after the first commit. Claim loops, inbox dedupe, and log connectors are the outbox and CDC pages.
What is the difference between a dual-read and a shadow read?
Answer
A shadow read compares and does not serve the new value. A dual-read that serves the new value is the cutover. Shadow first. Serving new on the first day you notice the stores disagree is how mismatches become customer-visible.
When is dual-write the wrong tool?
Answer
A pure column add on the same table: use expand/contract. A cross-region multi-master move with no conflict strategy. A target that CDC already syncs reliably. Also the wrong tool at very high QPS if you cannot batch or sample the second write without a lag budget you actually meet.
Why can the two stores apply updates in different orders?
Answer
There is no shared lock across stores. Request A and request B both dual-write key K. The old store commits A then B. The new store commits B then A. Last-write-wins without a version is not a strategy. Put a version or a per-key sequence on the payload and make the upsert conditional.
Why is read-your-writes easy to break during the migration?
Answer
The write hit the old primary. The following read hit a lagged replica, or the new store that has not applied yet. While the source of truth is still old, the read path for that user has to be old. Shadow traffic must not become the response.
What do you measure?
Answer
Partial dual-write failures, shadow mismatch rate, and the lag of the repair queue. A dashboard that only shows request success will stay green while the new store drifts.
Pitfalls
- Starting the backfill and forgetting to turn on dual-write for live rows.
- Serving the new store during shadow because the code path was "while we're here."
- A retry that inserts a second row on the new store because the upsert key was the request id of the retry, not a stable idempotency key.
- Stopping writes to the old store before reads have soaked, so rollback shows stale data.
- Re-implementing a Debezium connector inside the dual-write page. Point at the CDC cluster and move on.
The new store times out on 2 percent of writes. Product will not accept a matching error rate, and finance will not accept a silent miss. Which policy do you choose, what do you enqueue, and which metric pages you before the read flag moves?
Go deeper
Martin Fowler's parallel change is the phase model. microservices.io's transactional outbox page is the pattern name for "intent in the same transaction." Postgres logical replication is the alternative when the move is a second cluster fed from the log rather than a second RPC. Connector configuration stays on the CDC pages.
Dual-write is a distributed-systems problem wearing a database costume. Design the failure policy first. Next: backfills and reconciliation.