Rollback, Feature Flags & Cutover — Instant Reverse Without Data Loss
Cutover is when users start depending on the new path. Rollback flips the read flag in seconds, which only works while the old path still exists and dual-write kept it current. DROP COLUMN, destructive casts, and write-new-only after the old path has drained are the irreversible zone.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
The read flag is wrong in production
Prefer
Flip the flag, keep dual-write
Seconds. The binary that can read both shapes is already deployed. The old path is still receiving writes, so it is not stale.
- No restore, no binary rollback race with in-flight requests.
- Canary percent goes to zero. Shadow jobs stay on while you repair.
Alternative
Restore from backup
The last resort when you actually lost data. It is not how you undo a read switch.
- You roll the primary back in time and drop every write since the snapshot.
- Right only after a destructive contract or a transform you cannot invert.
Overview
Cutover is when users start depending on the new path. Rollback is how you undo that decision in seconds. Expand/contract makes rollback a code or feature-flag flip, not a schema restore. You have to know which steps are still reversible.
Reversible, and where you should stay until the soak is real:
- Which store or column a read uses
- Dual-write still enabled, so both paths receive updates
- The canary percentage on
read_new - Shadow compare jobs
Irreversible, or expensive to reverse:
DROP COLUMNorDROP TABLE- A destructive type coercion that loses precision
- Write-new-only after the old path has stopped, so the old path missed updates
- One-way encryption or hashing
- Hard deletes of the source rows after the copy
As long as the old shape exists and dual-write (or a CDC stream) kept it fresh, rolling reads back is instant. Schema restore is the wrong rollback tool for an online migration.
A flag beats a redeploy. Seconds, not minutes, and you do not race a binary rollback against requests that are already in flight. The binaries on the fleet must already be dual-compatible. The flag only chooses which path they serve. That compatibility is expand/contract. The proof you are allowed to flip the flag is checksums and shadow reads.
Flow
- 1
1. Dual-write and backfill are done
- next2. Dark launch, then shadow traffic
- 2
2. Dark launch, then shadow traffic
- next3. Canary read_new up to 100 percent
- 3
3. Canary read_new up to 100 percent
- next4. Soak errors and correctness
- 4
4. Soak errors and correctness
- next5. Stop writes on the old path
- 5
5. Stop writes on the old path
- next6. Contract later in a separate release
- 6
6. Contract later in a separate release
Lesson map
Rollback, Feature Flags & Cutover — Instant Reverse Without Data Loss
Cutover is when users start depending on the new path. Rollback flips the read flag in seconds, which only works while the old path still exists and dual-write kept it current. DROP COLUMN, destructive casts, and write-new-only after the old path has drained are the irreversible zone.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Dual-write and backfill are done"] b["2. Dark launch, then shadow traffic"] c["3. Canary read_new up to 100 percent"] d["4. Soak errors and correctness"] a -->|1. Dual-write and backfill are done| b b -->|2. Dark launch, then shadow traffic| c c -->|3. Canary read_new up to 100 percent| d
Forward, and the way back
A failed canary sets read_new to false and leaves dual-write on. Repair, then shadow again. Contract is not in this incident.
- 1
Dark launch and shadow
The new path executes. Results are discarded or compared. Users still see the old path. - 2
Mismatch or error SLO fails
Stay on old reads. Repair the data. Do not widen the canary to debug it. - 3
Canary the read flag
A percent, then 100. Watch the new path's latency and errors for that cohort. - 4
Rollback
read_new goes false. Dual-write stays on so the old path did not miss updates. - 5
Soak, then stop old writes
Only after the error budget and correctness hold. This is the step that makes rollback expensive. - 6
Contract later
Drop the old shape in a later release. Document retention. This is the irreversible zone.
Cutover checklist
- Schema is expanded. Dual-compatible code is on 100 percent of instances, including workers.
- Dual-write is healthy. The partial-failure rate is under its SLO, and repairs drain.
- The backfill resume token is at the end of the keyspace. Checksums are green.
- Shadow mismatch rate has been under the SLO for the soak you agreed, not for a five-minute glance.
- Dashboards and alerts exist for the new path's latency and errors before you send it traffic.
- On-call knows the window. A named person owns the rollback flag.
- The flag default is read-old. Enable a canary, then 100 percent.
- Only then disable writes to the old path. Contract is a later release.
- Write down which steps are irreversible, and how long you retain the old data for audit.
Dark launch, shadow, canary
- Dark launch: the new path executes. Results are discarded or only compared. Users still see the old path. You learn about errors and latency without serving the answer.
- Shadow traffic: mirror a percent of reads or writes to the new system for load and correctness. The response source of truth is still old.
- Canary: a percent of users are served the new path. This is cutover for that cohort. It is not a dark launch.
Do not call it cutover until the flag serves the new path as the response. Dark first, canary second. A canary that is actually 100 percent was a big-bang with extra vocabulary.
The brief lock at a table-swap cutover — gh-ost's postponeable swap, pt-osc's RENAME — is an online DDL fact. The flag does not remove that metadata lock. It removes the need to restore the data directory when the new path misbehaves after the swap.
Flag-gated reads
Rollback sets read_new to false. Nothing in the schema changes. Stopping old writes is a separate predicate: the read flag is on, the soak passed, and new writes are on. A canary uses a bucket, not a boolean, so 10 percent is a number you can set back to zero.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Bucket 5 is inside a 10 percent canary and sees the new value. Bucket 50 does not. After rollback the percent is 0, so bucket 5 sees the old value again. writeOld stays true in that function. Turning it off is can_stop_old_writes, and only after the soak.
Who is allowed to contract
Metrics that show zero reads and writes of the old shape, a time soak, and an explicit owner. Not the same pull request as the cutover. Not the same hour as the canary. Some data cannot be un-migrated — a lossy cast, a hash, a hard delete. Plan to retain both copies until contract, and write down the retention for audit before you drop.
Interview Q&A
How do you roll back a migration without restoring a backup?
Answer
Keep the old schema and the old path until the soak ends. Flip the read flag back. Keep dual-write so the old path stays current. A restore is the last resort for true data loss, because it throws away every write since the snapshot.
When is rollback impossible?
Answer
After a destructive contract (DROP), after write-new-only has run long enough that the old path diverged, or after an irreversible transform (precision loss, one-way hashing, hard delete of the source). Call these out in the design review, before the flag exists.
Why canary the read flag?
Answer
It limits blast radius. You watch errors and mismatches on a cohort before 100 percent. The percent is a number you can set to zero without a deploy.
What is the difference between a dark launch and a canary?
Answer
A dark launch does not serve the new path to users. A canary does, for a percent. Dark first, canary second. If you skip dark, the first users are your load test.
Who decides to contract?
Answer
The owner, from metrics that show zero old reads and writes, plus a time soak. Not the same pull request as the cutover. Contract is how you enter the irreversible zone on purpose.
Why does a flag beat redeploying the previous binary?
Answer
A flag changes behavior in seconds, in a process that already understands both shapes. A binary rollback takes a rollout, and requests in flight still hit a mix of versions. If the previous binary cannot read the expanded schema, you do not have a rollback. You have expand/contract done backwards.
What do you keep writing during a read rollback?
Answer
Both paths. If you already stopped writes to the old path, flipping reads back serves a stale store. The reversible region is "read flag off, dual-write still on." Leave that region only after the soak.
What belongs in the cutover checklist that people skip?
Answer
Workers, not just the web fleet, on the dual-compatible build. A named rollback owner. Alerts that exist before traffic moves. A shadow window long enough to mean something. A written list of irreversible steps. The resume token at the end of the keyspace, not "the job looked finished."
Pitfalls
- Calling a shadow test a cutover.
- Setting the canary to 100 percent because the graph looked quiet for a minute.
- Stopping old writes in the same change as the read flip.
- Dropping the column in the cutover pull request.
- Rolling back by restoring a backup while dual-write could have kept the old path fresh.
- A flag default of read-new, so a missing config serves the new path.
- Assuming the table-swap lock is the rollback story. It is a short stall. The data rollback is the flag.
You hashed emails in the new column with a one-way function, dual-wrote for a week, and then dropped the plaintext column. A product bug needs the old email format. Is this a flag flip, a re-backfill, or a restore? What should the design review have forbidden before the drop?
Go deeper
Martin Fowler's feature-toggles article is the flag taxonomy: a migration toggle is temporary, and it has an expiry. Parallel change is why the old shape still exists when you flip. Postgres ALTER TABLE and the MySQL online DDL docs are where you confirm that the final drop or rename still takes a brief lock even when the data movement was online. That lock is not your rollback plan.
If rollback needs a backup, the phases were too coarse. Prefer flags and dual paths until contract. The full map is the hub.