Progressive Delivery & Kill Switches — Canary, Ramp, Instant Rollback
Progressive delivery ramps a feature to a larger audience only while metrics stay healthy. A kill switch forces the safe path the moment they do not. This lesson contrasts flag ramps with Kubernetes traffic canaries, designs soak times and abort gates, and defines who may flip a kill switch and how wide the blast radius is.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does a flag ramp move?
Answer
Which users see the new behavior, inside a binary that is already running.
L2
What does a Kubernetes canary move?
Answer
Traffic or pods onto a new image. Rollback is an undo of that rollout, on the order of minutes, not a rule flip.
L3
Why soak between steps?
Answer
Some metrics are delayed. A jump from zero to 100 hides a slow burn that a two-hour soak at 1 percent would have shown.
L4
What aborts a ramp?
Answer
Error rate, latency, or sample ratio mismatch past a gate you wrote down. A business guardrail can abort even when the process is up.
L5
What makes a kill switch trustworthy?
Answer
It wins evaluation order, it has an owner and an audit trail, the safe variant is known, and caches that stored the bad variant get purged.
L6
Who may flip a global payments flag?
Answer
Dual-control, not anyone with console access. A service-level ops kill can be the on-call, still audited.
L7
When is a flag ramp the wrong rollback?
Answer
A crash loop needs the Deployment rolled back. A schema change needs the migration plan. A CVE needs a patch. A flag only helps when a safe path is already in the binary.
Failure modes
Zero to 100
No soak. The delayed metric arrives after everyone is on the new path.
Kill rule under the percentage
The switch is a normal targeting rule, so a later match turns the feature back on.
Cached HTML
The flag flipped and the CDN still serves the bad variant until something purges it.
Flag used on a crash
The process dies before evaluation. The rollback that works is the previous image.
Misconceptions
A flag canary replaces a Deployment canary.
They move different things. Use the rollout for the binary and the flag for the behavior.
Git is always fast enough for a kill.
Reviewed defaults belong in Git. An incident needs an audited remote override.
Global off is the responsible first move.
The smallest blast radius that stops the harm is the better first move: one tenant or one region, then global if you must.
Interviewer traps
Describe maxUnavailable, probes, and a PDB.
Say those belong to the Kubernetes canary. This answer is the percentage, the soak, and the kill precedence.
Ramp a schema change to 10 percent of rows.
That is a migration. Point at expand/contract and the cutover flag in the databases series.
Design scenario
Same prompt for every reader.
Requirements
Staff dogfood, then 1, 5, 25, 50, and 100 percent with soaks. Abort on error rate, p99, or sample ratio mismatch. A kill switch forces off and beats the percentage. Dual-control before the large payments step.
Traffic / scale
Sticky users, not a fresh coin flip. Internal segment first.
Latency
p99 above the gate aborts. The gate is a threshold you can say out loud, not a dashboard feeling.
Consistency
Killed means off even for bucket 0 at 100 percent. Every service uses that precedence.
Availability
The safe variant is already in the binary. A crash loop is a Deployment rollback, not a flag.
Failure assumptions
- The metric pipeline lags the ramp.
- A CDN may have cached the treatment HTML.
- The console is reachable by more people than should flip a global kill.
Constraints
- Do not reteach rolling update math.
- Do not invent a new error-budget formula.
Prompt
A new checkout path is in the binary. Pods are healthy. You may ramp a sticky percentage. Payments leadership wants a person in the loop before half of traffic moves.
API
What does evaluation return when killed is true and the percent is 100?
Data
Which three numbers abort the ramp?
Architecture
What is the smallest blast radius you try before a global force-off?
Checkout misbehaves and the pods are healthy
Prefer
Kill the flag
The safe path is already in this binary. One audited flip returns it for the blast radius you named.
- Seconds, if the rule is first in evaluation order.
- The previous image stays where it is.
- You still purge a cache that baked the bad HTML.
Alternative
Undo the Deployment
Right when the new image crashes, deadlocks, or never becomes Ready.
- Minutes, and it moves every behavior in that template.
- Wrong tool for a bad path behind a flag.
- Probes and surge belong to the Kubernetes lesson.
Ship, ramp, gate, kill or graduate
The binary and the behavior are different rollouts. Do not collapse them.
- 1
Ship the image
CI produces an artifact. The cluster rollout is healthy before any user-facing percentage moves. - 2
Dogfood, then 1 percent
Staff for a day. Then a sticky 1 percent for a few hours. Watch p99, errors, and the primary metric. - 3
Soak and raise
5, 25, 50, 100. Minimum hold times. A person approves the high-blast payments step. - 4
Abort or kill
Guardrail or sample ratio mismatch forces the safe variant. The kill rule beats the percentage. - 5
Graduate
At 100 percent and stable, make the winner the default and delete the flag. Cleanup is the last page.
Overview
Progressive delivery increases the audience only while the metrics stay healthy. A kill switch is the instant path back to the safe variant when they do not. Flags gate behavior. Orchestrators gate binaries.
Use both. The image moves through CI/CD Pipelines and the workload rollout on Kubernetes Workloads. How that rollout shifts pods is Rolling, Blue-Green & Canary. The risky UX on top of a healthy binary is this page.
A flag cannot fix a segfault. A Deployment rollback is the slow tool when the bug is a bad weight or a bad branch already loaded behind a flag. Kill the flag.
Decisions
- 1
1. Ship the binary
- next2. Pods look healthy
- 2
2. Pods look healthy
- next3. Sticky percent ramp
- 3
3. Sticky percent ramp
- next4. Guardrails hold
- ?
4. Guardrails hold
- yes5. Soak, then raise
- no6. Force the safe path
- 5
5. Soak, then raise
- next7. Graduate and delete
- 6
6. Force the safe path
- 7
7. Graduate and delete
Lesson map
Progressive Delivery & Kill Switches — Canary, Ramp, Instant Rollback
Progressive delivery ramps a feature to a larger audience only while metrics stay healthy. A kill switch forces the safe path the moment they do not. This lesson contrasts flag ramps with Kubernetes traffic canaries, designs soak times and abort gates, and defines who may flip a kill switch and how wide the blast radius is.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB ship["1. Ship the binary"] k8s["2. Pods look healthy"] ramp["3. Sticky percent ramp"] gate["4. Guardrails hold"] ship -->|1. Ship the binary| k8s k8s -->|2. Pods look healthy| ramp ramp -->|3. Sticky percent ramp| gate
Three ways to move a slice of traffic
| Mechanism | What moves | Rollback speed | Usual owner |
|---|---|---|---|
| Flag percentage | Users seeing new behavior in the same binary | Seconds, one flip | Product, ops, and engineering |
| Kubernetes rolling or canary | Pods on a new image | Minutes, rollout undo | Platform or SRE |
| Service mesh weight | A share of requests to version B | Seconds to minutes | Platform |
The Kubernetes docs describe the rolling update. Read them for the binary. Do not paste that model into a flag design.
A ramp you can defend
- Internal segment, staff, about a day of dogfood.
- Sticky 1 percent for two to four hours. Watch p99, error rate, and the primary metric.
- 5 percent, then 25, then 50, then 100, each with a minimum soak.
- Abort if a guardrail burns the error budget or sample ratio mismatch shows up.
Automate the gates when the flag vendor can see the metrics. Keep a human on the high-blast steps, especially past 50 percent on payments. Soak time exists because retention and batch pipelines are late. Instant 0 to 100 hides a slow burn.
The error-budget signal is SLOs, Error Budgets & Distributed Tracing. Sample ratio mismatch is defined on the experimentation page. This page only needs the abort.
ProblemA quiet metric snapshot continues. A 5 percent error rate aborts. Sample ratio mismatch aborts even when latency is fine.
ExpectedThe healthy snapshot is ok. The error snapshot and the SRM snapshot abort, and the reason names which gate fired.
Edge cases
- A latency breach with a clean error rate still aborts.
- Values on the gate are not a breach.
- Test: healthy snapshot continues
should_abort(Metrics(0.002, 400, 1.0)) == (False, 'ok') - Test: errors abort
should_abort(Metrics(0.05, 400, 1.0)) == (True, 'errors') - Test: SRM aborts before latency
should_abort(Metrics(0.002, 50, 9.0)) == (True, 'SRM') - Test: latency aborts
should_abort(Metrics(0.002, 1200, 1.0))[0] is True
Press Run. Snippets must be self-contained — no network, files, or native modules.
ProblemA killed flag is off even at 100 percent. An unkilled user in bucket 5 is on at 10 percent.
Expectedkilled with percent 100 is off. bucket 5 at 10 percent is on. bucket 20 at 10 percent is off.
Edge cases
- Bucket 0 at 100 percent is on only when the flag is not killed.
- Percent 0 is off for every bucket.
- Test: kill beats 100 percent
evaluate({ killed: true, bucket: 0, percent: 100 }) === 'off' - Test: bucket inside the percent is on
evaluate({ killed: false, bucket: 5, percent: 10 }) === 'on' - Test: bucket outside the percent is off
evaluate({ killed: false, bucket: 20, percent: 10 }) === 'off'
Press Run. Snippets must be self-contained — no network, files, or native modules.
What a kill switch needs
- Named and rehearsed. Not an ad-hoc rule buried under three segments.
- Fail-safe default. Evaluation prefers the safe variant when the switch is on.
- A blast-radius label. One tenant, one endpoint, or global. Try the small one first. Killing
new_checkoutfor one region beats a global force-off when the region is the blast. - Auth and audit. Who flipped it, when, and the ticket. Dual-control for a global payments or auth UX kill.
- Dependents. A CDN or a cache that baked variant HTML needs a purge. The flip alone does not evict them.
- A runbook. The incident channel knows the switch exists before the incident.
A kill switch that disables a dependency is a cousin of shedding load or opening a circuit. Overload still wants a coded bulkhead. That family is Resilience Patterns. Use the flag for feature blast radius, not as a substitute for those controls.
Who is allowed to flip production
| Flag class | Who can flip production | Extra control |
|---|---|---|
| Release or experiment | Feature owner and on-call | Automated guardrails |
| Ops kill for one service | On-call | Audit, and a note in the postmortem |
| Global payments or auth UX | Two people | Break-glass is documented |
GitOps for the defaults plus a console for the emergency is the usual shape. The console is not open to the company. Ordinary promotion still rides the pipeline.
When a flag ramp is the wrong tool
- Crash loops and a binary that never becomes Ready. Roll the Deployment back. The flag code is not running.
- Data migration correctness. Expand and contract, dual-write, and a deliberate cutover. A percentage of rows "feeling" a new schema without that plan will diverge. The cutover flag in that world is Rollback, Feature Flags & Cutover. It is a databases lesson. It is not a page in this series.
- A security CVE. Patch and deploy. A flag helps only when a safe alternate path is already shipped.
Interview Q&A
Flag canary versus Kubernetes canary?
Answer
A flag canary changes behavior for a percentage of users on the same build. A Kubernetes canary shifts traffic onto a new build. The failure modes and the rollback tools differ. Use the flag when the pods are healthy and the path is wrong. Use the rollout when the process itself is wrong.
What makes a kill switch trustworthy?
Answer
It wins over normal rules, access is audited, the runbook was rehearsed, the safe variant is explicit, and you know which caches still hold the old HTML. A switch nobody has flipped in a game day is a rumor.
How do SLOs affect a ramp?
Answer
Burn-rate alerts and the error budget should abort or freeze the next step. The budget math stays on the SLO page. This page only requires that a breach is an abort, not a debate.
Why soak between ramp steps?
Answer
Delayed metrics need time: batch jobs, next-day retention, a slow error that is not in the first minute. Jumping from zero to 100 hides that burn until it is everyone's problem.
Can the kill switch live only in Git?
Answer
Reviewed changes should. An incident often cannot wait for a merge. Keep an emergency remote override with strong IAM and an audit record. If the only off switch is a pull request, you do not have a kill switch.
Give a blast-radius example.
Answer
Turn off new_checkout for one region or one tenant before you force it off globally. Widen only if the smaller flip does not stop the harm. Write the radius on the switch so the incident channel does not guess.
Mesh weight versus a flag?
Answer
A mesh weight sends a share of requests to version B. It is still a binary split, owned by the platform, and it is fast. It does not select a behavior inside one binary. Use it when the thing you fear is the new build, not a branch inside the current one.
What do you ramp with a flag, and what do you refuse?
Answer
User-visible behavior that already has a safe path in this image. You refuse crash loops, schema migrations, and vulnerabilities that need a patch. Those have their own rollback tools.
Pitfalls
- Starting the percentage before the new pods are Ready.
- Putting the kill rule after the percentage so a segment turns the feature back on.
- Approving 100 percent because the first hour looked quiet.
- Forgetting the CDN after the flip.
- Letting anyone with the console URL kill payments globally.
- Using a flag to mean "10 percent of rows use the new column" with no dual-write story.
Say the next percentage, the soak, the three abort numbers, and who has to approve the step after 25 percent. Then say the one-region kill you would use if p99 blew through the gate, and the page you would open if the pods were crash-looping instead.
Go deeper
- Martin Fowler on feature toggles is where ops toggles sit next to release toggles.
- LaunchDarkly progressive rollouts and Unleash feature flags show automated steps and kill switches in product.
- Kubernetes Deployments are the contrast for the binary. Do not duplicate that lesson here.
- The SRE workbook chapter on canarying releases is the metrics-gated rollout view.
Next: Flag Architecture & Hygiene.