Rolling, Blue-Green & Canary — Strategies, PDBs & Blast Radius
RollingUpdate trades surge capacity for gradual replacement. Blue-green switches a Service between two Deployments. Canary shifts a fraction of traffic (Flagger/Argo Rollouts). PDBs bound voluntary disruption so you do not drain yourself offline.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Compatible API, spare capacity exists
Prefer
RollingUpdate
The built-in strategy replaces Pods inside maxUnavailable and maxSurge. Mixed versions run together. Rollback is another rollout.
- No second full stack unless you set surge high.
- Revision history is already on the Deployment.
- Default when old and new can serve the same requests.
Alternative
Blue-green
Two Deployments. Warm the idle color. Flip the Service selector. Rollback is flipping it back. The cut is all of the traffic at once.
- You pay about double capacity during the overlap.
- Instant switch, and instant switch back, if the old color is still up.
- Wrong when you needed 5 percent of users, not 100.
Abort path
Promotion is allowed to stop. The PDB still applies to the Pods you are about to evict.
- 1
New Pods become Ready
Readiness is the gate from the probe lesson. Unready Pods must not take the shift. - 2
Shift or scale
Rolling scales ReplicaSets. Blue-green flips a Service selector. A canary step moves a fraction. - 3
Read the SLIs
Error rate, latency, saturation. If they break, pause or undo. - 4
Restore
Scale the old ReplicaSet up or point the selector back. PDB still limits how many healthy Pods you may evict.
Overview
A deploy strategy answers two questions: how many users see the new version before you know it is safe, and how fast can you stop? Rolling, blue-green, and canary are different answers. A PodDisruptionBudget is the constraint that keeps a voluntary drain from taking the Service offline while you do it.
Capacity to surge comes from requests and limits. A surge Pod that cannot schedule is a stalled rollout, not a faster one.
RollingUpdate math
For replicas = N, maxUnavailable = U, maxSurge = S:
- Minimum available during the update is about
N - Uwhen U is an absolute count. - Maximum Pods during the update is
N + S. - Percentages: Kubernetes rounds maxUnavailable down and maxSurge up. Verify with the live objects. A sketch that floors both will under-count surge.
- U and S cannot both be 0. The rollout would be unable to move.
Example: N is 10, maxUnavailable 20 percent, maxSurge 25 percent. 20 percent of 10 rounds down to 2. 25 percent of 10 rounds up to 3. You tolerate about 2 unavailable and you may run 13 Pods. Higher surge finishes faster and costs more capacity. Higher unavailable finishes faster and risks the SLO.
maxUnavailable: 0 is surge-only. You never intentionally drop the ready count. You must have spare CPU and memory requests for the extra Pods.
The Deployment controller applies this by scaling the new ReplicaSet up and the old one down. That mechanism is Pods, ReplicaSets & Deployments. This page is the blast radius of the choice.
Blue-green is a Service selector
- Green runs the current version. The Service selects
version: green. - Blue is a second Deployment with the new version and
version: blue. It is not selected yet. - Wait until Blue is Ready. Run smoke checks. Finish migrations that are safe before the flip.
- Change the Service selector to
version: blue. In-cluster traffic moves in one API update. - Keep Green scaled for a fast flip back, then tear it down.
You get an instant cutover and an instant rollback, at about 2x capacity, with a 100 percent blast radius. Connection draining still matters: in-flight calls on Green must finish. That is readiness plus preStop on the probe page, not a second copy of edge-proxy design.
This is a Service label change. It is not an Ingress object and it is not a Gateway route. If the cut is implemented as weights on north-south traffic, that configuration lives in Ingress Controllers & North-South. Do not re-teach it here. Name the selector flip, then point at that page for the edge.
Canary is a fraction, and fractions lie
Three different mechanisms get called "canary":
- A small replica count on the new ReplicaSet, then watch metrics, then shift more. Clients are not guaranteed to hit new Pods in proportion to replica count. A Service usually spreads connections, and long-lived connections stick.
- A progressive-delivery controller (Flagger or Argo Rollouts) with steps and an analysis template that can abort.
- A real traffic percentage. That weight is applied by a mesh or by the edge. The weight is not this cluster's lesson. The link is Ingress Controllers & North-South and the Private Networking hub.
Pros: the blast radius can be small, and promotion can be tied to SLIs. Cons: the metric becomes part of deploy correctness. A bad dashboard promotes a bad build. Replica canary without a weight is not "5 percent of users."
Abort a native rolling update with kubectl rollout pause or kubectl rollout undo. Argo Rollouts can abort when analysis fails and shift back. Either way, say what happens to Pods already in the new revision.
PodDisruptionBudget
minAvailable: keep at least this many Pods available (count or percent).maxUnavailable: voluntary disruption may take at most this many down.- Pick one. They are mutually exclusive.
- The budget applies to voluntary disruptions: node drain, cluster upgrade eviction, and Deployment scale-down. A node crash or a kernel OOM is involuntary. The PDB does not stop it.
minAvailableequal to the replica count (one replica andminAvailable: 1, or 100 percent) blocks drain of that Pod. Upgrades stick until you scale up or change the budget.- Unready Pods count against the budget. A probe outage can freeze evictions even when you meant to allow them.
HPA scale-down is also a voluntary disruption. A strict PDB can leave replicas higher than the autoscaler wants. That interaction is the autoscaling lesson.
Decisions
- 1
1. Start rollout step
- next2. New Pods go Ready
- 2
2. New Pods go Ready
- next3. Shift or scale
- 3
3. Shift or scale
- next4. SLIs still healthy
- ?
4. SLIs still healthy
- yes5. Continue or finish
- no6. Pause or abort
- 5
5. Continue or finish
- next9. Scale old revision down
- 6
6. Pause or abort
- next7. Restore the old RS
- 7
7. Restore the old RS
- next8. PDB keeps capacity
- 8
8. PDB keeps capacity
- 9
9. Scale old revision down
- next10. Voluntary drain
- ?
10. Voluntary drain
- ok11. Evict inside PDB
- breach12. PDB denies eviction
- 11
11. Evict inside PDB
- 12
12. PDB denies eviction
Lesson map
Rolling, Blue-Green & Canary — Strategies, PDBs & Blast Radius
RollingUpdate trades surge capacity for gradual replacement. Blue-green switches a Service between two Deployments. Canary shifts a fraction of traffic (Flagger/Argo Rollouts). PDBs bound voluntary disruption so you do not drain yourself offline.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Start rollout step"] b["2. New Pods go Ready"] c["3. Shift or scale"] d["4. SLIs still healthy"] a -->|1. Start rollout step| b b -->|2. New Pods go Ready| c c -->|3. Shift or scale| d
Bounds, in memory
Unavailable percentages round down. Surge percentages round up. Absolute numbers pass through. This matches the Deployment rolling-update rules, including the 10-replica example above.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
For 10 replicas you should see max unavailable 2, max surge 3, min available 8, max Pods 13. For 4 with absolutes 1 and 1: min available 3, max Pods 5. For 3 with unavailable 0 and surge 1: min available stays 3, max Pods 4.
Interview Q&A
What does maxSurge control?
Answer
How many Pods above the desired count may exist during the update. Surge spends extra requests (CPU and memory) to bring the new revision up before the old one is fully gone. It is why a faster rollout needs free schedulable capacity.
Blue-green versus rolling, blast radius?
Answer
Rolling mixes versions and moves a few Pods at a time, bounded by unavailable and surge. Blue-green moves the Service selector, so the next connection is 100 percent new. Rolling's risk is mixed behavior. Blue-green's risk is everyone at once, plus the capacity bill.
Why isn't a replica canary equal to a traffic canary?
Answer
Replica count changes how many Pods exist. It does not assign a percentage of users. Connection reuse and load-spreading decide who lands where. A true percentage is a weight on a mesh or at the edge. That weight is Ingress Controllers & North-South, not a Deployment field.
What does a PDB block?
Answer
Voluntary disruptions that would break minAvailable or maxUnavailable: drain, eviction for upgrade, and scale-down. It does not block a node failure. A budget of "all Pods must stay" blocks the drain you need for the upgrade.
Can a PDB stop loss from a dead node?
Answer
No. Involuntary disruption is outside the budget. You still need replicas in more than one zone, a topology spread, and a strategy that can replace the Pod after the node is gone. The PDB only guards the disruptions Kubernetes asks permission for.
How do you abort a bad rolling update?
Answer
kubectl rollout undo or pause. Undo is itself a rollout back to the previous ReplicaSet. With Argo Rollouts, analysis failure can abort and shift traffic back. Say whether the new Pods are scaled down or left in place for debugging.
Why set maxUnavailable to 0?
Answer
You refuse to reduce the ready count on purpose. The update can only proceed by surging. If the cluster cannot schedule the extra Pod, the rollout waits. That is safer for a tight SLO and useless if you have no spare requests.
What does a 1-replica PDB with minAvailable 1 do?
Answer
It forbids voluntary eviction of that only Pod. Drains and upgrades stall until you add a replica or relax the budget. The PDB is doing what you asked. The ask was incompatible with a one-Pod Service.
How do percentages round?
Answer
maxUnavailable rounds down. maxSurge rounds up. 25 percent surge of 10 is 3, not 2. Both values at 0 are invalid. Re-check the Deployment status rather than trusting a mental floor on both sides.
Where do you put the edge weight if canary is not replica count?
Answer
On the traffic path that can actually split requests. In this site's map that is Ingress Controllers & North-South, inside Private Networking. This page stops at the Service and the ReplicaSets.
Pitfalls
- Surge set high on a cluster with no free requests. The new Pods stay Pending and the progress deadline fires.
- Calling a 1-of-10 new Pod a "10 percent canary" without measuring the share of requests.
- A PDB copied from a template that sets minAvailable to 100 percent.
- Blue-green that deletes Green in the same minute as the flip, so rollback has nothing to flip back to.
- Teaching host and path routing in the middle of a PDB answer.
You have 12 replicas, a breaking response change for 1 percent of clients, and spare capacity for 2 extra Pods. Choose rolling, blue-green, or canary. State the blast radius and the PDB you would set so a drain cannot take you below 10 available.