Pods, ReplicaSets & Deployments — Desired State & Controllers
The Deployment controller owns ReplicaSets that own Pods. Labels and selectors wire the tree. maxUnavailable/maxSurge and revision history define how a rollout moves. Prefer Deployments over hand-managing ReplicaSets.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Who should move the replicas
Prefer
A Deployment
Rolling update, revision history, rollback, pause, and a progress deadline are already implemented. You still have to understand the ReplicaSets it creates.
- Each template change is a revision you can undo.
- maxSurge and maxUnavailable bound the transition.
- Default for nearly every stateless service.
Alternative
ReplicaSets you edit by hand
You can see each generation, and you also reinvent rollout math, rollback, and cleanup. Stale ReplicaSets linger.
- Useful for a whiteboard, almost never for production.
- Easy to adopt the wrong Pods if selectors overlap.
- Progressive controllers (Argo Rollouts, Flagger) exist when a Deployment strategy is not enough.
One template change
The API server stores desired state. Two controllers close the gap. The scheduler and kubelet are downstream of that.
- 1
CI updates the template
New image or new replica count. The selector stays put. - 2
Deployment controller
Creates or scales the ReplicaSet for this revision. The previous ReplicaSet scales down. - 3
ReplicaSet controller
Creates or deletes Pods until its own replica count matches. - 4
Schedule and start
Scheduler binds using requests. Kubelet pulls the image and starts containers. - 5
Status
Ready replicas climb, or the rollout hits the progress deadline.
Overview
Most Kubernetes mysteries start here. Pods disappear, a Deployment recreates them, and kubectl get rs shows three rows. If you cannot narrate the ownership chain, you cannot debug a stuck rollout.
The hub is the map. This page is the tree: labels, selectors, ReplicaSet ownership, and the rollout knobs on the Deployment.
The Pod is the atom
A Pod is one or more containers that share a network namespace and volumes. It is the smallest thing you deploy. It is also ephemeral. Controllers replace Pods. SSH-and-fix is not a production policy, because the next reconcile deletes your edit or replaces the Pod.
Labels identify Pods. Selectors bind Services and controllers to those labels. A Service does not "own" the Pod the way a ReplicaSet does. It selects Ready Pods. How that Service is published north-south is Ingress Controllers & North-South, not this page.
Labels, selectors, and ownership
- The Pod template labels must match the Deployment
spec.selector.matchLabels. On a modern cluster the API rejects a mismatch. The selector is immutable after create. - The ReplicaSet selects those Pods and owns them with
ownerReferences. - Careless label edits orphan Pods (no owner) or let a ReplicaSet adopt Pods it should not touch.
- Two Deployments must not share a selector. Overlapping selectors make controllers fight, and ownership becomes undefined.
Deleting a ReplicaSet with orphan cascade leaves the Pods behind until something adopts or deletes them. That is an orphaned Pod: the owner is gone, or the labels no longer match any controller.
What a ReplicaSet actually does
Its job is narrow: keep the number of matching Pods at spec.replicas. It does not roll out a new template by itself. The Deployment orchestrates several ReplicaSets, one per revision.
Old ReplicaSets often stay at zero replicas so kubectl rollout undo can scale them back up. spec.revisionHistoryLimit is how many of those old objects you keep. The default is 10. Set it too low and you cannot roll back far. Set it huge and you accumulate empty ReplicaSets.
Rollout knobs you should know cold
| Field | What it decides |
|---|---|
spec.replicas | Desired Pod count for the active strategy |
spec.strategy.type | RollingUpdate (default) or Recreate |
maxUnavailable | How many can be down during the update |
maxSurge | How many extra Pods may exist above desired |
revisionHistoryLimit | How many old ReplicaSets to keep |
progressDeadlineSeconds | When a stalled rollout is marked failed |
Example: replicas 4, maxUnavailable 1, maxSurge 1. During the update you may have as few as 3 available and as many as 5 Pods total. Percentages, PDB interaction, and blue-green versus canary are the rollout lesson. Kubernetes rounds an unavailable percentage down and a surge percentage up. Both cannot be 0.
Recreate kills the old Pods before it starts the new ones. Use it when two versions cannot coexist: an exclusive volume, or a protocol that breaks if both versions run. You are accepting downtime. Say that out loud.
Pause freezes the strategy. That is how a human can watch a partial rollout without deleting anything. Resume continues. Undo scales a prior revision's ReplicaSet back up (kubectl rollout undo, optionally --to-revision).
progressDeadlineSeconds is the clock for a rollout that never finishes: ImagePullBackOff, quota, or a Pod that stays unschedulable. The condition is ProgressDeadlineExceeded. It does not fix the cause. It stops you from staring at "waiting" forever.
Commands that matter: kubectl rollout status, pause, resume, undo, and history.
Flow
- 1
1. CI updates Deployment
- next2. API stores desired
- 2
2. API stores desired
- next3. Deployment wakes
- 3
3. Deployment wakes
- next4. Scale new and old RS
- 4
4. Scale new and old RS
- next5. ReplicaSet makes Pods
- 5
5. ReplicaSet makes Pods
- next6. Scheduler binds Node
- 6
6. Scheduler binds Node
- next7. Kubelet starts them
- 7
7. Kubelet starts them
- next8. Status shows actual
- 8
8. Status shows actual
Lesson map
Pods, ReplicaSets & Deployments — Desired State & Controllers
The Deployment controller owns ReplicaSets that own Pods. Labels and selectors wire the tree. maxUnavailable/maxSurge and revision history define how a rollout moves. Prefer Deployments over hand-managing ReplicaSets.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. CI updates Deployment"] b["2. API stores desired"] c["3. Deployment wakes"] d["4. Scale new and old RS"] a -->|1. CI updates Deployment to| b b -->|2. API stores desired| c c -->|3. Deployment wakes| d
Flow
- 1
1. Desired replicas is 4
- next2. maxUnavailable is 1
- 2
2. maxUnavailable is 1
- next3. Keep at least 3 Ready
- 3
3. Keep at least 3 Ready
- next4. maxSurge is 1
- 4
4. maxSurge is 1
- next5. At most 5 Pods
- 5
5. At most 5 Pods
- next6. New RS up, old RS down
- 6
6. New RS up, old RS down
Desired versus actual, in memory
No cluster. The simulator only closes a count gap. A real ReplicaSet also selects by labels and will not delete Pods it does not own.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The first pass creates three Pods. Scale-up creates two more. Scale-down deletes down to two. A Deployment rollout is this loop on two ReplicaSets at once, with surge and unavailable caps.
Interview Q&A
What is the ownership chain for a Deployment?
Answer
Deployment, then one or more ReplicaSets, then Pods. The links are ownerReferences plus label selectors. The Deployment decides which ReplicaSet should hold the replicas. The ReplicaSet decides which Pod objects exist.
Why do I see multiple ReplicaSets for one Deployment?
Answer
Each template change creates a new revision. The old ReplicaSet scales toward zero and can remain for rollback until revisionHistoryLimit prunes it. Three rows in kubectl get rs usually means three templates, not three bugs.
What if selector labels do not match the template?
Answer
Current API validation rejects the Deployment. Historically a mismatch orphaned Pods or thrashed them. Do not "fix" a live selector. It is immutable. Create a new Deployment if the selector must change.
Walk maxUnavailable 1 and maxSurge 1 with 4 replicas.
Answer
You may drop to 3 available Pods, and you may run 5 Pods at once. The new ReplicaSet grows inside the surge budget. The old one shrinks inside the unavailable budget. Readiness, not mere Running, is what "available" means. Probes are the next lesson.
What does rollout pause do?
Answer
It freezes the strategy so you can inspect a partial shift. It does not delete the new ReplicaSet. Resume continues the same rollout. It is a manual observation hook, not a canary controller. Weighted traffic is a different problem and lives with Ingress or a progressive-delivery controller.
How do you roll back?
Answer
kubectl rollout undo scales the previous revision's ReplicaSet up and the bad one down. --to-revision picks an older history entry if you still have it. Rollback is another rollout. It still has to respect surge, unavailable, and probes.
What is an orphaned Pod?
Answer
A Pod whose owner was deleted with orphan cascade, or whose labels no longer match a controller selector. Nothing will recreate it when it dies, and nothing will delete it when you scale. That is usually a mistake.
When is Recreate the right strategy?
Answer
When two versions cannot run together: a single-writer volume, or a schema migration that forbids mixed binaries. You are choosing downtime. If they can coexist, RollingUpdate is the default.
What does progressDeadlineSeconds catch?
Answer
A rollout that makes no progress before the deadline: ImagePullBackOff, a quota that blocks new Pods, a scheduler that cannot place the requests. The Deployment condition becomes ProgressDeadlineExceeded. The Pod events still hold the real reason.
Can two Deployments select the same Pods?
Answer
They must not. Overlapping selectors cause two controllers to adopt and fight. Give each Deployment its own selector and keep the template labels equal to that selector.
Pitfalls
- Editing a live container and expecting the edit to survive the next Pod.
- Changing selector labels after create, or letting two apps share
app: web. - Leaving
revisionHistoryLimitat the default and then being surprised you cannot undo a revision from last month. Or the reverse: never pruning, and scrolling a sea of scaled-to-zero ReplicaSets. - Setting maxUnavailable and maxSurge both to 0. The API rejects that. You would be asking for a rollout that cannot move.
- Debugging "3 ReplicaSets" as a leak before you check rollout history.
You ship a new image. The Deployment shows 2 available of 4. kubectl get rs lists the new ReplicaSet at 2 replicas and the old one at 2. Say which controller set those numbers, and what you check if the new Pods stay at 0 Ready.