Kubernetes Internals - Declarative Reconciliation from kubectl apply to Running Pod
Hub: Kubernetes as a database of intentions plus independent watch-driven loops; imperative vs declarative reconciliation; the full kubectl apply -> running pod path (authn/authz, admission, etcd, Deployment/ReplicaSet controllers, scheduler, kubelet, CRI/CNI/CSI); component contracts, bottlenecks and a layer-picking decision chart. Goes one layer below the existing Workloads series.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does a successful kubectl apply actually guarantee?
Answer
Only that the API server accepted, validated and stored the object in etcd. Scheduling, image pulls, volumes and probes all happen afterwards.
L2
Imperative vs declarative: what is the difference in one line?
Answer
Imperative sends commands that are forgotten after they run. Declarative stores desired state that controllers keep enforcing.
L3
Which component is the only one that talks to etcd?
Answer
The kube-apiserver. Every other component is a client of the API server.
L4
Who decides which node a pod runs on, and who starts it?
Answer
The scheduler writes a binding. The kubelet on that node starts containers through the container runtime.
L5
What keeps running when the control plane is down?
Answer
Pods that are already running keep serving. You lose changes: deploys, rescheduling, scaling and recovery from node failure.
L6
Where should you block a bad object, and why there?
Answer
Validating admission, because the object is rejected before it reaches etcd, so no controller ever acts on it.
L7
When would you not choose Kubernetes?
Answer
A handful of stateless services and a small team are often better served by a PaaS or serverless containers, which have no control plane to run.
Failure modes
Apply succeeded but nothing runs
The object is stored, but pods stay Pending or ContainerCreating because of scheduling, image or volume problems.
A down admission webhook blocks every deploy
A webhook with failurePolicy Fail whose backend is unavailable rejects or times out all matching writes.
Hand edits are reverted
kubectl edit or kubectl scale changes are undone by the controller, Helm upgrade or GitOps agent that owns the field.
Misconceptions
kubectl apply deploys the app.
It stores intent. The rollout happens later, driven by controllers, the scheduler and the kubelet.
Components call each other in a pipeline.
They only talk to the API server and react to watch events.
A control plane outage takes the apps down.
Running pods keep serving. Only changes and recovery stop.
Interviewer traps
Saying the scheduler starts containers.
The scheduler only binds pods to nodes. The kubelet runs them.
Treating etcd like an application database.
etcd holds small metadata objects for the whole cluster. Big or high-churn data hurts every client.
Design scenario
Same prompt for every reader.
Requirements
Block :latest and unsigned images before they run, keep replicas at the desired count after node failure, and show engineers why a deploy is stuck.
Failure assumptions
- A node dies during business hours.
- The policy webhook backend is briefly unavailable.
- Someone runs kubectl scale by hand in production.
Constraints
- Managed Kubernetes; no access to etcd.
- No custom scheduler.
Prompt
Design the deploy path for 40 services onto a shared Kubernetes cluster so that bad images are blocked and a node loss heals without a human.
API
Which objects do teams write, and which admission policies validate them?
Data
What lives in etcd, and what status fields and events do engineers read to debug a rollout?
Architecture
Where do admission, controllers, the scheduler and the kubelet sit, and which failure does each one cover?
Overview
Kubernetes is easiest to understand as a database of intentions plus a crowd of loops that make reality match it. You never tell Kubernetes "start three containers". You write an object that says "three replicas of this image should exist", the API server validates it and stores it in etcd, and independent controllers each notice a gap between what is stored and what is real, then close their small part of that gap. The scheduler picks a node, the kubelet on that node asks the container runtime to start containers, the network plugin gives the pod an IP, and the storage plugin mounts its volumes. No component calls the next one directly. They all talk only to the API server and react to changes they watch.
That one design choice explains most of what you will meet in production: why a cluster self-heals, why it is eventually consistent rather than instant, why etcd and the API server are the scaling bottleneck, why a bad admission webhook can stop every deploy, why storage has its own binding rules, why Helm and GitOps exist at all, and why node autoscaling reacts to Pending pods rather than CPU graphs.
This series goes one layer below the existing Kubernetes Workloads series. That series teaches what Deployments, probes, requests and limits, HPA/VPA and rollout strategies do. This one teaches how the machine underneath works: the control plane, scheduling and reconciliation, persistent storage, packaging with Helm, and adding or removing nodes.
Imperative vs declarative: the mental model
| Question | Imperative system (a deploy script, docker run) | Declarative reconciliation (Kubernetes) |
|---|---|---|
| What do you send? | Commands: "start", "stop", "scale to 5" | Desired state: an object with a spec |
| Who remembers your intent? | Nobody after the script exits | etcd, as the stored object |
| What happens when a node dies? | Nothing until a human notices | A controller sees fewer pods than desired and creates more |
| What happens if a step is retried? | Possibly double work | Nothing extra: reconcile is idempotent |
| How do you know it worked? | Exit code of each step | status fields and conditions, written by controllers |
| Cost | Simple, fast, predictable ordering | Eventual consistency, more moving parts, harder debugging |
Imperative scripts or declarative reconciliation?
Prefer
Declarative reconciliation
Store the desired state and let loops keep closing the gap.
- After a node failure, reconcile #2 recreated the missing pods.
- Manual drift from 3 to 5 pods was deleted back to 3.
- Running reconcile again is a no-op, so retries are safe.
Alternative
An imperative deploy script
Run the steps once and exit; nobody remembers the intent.
- After the node failure, 1 pod was left and nobody noticed.
- Retrying a step can do the work twice.
- Success means each command exited, not that the system is healthy.
From kubectl apply to a running pod
Diagram 1 condensed. Sibling pages go deep on each layer.
- 1
Admit the request
Authentication, RBAC, mutating admission, schema and validating admission. - 2
Store the intent
The object lands in etcd with a new resourceVersion. - 3
Controllers fan out
Deployment and ReplicaSet controllers create Pods with no node. - 4
Schedule and run
The scheduler binds each Pod and the kubelet starts it through CRI, CNI and CSI. - 5
Fail early or stay Pending
A policy rejection writes nothing. No feasible node leaves the Pod Pending.
The tradeoff is real. Declarative control buys self-healing and safe retries, but you give up "it is done when the command returns". A kubectl apply that succeeds only means the API server accepted and stored your intent. Whether a pod ever runs depends on a scheduler decision, image pulls, volume attaches and probes that happen afterwards.
// Simulation: an imperative deploy script vs a declarative reconcile loop.
// Both start 3 replicas. Then the world changes behind their backs:
// a node dies (2 pods lost) and someone scales the app by hand.
type World = { pods: Set<string>; nextId: number };
const world: World = { pods: new Set(), nextId: 0 };
const startPod = (w: World): string => { const id = `web-${w.nextId++}`; w.pods.add(id); return id; };
// Imperative: "run these steps once". It has no memory of intent.
function imperativeDeploy(w: World, n: number): void {
for (let i = 0; i < n; i++) startPod(w);
}
// Declarative: store intent, and keep closing the gap between intent and reality.
const desired = { replicas: 3 };
function reconcile(w: World): string {
const have = w.pods.size;
if (have < desired.replicas) {
const made: string[] = [];
while (w.pods.size < desired.replicas) made.push(startPod(w));
return `have ${have}, want ${desired.replicas}: created ${made.join(", ")}`;
}
if (have > desired.replicas) {
const extra = [...w.pods].slice(desired.replicas);
extra.forEach((p) => w.pods.delete(p));
return `have ${have}, want ${desired.replicas}: deleted ${extra.join(", ")}`;
}
return `have ${have}, want ${desired.replicas}: no-op`;
}
// --- imperative run ---
imperativeDeploy(world, 3);
console.log("imperative after deploy:", [...world.pods].join(", "));
world.pods.delete("web-0"); world.pods.delete("web-1"); // node failure
console.log("imperative after node failure:", world.pods.size, "pod(s); nobody notices");
// --- declarative run on a fresh world ---
const w2: World = { pods: new Set(), nextId: 0 };
console.log("\nreconcile #1:", reconcile(w2));
w2.pods.delete("web-0"); w2.pods.delete("web-1"); // node failure
console.log("reconcile #2:", reconcile(w2));
startPod(w2); startPod(w2); // manual scale-up drift
console.log("reconcile #3:", reconcile(w2));
console.log("reconcile #4:", reconcile(w2), "(idempotent: safe to run any number of times)");Output:
imperative after deploy: web-0, web-1, web-2
imperative after node failure: 1 pod(s); nobody notices
reconcile #1: have 0, want 3: created web-0, web-1, web-2
reconcile #2: have 1, want 3: created web-3, web-4
reconcile #3: have 5, want 3: deleted web-5, web-6
reconcile #4: have 3, want 3: no-op (idempotent: safe to run any number of times)Expectedimperative after deploy: web-0, web-1, web-2 imperative after node failure: 1 pod(s); nobody notices reconcile #1: have 0, want 3: created web-0, web-1, web-2 reconcile #2: have 1, want 3: created web-3, web-4 reconcile #3: have 5, want 3: deleted web-5, web-6 reconcile #4: have 3, want 3: no-op (idempotent: safe to run any number of times)
Press Run. Snippets must be self-contained — no network, files, or native modules.
From kubectl apply to a running pod
The diagram below is the end-to-end path for creating a Deployment. Every arrow into etcd goes through the API server, and every arrow out of it is a watch notification, not a direct call.
Flow
- 1
Step 1: kubectl apply sends the Deployment to kube-apiserver
- nextStep 2: authentication then RBAC authorization
- 2
Step 2: authentication then RBAC authorization
- nextStep 3: mutating admission sets defaults and labels
- 3
Step 3: mutating admission sets defaults and labels
- nextStep 4: schema validation then validating admission policies
- 4
Step 4: schema validation then validating admission policies
- nextStep 5: object stored in etcd with a new resourceVersion
- nextFailure path: request denied, nothing written, controllers never see it
- 5
Step 5: object stored in etcd with a new resourceVersion
- nextStep 6: Deployment controller watches, creates a ReplicaSet
- 6
Step 6: Deployment controller watches, creates a ReplicaSet
- nextStep 7: ReplicaSet controller creates Pods with empty nodeName
- 7
Step 7: ReplicaSet controller creates Pods with empty nodeName
- nextStep 8: scheduler filters and scores nodes, writes a Binding
- 8
Step 8: scheduler filters and scores nodes, writes a Binding
- nextStep 9: kubelet on that node sees the bound Pod
- nextFailure path: Pod stays Pending, preemption or node autoscaler may act
- 9
Step 9: kubelet on that node sees the bound Pod
- nextStep 10: CRI pulls image and starts containers, CNI assigns IP, CSI mounts volumes
- nextFailure path: ImagePullBackOff or ContainerCreating with attach or mount errors
- 10
Step 10: CRI pulls image and starts containers, CNI assigns IP, CSI mounts volumes
- nextStep 11: probes pass, kubelet reports Ready, endpoints updated
- 11
Step 11: probes pass, kubelet reports Ready, endpoints updated
- 12
Failure path: request denied, nothing written, controllers never see it
- 13
Failure path: Pod stays Pending, preemption or node autoscaler may act
- 14
Failure path: ImagePullBackOff or ContainerCreating with attach or mount errors
Lesson map
Kubernetes Internals - Declarative Reconciliation from kubectl apply to Running Pod
Hub: Kubernetes as a database of intentions plus independent watch-driven loops; imperative vs declarative reconciliation; the full kubectl apply -> running pod path (authn/authz, admission, etcd, Deployment/ReplicaSet controllers, scheduler, kubelet, CRI/CNI/CSI); component contracts, bottlenecks and a layer-picking decision chart. Goes one layer below the existing Workloads series.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB u["Step 1: kubectl apply sends the Deployment to kube-apiserver"] a["Step 2: authentication then RBAC authorization"] m["Step 3: mutating admission sets defaults and labels"] v["Step 4: schema validation then validating admission policies"] e["Step 5: object stored in etcd with a new resourceVersion"] d["Step 6: Deployment controller watches, creates a ReplicaSet"] r["Step 7: ReplicaSet controller creates Pods with empty nodeName"] s["Step 8: scheduler filters and scores nodes, writes a Binding"] k["Step 9: kubelet on that node sees the bound Pod"] c["Step 10: CRI pulls image and starts containers, CNI assigns IP, CSI mounts volumes"] p["Step 11: probes pass, kubelet reports Ready, endpoints updated"] x["Failure path: request denied, nothing written, controllers never see it"] y["Failure path: Pod stays Pending, preemption or node autoscaler may act"] z["Failure path: ImagePullBackOff or ContainerCreating with attach or mount errors"] u -->|continues| a a -->|continues| m m -->|continues| v v -->|continues| e e -->|continues| d d -->|continues| r r -->|continues| s s -->|continues| k k -->|continues| c c -->|continues| p v -->|continues| x s -->|continues| y k -->|continues| z
| Step | Component | What it reads | What it writes | Where this series covers it |
|---|---|---|---|---|
| 1 to 5 | kube-apiserver, admission, etcd | Your request | The object, a resourceVersion | Control plane page |
| 6 to 7 | kube-controller-manager (Deployment, ReplicaSet controllers) | Deployments, ReplicaSets, Pods via informers | ReplicaSets, Pods | Scheduler, controllers and kubelet page; behavior in Pods, ReplicaSets & Deployments |
| 8 | kube-scheduler | Unbound Pods, Nodes, PVCs | Pod binding (spec.nodeName) | Scheduler page; requests in Requests, Limits & QoS |
| 9 to 10 | kubelet, container runtime (CRI), CNI plugin, CSI node plugin | Pods bound to its node | Pod status, node status | Scheduler/kubelet page and Storage page |
| 11 | kubelet probes, EndpointSlice controller | Probe results | Ready condition, endpoints | Probes & Pod Lifecycle |
Here is the same path as a runnable toy. It models the documented order of API server stages and shows that controllers are just API clients whose writes go through the same pipeline as yours.
"""Simulation: the path of `kubectl apply` through a toy control plane.
Not real Kubernetes code. It models the order of stages documented at
kubernetes.io (authn -> authz -> mutating admission -> schema validation ->
validating admission -> etcd), then shows how watches drive each controller.
"""
import copy
STORE = {} # key -> object (our "etcd")
REV = [100] # global revision counter, like etcd's revision
WATCHERS = [] # (kind, callback)
LOG = []
def log(step, msg):
LOG.append(f"{len(LOG)+1:>2}. [{step}] {msg}")
# ---- API server request pipeline -----------------------------------------
def authenticate(req):
if req["token"] != "ajay-token":
raise PermissionError("401 Unauthorized")
return "ajay"
def authorize(user, verb, kind):
allowed = {("ajay", "create", "Deployment"), ("ajay", "update", "Deployment")}
# controllers run as service accounts with their own RBAC
if user.startswith("system:") or (user, verb, kind) in allowed:
return True
raise PermissionError(f"403 Forbidden: {user} cannot {verb} {kind}")
def mutate(obj):
# like a defaulting webhook / MutatingAdmissionPolicy: fill unset fields
spec = obj.setdefault("spec", {})
if obj["kind"] == "Pod":
spec.setdefault("restartPolicy", "Always")
else:
spec.setdefault("replicas", 1)
obj.setdefault("metadata", {}).setdefault("labels", {})["team"] = "payments"
return obj
def schema_validate(obj):
if obj["kind"] != "Pod" and not isinstance(obj["spec"]["replicas"], int):
raise ValueError("400 BadRequest: spec.replicas must be an integer")
def validate(obj):
# like a ValidatingAdmissionPolicy: forbid mutable :latest tags
if obj["kind"] in ("Deployment", "Pod"):
img = obj["spec"]["image"]
if img.endswith(":latest"):
raise ValueError(f"admission denied by policy no-latest-tag: {img}")
def api_write(req, user=None):
obj = copy.deepcopy(req["obj"])
user = user or authenticate(req)
authorize(user, req["verb"], obj["kind"])
mutate(obj); schema_validate(obj); validate(obj)
if user.startswith("system:"):
# controllers are clients too: every write takes the same pipeline
log("api", f"{user}: authn, authz, admission, schema all ok")
else:
log("authn", f"user={user}")
log("authz", f"RBAC allows {user} to {req['verb']} {obj['kind']}")
log("mutating", f"defaults applied, labels={obj['metadata']['labels']}")
log("schema", "OpenAPI schema ok")
log("validating", "policy no-latest-tag ok")
REV[0] += 1
obj["metadata"]["resourceVersion"] = str(REV[0])
key = f"{obj['kind']}/{obj['metadata']['name']}"
STORE[key] = obj
log("etcd", f"stored {key} at resourceVersion {REV[0]}")
for kind, cb in WATCHERS: # watch fan-out (served from watch cache)
if kind == obj["kind"]:
cb(obj)
return obj
# ---- controllers, scheduler and kubelet, each reacting to a watch ---------
def deployment_controller(dep):
log("deploy-ctrl", f"saw {dep['metadata']['name']}, creating ReplicaSet")
rs = {"kind": "ReplicaSet", "metadata": {"name": dep["metadata"]["name"] + "-7d9f"},
"spec": {"replicas": dep["spec"]["replicas"], "image": dep["spec"]["image"]}}
api_write({"verb": "create", "obj": rs}, user="system:deployment-controller")
def replicaset_controller(rs):
have = sum(1 for k in STORE if k.startswith("Pod/" + rs["metadata"]["name"]))
for i in range(have, rs["spec"]["replicas"]):
pod = {"kind": "Pod", "metadata": {"name": f"{rs['metadata']['name']}-{i}"},
"spec": {"image": rs["spec"]["image"], "nodeName": None}}
log("rs-ctrl", f"creating {pod['metadata']['name']} (nodeName empty)")
api_write({"verb": "create", "obj": pod}, user="system:replicaset-controller")
NODES = {"node-a": 1, "node-b": 2} # free slots, toy capacity
def scheduler(pod):
if pod["spec"]["nodeName"]:
return
node = max(NODES, key=NODES.get) # filter (has room) then score (most free)
NODES[node] -= 1
pod["spec"]["nodeName"] = node
log("scheduler", f"bind {pod['metadata']['name']} -> {node}")
REV[0] += 1
pod["metadata"]["resourceVersion"] = str(REV[0])
kubelet(pod) # the kubelet watches pods bound to its node
def kubelet(pod):
log("kubelet", f"{pod['spec']['nodeName']}: CRI pull+start, CNI IP, CSI mounts -> Running")
WATCHERS += [("Deployment", deployment_controller), ("ReplicaSet", replicaset_controller),
("Pod", scheduler)]
dep = {"kind": "Deployment", "metadata": {"name": "checkout"},
"spec": {"replicas": 2, "image": "registry.example/checkout:1.4.2"}}
api_write({"token": "ajay-token", "verb": "create", "obj": dep})
print("\n".join(LOG))
print("\n-- failure path: same apply with a :latest image --")
LOG.clear()
bad = copy.deepcopy(dep); bad["metadata"]["name"] = "checkout-canary"
bad["spec"]["image"] = "registry.example/checkout:latest"
try:
obj = mutate(copy.deepcopy(bad)); log("authn+authz", "ok"); log("mutating", "ok")
schema_validate(obj); log("schema", "ok")
api_write({"token": "ajay-token", "verb": "create", "obj": bad})
except ValueError as e:
print("\n".join(LOG))
print("REJECTED at validating admission:", e)
print("objects in etcd:", sorted(STORE))Output:
1. [authn] user=ajay
2. [authz] RBAC allows ajay to create Deployment
3. [mutating] defaults applied, labels={'team': 'payments'}
4. [schema] OpenAPI schema ok
5. [validating] policy no-latest-tag ok
6. [etcd] stored Deployment/checkout at resourceVersion 101
7. [deploy-ctrl] saw checkout, creating ReplicaSet
8. [api] system:deployment-controller: authn, authz, admission, schema all ok
9. [etcd] stored ReplicaSet/checkout-7d9f at resourceVersion 102
10. [rs-ctrl] creating checkout-7d9f-0 (nodeName empty)
11. [api] system:replicaset-controller: authn, authz, admission, schema all ok
12. [etcd] stored Pod/checkout-7d9f-0 at resourceVersion 103
13. [scheduler] bind checkout-7d9f-0 -> node-b
14. [kubelet] node-b: CRI pull+start, CNI IP, CSI mounts -> Running
15. [rs-ctrl] creating checkout-7d9f-1 (nodeName empty)
16. [api] system:replicaset-controller: authn, authz, admission, schema all ok
17. [etcd] stored Pod/checkout-7d9f-1 at resourceVersion 105
18. [scheduler] bind checkout-7d9f-1 -> node-a
19. [kubelet] node-a: CRI pull+start, CNI IP, CSI mounts -> Running
-- failure path: same apply with a :latest image --
1. [authn+authz] ok
2. [mutating] ok
3. [schema] ok
REJECTED at validating admission: admission denied by policy no-latest-tag: registry.example/checkout:latest
objects in etcd: ['Deployment/checkout', 'Pod/checkout-7d9f-0', 'Pod/checkout-7d9f-1', 'ReplicaSet/checkout-7d9f']Notice the failure path: the :latest request was rejected at validating admission, so nothing reached etcd, and no controller ever learned it existed. Admission is the only place to stop a bad object before the loops start acting on it.
The pieces and their contracts
| Component | Job | Talks to | If it is down |
|---|---|---|---|
| kube-apiserver | The only door to cluster state: authn, authz, admission, validation, storage, watches | etcd, every client | No changes, no watches; running pods keep running |
| etcd | Consistent key-value store for all API objects (Raft replicated) | API server only | API server cannot read or write; cluster frozen in place |
| kube-scheduler | Assigns unbound pods to nodes | API server | New pods stay Pending; existing pods unaffected |
| kube-controller-manager | Runs built-in controllers (Deployment, ReplicaSet, Job, node lifecycle, endpoints, garbage collector, PV binder, attach/detach and more) | API server | No self-healing, no rollouts, no garbage collection |
| cloud-controller-manager (optional) | Cloud-specific loops: nodes, routes, load balancers | API server, cloud APIs | New LoadBalancer Services and node cleanup stall |
| kubelet | Node agent: runs pods bound to its node, reports status, evicts under pressure | API server, CRI, CNI, CSI | Pods on that node keep running but are unmanaged; node goes NotReady |
| Container runtime (CRI) | Pulls images, creates containers (containerd, CRI-O) | kubelet | No new containers on the node |
| CNI plugin | Pod networking and IP allocation | kubelet/runtime | Pods stuck ContainerCreating |
| CSI driver | Provisions, attaches and mounts volumes | Controller sidecars and kubelet | PVCs Pending, pods stuck ContainerCreating |
A useful rule: the data plane keeps running when the control plane fails. Pods already running keep serving traffic during an API server or etcd outage. What you lose is change: no deploys, no rescheduling after node failure, no scaling. That is why a control plane outage is often noticed late, and why static stability matters.
Where the bottlenecks and sharp edges are
| Area | What goes wrong | Root cause | Covered on |
|---|---|---|---|
| etcd size and latency | Slow or failing writes, mvcc: database space exceeded | Huge objects, too many objects, slow disks, no defrag | Control plane page |
| Watch fan-out and LISTs | API server memory spikes, 429s | Controllers that LIST everything instead of using informers | Control plane page |
| Admission webhooks | Every create fails or times out cluster-wide | A webhook with failurePolicy: Fail whose backend is down | Control plane page |
| Scheduling | Pods Pending forever | Over-constrained affinity, taints, zone-pinned volumes | Scheduler page and Storage page |
| Finalizers | Namespace or PVC stuck Terminating | The controller that owns the finalizer is gone | Scheduler/controllers page |
| Storage topology | Pod cannot start after reschedule | Volume pinned to one zone, RWO multi-attach | Storage page |
| Packaging drift | Rollback does not restore state, CRDs never upgrade | Merge semantics, Helm CRD rules | Helm page |
| Node scaling latency | Spike traffic waits minutes for capacity | VM boot time dominates | Node scaling page |
Decision chart: which layer do I reach for?
Decisions
- 1
Q1
- nextExisting Workloads series: Deployments, probes, HPA, rollouts
- nextStep 2: a pod is Pending?
- 2
Existing Workloads series: Deployments, probes, HPA, rollouts
- ?
Step 2: a pod is Pending?
- nextScheduler page, then Node scaling page if capacity is the gap
- nextStorage page: binding modes and zones
- nextStep 3: a write or deploy fails before any pod exists?
- 4
Scheduler page, then Node scaling page if capacity is the gap
- 5
Storage page: binding modes and zones
- ?
Step 3: a write or deploy fails before any pod exists?
- nextControl plane page: RBAC, admission, optimistic concurrency
- nextStep 4: the problem is packaging, versioning or drift?
- 7
Control plane page: RBAC, admission, optimistic concurrency
- ?
Step 4: the problem is packaging, versioning or drift?
- nextHelm page: values, releases, 3-way merge, GitOps
- nextControllers page: finalizers, owner references, garbage collection
- 9
Helm page: values, releases, 3-way merge, GitOps
- 10
Controllers page: finalizers, owner references, garbage collection
What happens if you choose an alternative to Kubernetes
| Alternative | You gain | You lose | Choose it when |
|---|---|---|---|
| A PaaS or serverless containers (Cloud Run, App Runner, ECS Fargate) | No control plane to run, scale-to-zero, simpler mental model | Custom controllers, fine-grained scheduling, portable CRD ecosystem | A handful of stateless services and a small team |
| ECS or Nomad | Smaller surface, simpler scheduler | Huge ecosystem (Helm charts, operators, CSI drivers) | You are deep in one cloud or want a lighter scheduler |
| VMs plus configuration management | Familiar, no container networking layer | Bin-packing, self-healing, fast rollouts | Few large stateful apps, strict compliance on hosts |
| Managed Kubernetes (EKS, GKE, AKS) | Provider runs etcd and the API server | Some control plane knobs; provider limits and upgrade cadence | Most teams that want Kubernetes |
| Self-managed Kubernetes | Every knob | You own etcd backups, upgrades, certificates | Air-gapped, edge, or platform teams with the staff |
Pitfalls
- "apply succeeded" is not "deployed". Watch rollout status and pod conditions; the API server only accepted intent.
- Treating the API server like a database for app data. etcd is designed for small metadata objects; big ConfigMaps and high-churn custom resources hurt every client.
- Fighting the reconcilers. Hand edits (
kubectl edit,kubectl scale) are reverted by whoever owns that field: a controller, Helm on the next upgrade, or a GitOps agent within minutes. - Forgetting that controllers are eventually consistent. Scripts that create an object then immediately read a derived object (ReplicaSet, endpoints) will race.
- One namespace, one cluster, everything. Blast radius: a runaway controller or webhook affects every tenant. Split by environment and criticality.
The cluster map
- Kubernetes Control Plane - API Server Request Path, etcd, Watches, Informers, Optimistic Concurrency & CRDs: the API server request path, API Priority and Fairness, admission, etcd limits, resourceVersion, watches and 410 Gone, 409 conflicts, server-side apply, informers, CRDs and operators.
- Scheduler, Controllers & the Kubelet - Filter and Score, Taints, Preemption, Reconcile Loops, Garbage Collection & Pod Sync: filter and score, taints, affinity and spread, preemption, level-triggered reconcile, owner references, finalizers, kubelet pod sync and node-pressure eviction.
- Kubernetes Storage - Volumes, PV, PVC, StorageClass, CSI, StatefulSets & Failure Modes: PV, PVC and StorageClass, binding modes, CSI calls, access and reclaim modes, StatefulSets, expansion, snapshots and the node-loss timeline.
- Helm & Kubernetes Packaging - Charts, Values, Releases, Rollback, Hooks vs Kustomize, Operators & GitOps: charts, values precedence, releases and rollback, the 3-way merge, hooks and CRDs, and Helm vs Kustomize vs operators vs GitOps.
- Node-Level Scaling - Node Groups, Cluster Autoscaler vs Karpenter, Over-Provisioning, PDBs & Drains: node groups, Cluster Autoscaler vs Karpenter, consolidation, the scale-up latency budget, over-provisioning, PDBs, drains and cost levers.
Interview Q&A
Walk me through what happens when you run `kubectl apply -f deployment.yaml`.
Answer
kubectl sends the object to the API server. The request is authenticated, authorized by RBAC, passed through mutating admission (defaults, injected fields), validated against the schema and validating admission policies, then written to etcd with a new resourceVersion. The Deployment controller, watching Deployments, creates a ReplicaSet; the ReplicaSet controller creates Pods; the scheduler binds each Pod to a node; the kubelet on that node starts containers via the CRI, gets networking from CNI and volumes from CSI; probes mark the Pod Ready and the endpoints controller adds it to Services.
Why is Kubernetes called level-triggered or declarative?
Answer
Controllers act on the current difference between desired and observed state, not on the history of events. If an event is missed or a controller restarts, the next reconcile still computes the right action from current state, which makes the system self-healing and retries idempotent.
What keeps running if etcd goes down?
Answer
Running containers keep running and serving. You lose all changes: no scheduling, no scaling, no rollouts, no recovery from node failure, and kubectl reads may fail. That is why etcd backups and quorum health are critical.
Where would you enforce "no `:latest` images" and why there?
Answer
Validating admission (a ValidatingAdmissionPolicy with CEL or a policy engine webhook), because it blocks the object before it is persisted, so no controller acts on it. Checking later, for example in a controller, means pods may already be running.
Why does the scheduler not start the pod itself?
Answer
Separation of concerns and failure isolation. The scheduler only writes a binding; the kubelet that owns the node runs containers. Each component can fail or be replaced independently, and all coordination goes through the API server.
What is the difference between spec and status?
Answer
spec is the desired state you write. status is the observed state that controllers and the kubelet write back.
Why do Kubernetes components never call each other directly?
Answer
Each component watches the API server and writes back through it, so any component can restart or be replaced without the others noticing, and missed calls cannot lose work.
Why is the system eventually consistent?
Answer
Each loop reacts to watch events on its own schedule, so a derived object such as a ReplicaSet or an endpoint appears a little after the object that caused it.
Check yourself
Run kubectl apply on a small Deployment and watch kubectl get events --watch in another terminal. Label each event with the component that produced it and the step number from Diagram 1.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Kubernetes Workloads — Deployments, Probes, Resources & Progressive Delivery, Pods, ReplicaSets & Deployments — Desired State & Controllers, HPA, VPA & Autoscaling Gotchas — Metrics, Stabilization & Thrash, Ingress Controllers & North-South — K8s Ingress/Gateway API, L7 vs L4, TLS, Infrastructure as Code — Terraform State, Modules & Safe Change, CI/CD Pipelines — Stages, Artifacts, Caching & Supply Chain, Build Systems & Monorepos - Graphs, Caching, Hermeticity & Reproducible Environments, Disaster Recovery & Multi-Region - RTO/RPO, Backups, Pilot Light to Active-Active, Raft Consensus — Leader Election, Log Replication & Safety.