Scheduler, Controllers & the Kubelet - Filter and Score, Taints, Preemption, Reconcile Loops, Garbage Collection & Pod Sync
Scheduler, controllers and kubelet: scheduling framework (PreFilter, Filter, Score, Reserve, Permit, Bind), taints/tolerations, affinity and topology spread, priority and preemption; level-triggered reconcile loops, owner references, finalizers and garbage collection; kubelet pod sync, node-pressure eviction and a Pending/ContainerCreating triage chart.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What are the two phases of choosing a node?
Answer
Filter removes nodes that cannot run the pod; Score ranks the rest.
L2
Taint vs node affinity?
Answer
A taint repels pods that do not tolerate it. Affinity attracts a pod to labeled nodes. Dedicated pools need both.
L3
When does preemption run?
Answer
Only when no node is feasible. It evicts lower-priority pods and sets nominatedNodeName.
L4
What is level-triggered reconciliation?
Answer
Act on the current difference between desired and actual state, not on individual events.
L5
Why is a namespace stuck Terminating?
Answer
A finalizer that no running controller will remove.
L6
Does node-pressure eviction respect PDBs?
Answer
No. PDBs only govern voluntary evictions through the Eviction API.
L7
How long until pods leave an unreachable node by default?
Answer
About 50 seconds to mark it unhealthy plus a 300-second NoExecute toleration.
Failure modes
Pods Pending forever
Over-constrained affinity, untolerated taints or DoNotSchedule spread with a full zone leave no feasible node.
Drift after a missed event
An edge-triggered controller keeps a private count and never repairs a pod lost during a watch reconnect.
Objects stuck Terminating
The controller that owns a finalizer was removed or is crash-looping.
Misconceptions
A toleration attracts a pod to a node.
It only allows it. Pair it with node affinity.
Affinity rules move running pods when labels change.
IgnoredDuringExecution means they are checked only at scheduling time.
The scheduler packs by real usage.
It packs by requests, so zero-request pods all land together.
Interviewer traps
Setting spec.nodeName to skip the scheduler.
It bypasses filters, including taints and resource checks. Use affinity or a scheduler profile.
Removing finalizers by hand as the default fix.
It skips the cleanup they protect. Fix the owning controller first.
Design scenario
Same prompt for every reader.
Requirements
Payments must spread across zones and win capacity under pressure; GPU nodes only run GPU work; batch work fills spare capacity.
Failure assumptions
- A zone loses capacity.
- A node becomes unreachable.
- A batch job hits node memory pressure.
Constraints
- No custom scheduler.
- Node autoscaling exists but takes minutes.
Prompt
Design scheduling for a cluster that mixes a payments API, GPU training jobs and low-priority batch work across three zones.
API
Which PriorityClasses, taints, tolerations, affinity and spread constraints do the specs carry?
Data
What scheduler events and pod conditions tell you why a pod is Pending?
Architecture
How do the scheduler, preemption, the node controller and the kubelet interact when capacity runs out?
Overview
Three kinds of loops do the actual work in a cluster. The scheduler takes pods with no node and picks one: first it filters out nodes that cannot run the pod, then it scores the survivors, and if nothing fits it may preempt lower-priority pods. Controllers compare desired state with observed state and issue API writes to close the gap, again and again, without trusting that they saw every event. The kubelet on each node does the same for "pods bound to me": it makes the container runtime, network plugin and volume plugin match each pod spec, reports status, and evicts pods when the node runs out of memory or disk.
The connecting idea is level-triggered reconciliation: act on what is true now, not on the event that woke you. This page shows how each loop works, how they cooperate through the API server, and where they go wrong. Requests, limits and QoS classes themselves are taught in Requests, Limits & QoS; here we use them as scheduler and kubelet inputs.
The scheduler: filter, score, bind
The kube-scheduler is built on the scheduling framework (stable since v1.19). Each pod attempt has a scheduling cycle (pick a node; cycles run one pod at a time) and a binding cycle (persist the decision; binding cycles can run concurrently).
Flow
- 1
Step 1: pod with empty nodeName enters the active queue (QueueSort by priority)
- nextStep 2: PreFilter computes pod-wide state (requests, affinity terms, spread counts)
- 2
Step 2: PreFilter computes pod-wide state (requests, affinity terms, spread counts)
- nextStep 3: Filter runs per node: resources fit, taints, node affinity, pod anti-affinity, volume topology, ports
- 3
Step 3: Filter runs per node: resources fit, taints, node affinity, pod anti-affinity, volume topology, ports
- nextStep 4: PreScore and Score rank feasible nodes, NormalizeScore, then weighted sum
- nextFailure path: PostFilter tries preemption of lower-priority pods
- 4
Step 4: PreScore and Score rank feasible nodes, NormalizeScore, then weighted sum
- nextStep 5: Reserve the node in the scheduler cache, Permit can approve, deny or wait
- 5
Step 5: Reserve the node in the scheduler cache, Permit can approve, deny or wait
- nextStep 6: PreBind (for example volume binding), then Bind writes the Binding to the API server
- 6
Step 6: PreBind (for example volume binding), then Bind writes the Binding to the API server
- nextStep 7: kubelet on that node watches and starts the pod
- 7
Step 7: kubelet on that node watches and starts the pod
- 8
Failure path: PostFilter tries preemption of lower-priority pods
- nextVictims get graceful termination, pod gets nominatedNodeName and retries
- nextPod marked Unschedulable and parked until a cluster event may help; node autoscaler may add capacity
- 9
Victims get graceful termination, pod gets nominatedNodeName and retries
- 10
Pod marked Unschedulable and parked until a cluster event may help; node autoscaler may add capacity
Lesson map
Scheduler, Controllers & the Kubelet - Filter and Score, Taints, Preemption, Reconcile Loops, Garbage Collection & Pod Sync
Scheduler, controllers and kubelet: scheduling framework (PreFilter, Filter, Score, Reserve, Permit, Bind), taints/tolerations, affinity and topology spread, priority and preemption; level-triggered reconcile loops, owner references, finalizers and garbage collection; kubelet pod sync, node-pressure eviction and a Pending/ContainerCreating triage chart.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB q["Step 1: pod with empty nodeName enters the active queue (QueueSort by priority)"] pf["Step 2: PreFilter computes pod-wide state (requests, affinity terms, spread counts)"] f["Step 3: Filter runs per node: resources fit, taints, node affinity, pod anti-affinity, volume topology, ports"] sc["Step 4: PreScore and Score rank feasible nodes, NormalizeScore, then weighted sum"] rs["Step 5: Reserve the node in the scheduler cache, Permit can approve, deny or wait"] pb["Step 6: PreBind (for example volume binding), then Bind writes the Binding to the API server"] k["Step 7: kubelet on that node watches and starts the pod"] po["Failure path: PostFilter tries preemption of lower-priority pods"] nn["Victims get graceful termination, pod gets nominatedNodeName and retries"] un["Pod marked Unschedulable and parked until a cluster event may help node autoscaler may add capacity"] q -->|continues| pf pf -->|continues| f f -->|continues| sc sc -->|continues| rs rs -->|continues| pb pb -->|continues| k f -->|continues| po po -->|continues| nn po -->|continues| un
| Extension point | Typical in-tree plugins | Purpose |
|---|---|---|
| QueueSort | PrioritySort | Higher-priority pods are tried first |
| PreFilter / Filter | NodeResourcesFit, NodeAffinity, TaintToleration, InterPodAffinity, PodTopologySpread, VolumeBinding, VolumeZone, NodePorts | Remove nodes that cannot run the pod |
| PostFilter | DefaultPreemption | Runs only when no node is feasible |
| PreScore / Score | NodeResourcesFit (LeastAllocated by default, or MostAllocated), NodeResourcesBalancedAllocation, InterPodAffinity, PodTopologySpread, ImageLocality, TaintToleration | Rank feasible nodes |
| Reserve / Permit | VolumeBinding; gang-scheduling style plugins use Permit | Hold resources, delay binding |
| PreBind / Bind | VolumeBinding, DefaultBinder | Provision or bind volumes, then write the Binding |
How should replicas spread across nodes and zones?
Prefer
Topology spread constraints
Score nodes by skew so replicas even out across zones and hosts.
- The spread score sent api-3 to zone b (n3=233 vs n2=112).
- maxSkew keeps zones even without capping replicas per node.
- ScheduleAnyway turns it into a preference when capacity is tight.
Alternative
Required anti-affinity on hostname
Forbid two replicas on one node as a hard filter.
- It removed n1 outright for api-3.
- Replicas are capped at the node count.
- It is the slowest predicate and stops the autoscaler from packing.
From an unbound pod to a running container
Diagram 1 condensed: the scheduling cycle, binding and the kubelet's sync.
- 1
Queue by priority
PrioritySort tries higher-priority pods first. - 2
Filter then score
Drop nodes that fail resources, taints, affinity or volume checks, then rank the rest. - 3
Preempt if nothing fits
Evict lower-priority victims and record nominatedNodeName. - 4
Bind and sync
The Binding is written and the kubelet sets up volumes, sandbox and containers. - 5
Stay Pending
No victims and no feasible node: only a node autoscaler can help.
Big clusters do not score every node. percentageOfNodesToScore stops the feasibility search once enough nodes are found. If unset, Kubernetes uses a linear formula that gives 50% for a 100-node cluster and 10% for 5,000 nodes, never below 5%, and always checks at least 100 nodes. Nodes are visited round-robin across zones so every zone gets considered.
Constraints: who can land where
| Mechanism | Declared on | Hard or soft | Typical use | Common mistake |
|---|---|---|---|---|
nodeSelector / node affinity | Pod | Required or preferred | "Only on arm64" or "only on GPU nodes" | Nobody labels the new node group |
| Taints and tolerations | Node (taint), pod (toleration) | NoSchedule, PreferNoSchedule, NoExecute | Dedicated or special nodes repel everyone else | A toleration lets a pod in, it does not attract it; pair with affinity |
| Pod affinity / anti-affinity | Pod | Required or preferred, by topologyKey | Co-locate with a cache; keep replicas on different hosts | Required anti-affinity on hostname caps replicas at node count; it is also the slowest predicate |
| Topology spread constraints | Pod | whenUnsatisfiable: DoNotSchedule or ScheduleAnyway, maxSkew | Even spread across zones and nodes | DoNotSchedule plus a zone with no capacity leaves pods Pending |
| PriorityClass | Pod | Ordering plus preemption | Protect critical services | Everything marked high priority, so nothing is |
IgnoredDuringExecution in the affinity field names matters: constraints are checked at scheduling time only. A running pod is not moved if labels change later. NoExecute taints are the exception: pods that do not tolerate them are evicted, and pods that tolerate them with tolerationSeconds stay only that long. That is how node failure eviction works: the node controller taints an unreachable node with node.kubernetes.io/unreachable:NoExecute, and the DefaultTolerationSeconds admission plugin gives every pod a 300-second toleration unless it sets its own.
Preemption
When no node fits, DefaultPreemption looks for nodes where evicting lower-priority pods would make room, prefers candidates whose victims matter least, gives victims their graceful termination period, and records nominatedNodeName on the waiting pod. PodDisruptionBudgets are respected best effort only: if no PDB-respecting victim set exists, preemption may still go ahead. The pod is not guaranteed the nominated node; if another node frees up first, it may land there.
"""Toy scheduler modelled on the kube-scheduler framework: Filter, then Score,
then (only if nothing fits) PostFilter preemption. Plugin names mirror real
in-tree plugins, but the logic is simplified for teaching.
"""
NODES = {
# name zone cpu_free(m) mem_free(Mi) taints pods (name, app, priority, cpu)
"n1": dict(zone="a", cpu=2000, mem=6000, taints=[], pods=[("api-1", "api", 0, 500)]),
"n2": dict(zone="a", cpu=1500, mem=3000, taints=[], pods=[("batch-1", "batch", -10, 1500)]),
"n3": dict(zone="b", cpu=3200, mem=8000, taints=[], pods=[]),
"n4": dict(zone="b", cpu=8000, mem=30000, taints=[("gpu", "NoSchedule")], pods=[]),
"n5": dict(zone="c", cpu=400, mem=1000, taints=[], pods=[("api-2", "api", 0, 500)]),
}
def f_resources(pod, n): # NodeResourcesFit
return (pod["cpu"] <= n["cpu"] and pod["mem"] <= n["mem"]) or "Insufficient cpu/memory"
def f_taints(pod, n): # TaintToleration (filter part)
for key, eff in n["taints"]:
if eff == "NoSchedule" and key not in pod["tolerates"]:
return f"untolerated taint {key}:{eff}"
return True
def f_anti_affinity(pod, n): # InterPodAffinity, required anti-affinity on hostname
if pod.get("anti_affinity_app") and any(a == pod["anti_affinity_app"] for _, a, _, _ in n["pods"]):
return "anti-affinity: same app already on node"
return True
FILTERS = [f_resources, f_taints, f_anti_affinity]
def s_least_allocated(pod, n): # NodeResourcesFit scoring with LeastAllocated: prefer free nodes
return int(100 * (n["cpu"] - pod["cpu"]) / 8000)
def s_spread(pod, n, feasible): # PodTopologySpread (soft): prefer zones with fewer matching pods
count = {z: 0 for z in {NODES[x]["zone"] for x in NODES}}
for nn in NODES.values():
count[nn["zone"]] += sum(1 for _, a, _, _ in nn["pods"] if a == pod["app"])
return 100 - 50 * count[n["zone"]]
def schedule(pod):
print(f"\nPOD {pod['name']} cpu={pod['cpu']}m priority={pod['priority']}")
feasible = []
for name, n in NODES.items():
reason = next((r for f in FILTERS if (r := f(pod, n)) is not True), None)
print(f" filter {name}: {'ok' if reason is None else 'X ' + reason}")
if reason is None:
feasible.append(name)
if feasible:
scored = {x: s_least_allocated(pod, NODES[x]) + 2 * s_spread(pod, NODES[x], feasible) for x in feasible}
print(" score " + ", ".join(f"{k}={v}" for k, v in scored.items()) + " (spread weight 2)")
best = max(scored, key=scored.get)
bind(pod, best); return best
return preempt(pod)
def preempt(pod): # DefaultPreemption (PostFilter): evict lower-priority pods
candidates = []
for name, n in NODES.items():
if f_taints(pod, n) is not True:
continue
victims = sorted([p for p in n["pods"] if p[2] < pod["priority"]], key=lambda p: p[2])
freed, chosen = 0, []
for v in victims: # remove the lowest-priority pods first
freed += v[3]; chosen.append(v)
if pod["cpu"] <= n["cpu"] + freed:
# prefer the node whose most important victim is least important, then fewest victims
candidates.append((max(c[2] for c in chosen), len(chosen), name, chosen))
break
if not candidates:
print(" postfilter: no victims found -> Pending (Unschedulable); a node autoscaler may react")
return None
for c in sorted(candidates):
print(f" postfilter candidate {c[2]}: victims {[v[0] for v in c[3]]} (max victim priority {c[0]})")
_, _, name, chosen = min(candidates)
n = NODES[name]
for c in chosen:
n["pods"].remove(c); n["cpu"] += c[3]
print(f" preempt {[c[0] for c in chosen]} (graceful termination), nominatedNodeName={name}")
bind(pod, name); return name
def bind(pod, node):
n = NODES[node]; n["cpu"] -= pod["cpu"]; n["mem"] -= pod["mem"]
n["pods"].append((pod["name"], pod["app"], pod["priority"], pod["cpu"]))
print(f" bind -> {node} (zone {n['zone']})")
schedule(dict(name="api-3", app="api", cpu=500, mem=512, priority=0, tolerates=[], anti_affinity_app="api"))
schedule(dict(name="trainer", app="ml", cpu=6000, mem=16000, priority=0, tolerates=["gpu"]))
schedule(dict(name="payments", app="pay", cpu=2800, mem=2000, priority=1000, tolerates=[]))
schedule(dict(name="report", app="rep", cpu=7000, mem=1000, priority=0, tolerates=[]))Output:
POD api-3 cpu=500m priority=0
filter n1: X anti-affinity: same app already on node
filter n2: ok
filter n3: ok
filter n4: X untolerated taint gpu:NoSchedule
filter n5: X Insufficient cpu/memory
score n2=112, n3=233 (spread weight 2)
bind -> n3 (zone b)
POD trainer cpu=6000m priority=0
filter n1: X Insufficient cpu/memory
filter n2: X Insufficient cpu/memory
filter n3: X Insufficient cpu/memory
filter n4: ok
filter n5: X Insufficient cpu/memory
score n4=225 (spread weight 2)
bind -> n4 (zone b)
POD payments cpu=2800m priority=1000
filter n1: X Insufficient cpu/memory
filter n2: X Insufficient cpu/memory
filter n3: X Insufficient cpu/memory
filter n4: X Insufficient cpu/memory
filter n5: X Insufficient cpu/memory
postfilter candidate n2: victims ['batch-1'] (max victim priority -10)
postfilter candidate n3: victims ['api-3'] (max victim priority 0)
preempt ['batch-1'] (graceful termination), nominatedNodeName=n2
bind -> n2 (zone a)
POD report cpu=7000m priority=0
filter n1: X Insufficient cpu/memory
filter n2: X Insufficient cpu/memory
filter n3: X Insufficient cpu/memory
filter n4: X Insufficient cpu/memory
filter n5: X Insufficient cpu/memory
postfilter: no victims found -> Pending (Unschedulable); a node autoscaler may reactRead the output like a scheduler event log. api-3 was kept off n1 by anti-affinity and pushed to zone b by the spread score. trainer only fit on the GPU node because it tolerated the taint. payments preempted the -10 priority batch pod rather than the 0 priority API pod. report asked for more than any node has, so only a node autoscaler (see the node scaling page) can help.
Controllers: reconcile, don't react
Every built-in controller in kube-controller-manager follows the same shape: informers feed a work queue of keys; a worker pops a key, reads current state from the cache, computes the difference from desired state, and writes changes through the API server. Replicated control-plane components use a Lease object for leader election (--leader-elect-lease-duration defaults to 15s), so only one active instance of each controller acts at a time.
| Edge-triggered handler | Level-triggered reconcile | |
|---|---|---|
| Input | The event: "pod deleted" | The key: "ReplicaSet web, go look" |
| State | Private counters updated per event | Recomputed from the cache every time |
| Missed event (watch reconnect, controller restart) | Permanent drift | Fixed on the next reconcile or periodic resync |
| Duplicate event | Double action | No-op |
| Testing | Need exact event sequences | Give it a state, check the writes |
// Simulation: a mini ReplicaSet-style controller.
// Part 1: edge-triggered (react to each event's delta) vs level-triggered
// (look at current state every time) when one event is dropped.
// Part 2: owner references, garbage collection and a finalizer.
type Obj = { kind: string; name: string; owner?: string; finalizers: string[]; deleting: boolean };
const store = new Map<string, Obj>();
let seq = 0;
const create = (kind: string, owner?: string): Obj => {
const o = { kind, name: `${kind.toLowerCase()}-${seq++}`, owner, finalizers: [], deleting: false };
store.set(o.name, o); return o;
};
const podsOf = (owner: string) => [...store.values()].filter((o) => o.kind === "Pod" && o.owner === owner && !o.deleting);
// ---------- Part 1 ----------
const rs = create("ReplicaSet");
const desired = 3;
let edgeBelief = 0; // edge-triggered controller's private counter
type Ev = { type: "PodAdded" | "PodDeleted"; delivered: boolean };
function edgeHandle(ev: Ev) { // trusts the event stream to be complete
if (!ev.delivered) return;
edgeBelief += ev.type === "PodAdded" ? 1 : -1;
}
function edgeAct() { while (edgeBelief < desired) { create("Pod", rs.name); edgeBelief++; } }
function levelReconcile(): string { // ignores event content, reads current state
const have = podsOf(rs.name).length;
for (let i = have; i < desired; i++) create("Pod", rs.name);
for (const p of podsOf(rs.name).slice(desired)) store.delete(p.name);
return `have ${have} -> ${podsOf(rs.name).length}`;
}
edgeAct();
console.log("edge: start ->", podsOf(rs.name).length, "pods");
const victims = podsOf(rs.name).slice(0, 2);
victims.forEach((p) => store.delete(p.name)); // a node dies, 2 pods gone
edgeHandle({ type: "PodDeleted", delivered: true });
edgeHandle({ type: "PodDeleted", delivered: false }); // watch reconnect: one event lost
edgeAct();
console.log("edge: after lost event ->", podsOf(rs.name).length, "pods, believes", edgeBelief, "(drift persists)");
console.log("level: reconcile ->", levelReconcile(), "(any trigger, even a periodic resync, repairs it)");
// ---------- Part 2 ----------
console.log("\nownerReferences + garbage collection");
const dep = create("Deployment");
const rs2 = create("ReplicaSet", dep.name);
const pods = [create("Pod", rs2.name), create("Pod", rs2.name)];
pods[0].finalizers.push("example.com/flush-logs"); // a controller must clean up first
function deleteObj(name: string, propagation: "Background" | "Orphan") {
const o = store.get(name)!;
if (propagation === "Orphan") { // dependents survive, ownerRef removed
[...store.values()].filter((d) => d.owner === name).forEach((d) => (d.owner = undefined));
}
o.deleting = true; // deletionTimestamp set
if (o.finalizers.length === 0) store.delete(name); // removed only when finalizers are empty
if (propagation === "Background") gc();
}
function gc() { // GC deletes objects whose owner is gone
let changed = true;
while (changed) {
changed = false;
for (const o of [...store.values()]) {
if (o.owner && !store.has(o.owner) && !o.deleting) {
o.deleting = true; changed = true;
if (o.finalizers.length === 0) store.delete(o.name);
}
}
}
}
const show = () => [...store.values()].filter((o) => o.owner === rs2.name || o.name === rs2.name || o.name === dep.name)
.map((o) => `${o.name}${o.deleting ? "(Terminating, finalizers=" + o.finalizers.join(",") + ")" : ""}`).join(" ");
deleteObj(dep.name, "Background");
console.log("after delete Deployment (Background):", show() || "(all gone)");
pods[0].finalizers = []; // log-flusher finishes, removes its finalizer
if (pods[0].finalizers.length === 0 && pods[0].deleting) store.delete(pods[0].name);
console.log("after finalizer removed:", show() || "(all gone)");
const dep2 = create("Deployment"); const rs3 = create("ReplicaSet", dep2.name); create("Pod", rs3.name);
deleteObj(dep2.name, "Orphan");
console.log("delete with --cascade=orphan keeps:", [...store.values()].filter((o) => o.name === rs3.name || o.owner === rs3.name).map((o) => o.name).join(", "));Output:
edge: start -> 3 pods
edge: after lost event -> 2 pods, believes 3 (drift persists)
level: reconcile -> have 2 -> 3 (any trigger, even a periodic resync, repairs it)
ownerReferences + garbage collection
after delete Deployment (Background): pod-8(Terminating, finalizers=example.com/flush-logs)
after finalizer removed: (all gone)
delete with --cascade=orphan keeps: replicaset-11, pod-12Expectededge: start -> 3 pods edge: after lost event -> 2 pods, believes 3 (drift persists) level: reconcile -> have 2 -> 3 (any trigger, even a periodic resync, repairs it) ownerReferences + garbage collection after delete Deployment (Background): pod-8(Terminating, finalizers=example.com/flush-logs) after finalizer removed: (all gone) delete with --cascade=orphan keeps: replicaset-11, pod-12
Press Run. Snippets must be self-contained — no network, files, or native modules.
Owner references, garbage collection and finalizers
- ownerReferences link a dependent to its owner (Pod to ReplicaSet to Deployment). The garbage collector controller deletes dependents whose owner is gone.
- Cascading deletion has three modes: background (default; owner deleted immediately, GC removes dependents afterwards), foreground (owner stays with a deletionTimestamp and a
foregroundDeletionfinalizer until blocking dependents are gone), and orphan (dependents kept, ownerReferences removed;kubectl delete --cascade=orphan). - Finalizers are strings in
metadata.finalizers. A delete request only setsdeletionTimestamp; the object is removed from etcd when the list is empty. Controllers use them to clean up external resources first (a cloud load balancer, a bucket, a volume). Order among finalizers is deliberately not enforced, to avoid deadlocks.
The famous failure: a namespace or PVC stuck Terminating because the controller that owns a finalizer was uninstalled or is crash-looping. Removing the finalizer by hand works, but it skips the cleanup the finalizer was protecting, which can leak cloud resources or data.
The kubelet: the node's reconcile loop
The kubelet watches pods bound to its node (plus static pods from local manifests) and runs a per-pod sync: compute the desired containers, compare with what the runtime reports, and act.
| Step | Interface | What happens |
|---|---|---|
| Admit | kubelet | Checks resources, node features, and that the pod can run here |
| Volumes | CSI node plugin (and in-tree for a few types) | Waits for attach, then stage and publish mounts |
| Sandbox | CRI RunPodSandbox | Runtime creates the pod sandbox; the runtime invokes the CNI plugin, which assigns the IP |
| Containers | CRI image and runtime services | Pull images, create and start init then app containers |
| Resources | cgroups (v2 on modern distros) | Requests and limits become cgroup settings, grouped by QoS class |
| Health | Probes | Liveness restarts containers; readiness gates traffic |
| Status | API server | Pod phase and conditions; node status plus a Lease in kube-node-lease as heartbeat |
Heartbeats and node failure. Each kubelet renews a Lease object for its node. If the node controller sees no heartbeat for --node-monitor-grace-period (default 50s), it marks the node unhealthy and applies the NoExecute taints that start the 300-second eviction clock described above.
Node-pressure eviction. When the node runs short, the kubelet evicts pods itself, and it does not respect PodDisruptionBudgets or, for hard thresholds, termination grace periods. Default hard thresholds on Linux are memory.available<100Mi, nodefs.available<10%, imagefs.available<15% and nodefs.inodesFree<5%. It first tries to reclaim node resources (for example, deleting unused images), then evicts pods ranked by whether their usage exceeds their requests, then by priority, then by how far usage exceeds requests. BestEffort and over-request Burstable pods go first; Guaranteed pods and pods under their requests go last. That is one more reason requests matter, as covered in Requests, Limits & QoS. The Linux memory side is in page faults, swapping and thrashing.
Decision chart: why is my pod not running?
Decisions
- 1
A
- nextStep 2: what does the FailedScheduling event say?
- nextStep 3: what is the container state?
- ?
Step 2: what does the FailedScheduling event say?
- nextLower requests, add capacity, or let a node autoscaler react
- nextAdd a toleration plus node affinity, or use another node group
- nextRelax required rules to preferred, or ScheduleAnyway
- nextZone-pinned PV: see the storage page
- 3
Lower requests, add capacity, or let a node autoscaler react
- 4
Add a toleration plus node affinity, or use another node group
- 5
Relax required rules to preferred, or ScheduleAnyway
- 6
Zone-pinned PV: see the storage page
- ?
Step 3: what is the container state?
- nextCheck CNI IP allocation and volume attach or mount events
- nextRegistry auth, tag, or node egress
- nextApp or liveness probe problem: see Probes and Pod Lifecycle
- nextNode pressure: requests too low or node too small
- 8
Check CNI IP allocation and volume attach or mount events
- 9
Registry auth, tag, or node egress
- 10
App or liveness probe problem: see Probes and Pod Lifecycle
- 11
Node pressure: requests too low or node too small
What happens if you choose an alternative
| Choice | Instead of | Consequence |
|---|---|---|
| Required pod anti-affinity on hostname | Topology spread with maxSkew | Hard cap of one replica per node, slower scheduling, autoscaler cannot pack |
Set spec.nodeName directly | Let the scheduler bind | Bypasses filters, including NoSchedule taints and resource checks |
| A second custom scheduler | Scheduler profiles or plugins | Two schedulers can race for the same node resources |
| Edge-triggered controller | Level-triggered reconcile | Drift after any missed event or restart |
| Manually removing finalizers | Fixing the owning controller | Leaked external resources or lost cleanup |
| No PriorityClasses | A small set of classes | Critical pods cannot preempt batch work when capacity is tight |
Pitfalls
- Requests of zero. The scheduler packs pods by requests, not usage. Zero-request pods all land on one node and get evicted first under pressure.
- Spread constraints without matching labels:
labelSelectormust match the pods you want counted, or the spread is computed over nothing. - Over-tolerant DaemonSets that tolerate every taint can keep a node busy even after you drain it for maintenance.
- Finalizers on high-churn objects slow deletes and pile up when the controller lags.
- Assuming the node controller evicts quickly. With defaults, pods on an unreachable node are evicted after roughly 50 seconds plus a 300-second toleration.
Interview Q&A
How does the scheduler choose a node?
Answer
For each pod it runs Filter plugins to drop infeasible nodes (resources, taints, affinity, volume topology), runs Score plugins on the survivors, normalizes and weights the scores, reserves the best node, and binds the pod with a Binding write. If no node is feasible, PostFilter preemption may evict lower-priority pods.
Taints vs node affinity?
Answer
Taints repel: a tainted node accepts only pods that tolerate the taint. Node affinity attracts: a pod requires or prefers nodes with certain labels. Dedicated node pools usually need both, so outsiders stay off and the intended pods actually go there.
What does level-triggered mean and why does Kubernetes use it?
Answer
The controller acts on the current difference between desired and actual state, not on individual events. It tolerates missed, duplicated or reordered events, controller restarts and partial failures, because every reconcile recomputes the right action.
Why is a namespace stuck Terminating?
Answer
Some object in it, or the namespace itself, has a finalizer that no controller is removing, often because the controller or the API service it depends on was deleted. Fix the controller or API service; remove finalizers by hand only if you accept skipping their cleanup.
Does the kubelet respect PDBs when it evicts?
Answer
No. Node-pressure eviction ignores PodDisruptionBudgets, and with hard thresholds it uses a zero grace period. PDBs only govern voluntary evictions through the Eviction API, such as drains.
What is preemption's relationship with PodDisruptionBudgets?
Answer
The scheduler tries to pick victims whose PDBs would not be violated, but this is best effort; if no such set exists, it may preempt anyway.
What does percentageOfNodesToScore do?
Answer
It stops the feasibility search once enough nodes are found, so big clusters do not score every node. By default it scales from 50% at 100 nodes to 10% at 5,000 nodes.
What are the three cascading deletion modes?
Answer
Background deletes the owner first and GC removes dependents, foreground keeps the owner until blocking dependents are gone, and orphan keeps dependents and strips their ownerReferences.
Check yourself
Create a pod that requests more CPU than any node has, then run kubectl describe pod and map each line of the FailedScheduling event to a Filter plugin from the table.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Pods, ReplicaSets & Deployments — Desired State & Controllers, Probes & Pod Lifecycle — Liveness, Readiness, Startup & PreStop, Requests, Limits & QoS — CPU Throttling, Memory OOM & Scheduling, Lease Renewal, Clock Skew & Heartbeat Failure Modes, Faults, Swapping, and Thrashing — Working Set and the Clock, Sidecar Pattern & Service Mesh — Out-of-Process Proxies, Architecture & Tradeoffs, Plan, Apply, Drift Detection & Import.