Node-Level Scaling - Node Groups, Cluster Autoscaler vs Karpenter, Over-Provisioning, PDBs & Drains
Node scaling: node groups/ASGs/managed node pools, Cluster Autoscaler simulation and expanders, Karpenter NodePools, consolidation and disruption budgets, CA vs Karpenter, scale-up latency budget, over-provisioning with pause pods and CapacityBuffer, PDBs and drains, cost levers. HPA/VPA linked, not re-taught.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What triggers a node scale-up?
Answer
Pods the scheduler marked unschedulable, not CPU usage.
L2
When does Cluster Autoscaler remove a node?
Answer
Requests below 50% of allocatable, all pods movable, unneeded for 10 minutes.
L3
How is Karpenter different?
Answer
It is groupless: it chooses instance types per batch of pods, launches machines directly and consolidates continuously.
L4
Why does scale-up take minutes?
Answer
VM boot and node registration dominate, plus image pulls and readiness.
L5
How do pause pods help?
Answer
Low-priority placeholders hold warm capacity; real pods preempt them instantly and the evicted placeholders trigger a node add.
L6
What do PDBs protect against?
Answer
Voluntary disruptions through the Eviction API, such as drains, scale-down and consolidation.
L7
Why should the scheduler use MostAllocated with Karpenter?
Answer
LeastAllocated spreads pods out, which under-packs nodes and makes consolidation churn.
Failure modes
Pending pods nobody can help
A pod larger than any allowed shape, or with selectors, taints or zones no group can satisfy, stays Pending.
Drains that never finish
PDBs with maxUnavailable 0 or minAvailable equal to replicas block every voluntary eviction.
Consolidation churn
Aggressive consolidation repeatedly restarts long jobs or stateful pods.
Misconceptions
Node autoscalers react to high CPU.
They react to Pending pods and simulate scheduling from requests.
PDBs protect against node crashes.
They only rate-limit voluntary disruptions.
The autoscaler is the slow part.
It reacts in seconds; VM boot takes most of the time.
Interviewer traps
Adding CPU-based ASG policies next to Cluster Autoscaler.
They add unusable nodes or remove nodes running critical pods. Use one pending-pod-aware autoscaler.
Setting maxUnavailable 0 to be safe.
It blocks drains, upgrades and scale-down forever.
Design scenario
Same prompt for every reader.
Requirements
Serve spikes within about 90 seconds, keep cost down off-peak, and upgrade nodes without downtime.
Failure assumptions
- Spot capacity is interrupted.
- Cloud capacity for one instance type runs out.
- A service has no PDB.
Constraints
- AWS EKS.
- Budget for a small warm buffer only.
Prompt
Design node scaling for a cluster whose traffic spikes 5x within two minutes several times a day, with batch jobs at night.
API
Which NodePools, PriorityClasses, PDBs and pause-pod Deployments do you define?
Data
Which metrics show Pending time, buffer usage and consolidation churn?
Architecture
How do HPA, the scheduler, the node autoscaler and the cloud provider interact during a spike and a drain?
Overview
Pod autoscaling (HPA and VPA, covered in HPA, VPA & Autoscaling Gotchas) changes how many pods you want. Node autoscaling changes how many machines exist to run them. The key idea is that a good node autoscaler does not watch CPU graphs. It watches the scheduler: when pods are Pending because nothing fits, it simulates "would a new node let these pods schedule?" and adds capacity. When nodes are underused and their pods could fit elsewhere, it drains them and removes them.
Two designs dominate. Cluster Autoscaler (CA) scales pre-defined node groups: autoscaling groups, managed node groups, node pools or VM scale sets, each with one machine shape. Karpenter is groupless: it reads the pending pods' requirements, picks instance types from a broad allowed set, launches machines directly, and keeps consolidating the fleet toward cheaper packing. Both are bounded by the same physics: a new VM takes minutes, draining must respect PodDisruptionBudgets, and the scheduler decides by requests, not real usage.
Node groups: the substrate
| Cloud construct | What it is | How the autoscaler uses it |
|---|---|---|
| AWS EC2 Auto Scaling group (ASG) | Fleet of instances from a launch template, desired/min/max, across subnets and AZs | CA changes the desired count |
| EKS managed node group | EKS-managed ASG plus lifecycle (AMI updates, graceful drain on update), spanning the AZs you choose | CA changes desired count; EKS handles rolling updates |
| GKE node pool, AKS node pool (VM Scale Set) | Equivalent per-shape groups | Built-in CA in each managed service |
| Karpenter NodePool + NodeClass (EC2NodeClass, AKSNodeClass) | Constraints (allowed instance families, zones, capacity types, limits) plus provider settings (image, subnets, security groups) | Karpenter creates NodeClaims and launches machines directly, no ASG |
Managed offerings now package the groupless model too: EKS Auto Mode runs a Karpenter-based system managed by AWS, and AKS node auto-provisioning deploys Karpenter with the AKS provider.
Cluster Autoscaler or Karpenter for a bursty cluster?
Prefer
Karpenter (groupless)
Pick machines per batch of Pending pods and keep re-packing.
- The burst cost $1.130/h with mixed shapes.
- Consolidation later re-packed to $0.853/h.
- It needs good PDBs because nodes change more often.
Alternative
Cluster Autoscaler (node groups)
Grow the groups someone created earlier, chosen by an expander.
- Least-waste picked 4 x c-8x16 at $1.360/h.
- Scale-down is gated by the 50% threshold and a 10-minute timer.
- At best it shrank to 3 x c-8x16 at $1.020/h.
From a Pending pod to new capacity
Diagram 1 condensed: the scale-up path and where time goes.
- 1
Pod goes Pending
The scheduler marks it unschedulable. - 2
Simulate and choose
The autoscaler simulates template nodes and picks a group or instance type. - 3
Boot the node
VM boot, kubelet registration and CNI take most of the time. - 4
Bind and start
The scheduler binds the pod; image pull and readiness follow. - 5
Nothing can help
No group or NodePool can fit the pod, or the cloud has no capacity.
How Cluster Autoscaler works
Flow
- 1
Step 1: scheduler marks a pod Unschedulable (PodScheduled condition false)
- nextStep 2: every scan-interval (default 10s) CA lists unschedulable pods
- 2
Step 2: every scan-interval (default 10s) CA lists unschedulable pods
- nextStep 3: for each node group CA builds a template node and simulates scheduling the pods onto it
- 3
Step 3: for each node group CA builds a template node and simulates scheduling the pods onto it
- nextStep 4: expander picks a group: least-waste by default; random, most-pods, least-nodes, price, priority, grpc
- nextFailure path: pod stays Pending (too big, wrong labels, untolerated taints, max size or quota)
- 4
Step 4: expander picks a group: least-waste by default; random, most-pods, least-nodes, price, priority, grpc
- nextStep 5: CA raises the group's desired size via the cloud API
- 5
Step 5: CA raises the group's desired size via the cloud API
- nextStep 6: VM boots, kubelet registers, node Ready, scheduler binds the pending pods
- nextFailure path: CA gives up on those nodes, may try another group, removes unregistered instances
- 6
Step 6: VM boots, kubelet registers, node Ready, scheduler binds the pending pods
- nextStep 7: scale-down loop: node requests below 50 percent of allocatable and pods movable for 10 min
- 7
Step 7: scale-down loop: node requests below 50 percent of allocatable and pods movable for 10 min
- nextStep 8: taint node, evict pods through the Eviction API honoring PDBs, terminate instance
- 8
Step 8: taint node, evict pods through the Eviction API honoring PDBs, terminate instance
- nextFailure path: node kept; retried later
- 9
Failure path: pod stays Pending (too big, wrong labels, untolerated taints, max size or quota)
- 10
Failure path: CA gives up on those nodes, may try another group, removes unregistered instances
- 11
Failure path: node kept; retried later
Lesson map
Node-Level Scaling - Node Groups, Cluster Autoscaler vs Karpenter, Over-Provisioning, PDBs & Drains
Node scaling: node groups/ASGs/managed node pools, Cluster Autoscaler simulation and expanders, Karpenter NodePools, consolidation and disruption budgets, CA vs Karpenter, scale-up latency budget, over-provisioning with pause pods and CapacityBuffer, PDBs and drains, cost levers. HPA/VPA linked, not re-taught.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB p["Step 1: scheduler marks a pod Unschedulable (PodScheduled condition false)"] l["Step 2: every scan-interval (default 10s) CA lists unschedulable pods"] t["Step 3: for each node group CA builds a template node and simulates scheduling the pods onto it"] x["Step 4: expander picks a group: least-waste by default random, most-pods, least-nodes, price, priority, grpc"] i["Step 5: CA raises the group's desired size via the cloud API"] r["Step 6: VM boots, kubelet registers, node Ready, scheduler binds the pending pods"] d["Step 7: scale-down loop: node requests below 50 percent of allocatable and pods movable for 10 min"] e["Step 8: taint node, evict pods through the Eviction API honoring PDBs, terminate instance"] f1["Failure path: pod stays Pending (too big, wrong labels, untolerated taints, max size or quota)"] f2["Failure path: CA gives up on those nodes, may try another group, removes unregistered instances"] f3["Failure path: node kept retried later"] p -->|continues| l l -->|continues| t t -->|continues| x x -->|continues| i i -->|continues| r r -->|continues| d d -->|continues| e t -->|continues| f1 i -->|continues| f2 e -->|continues| f3
Facts from the Cluster Autoscaler FAQ worth memorizing:
| Behavior | Default |
|---|---|
| Check for unschedulable pods and for unneeded nodes | Every 10s (--scan-interval) |
| Scale-down candidate | Sum of CPU and memory requests below 50% of allocatable (--scale-down-utilization-threshold) and all pods can move |
| Time a node must be unneeded | 10 min (--scale-down-unneeded-time) |
| Pause after a scale-up before scale-down resumes | 10 min (--scale-down-delay-after-add) |
| Deletions | Non-empty nodes one at a time; empty nodes in bulk (up to 10 by default) |
| Graceful termination given to pods on scale-down | Up to 10 min (--max-graceful-termination-sec) |
| Expendable pods | Priority below -10 do not trigger scale-up and do not block scale-down |
| Reaction SLO | No more than 30s on clusters under 100 nodes (about 5s average); no more than 60s on 100 to 1,000 nodes (about 15s average), assuming no pod affinity |
Pods that block scale-down include: pods with a restrictive PDB, kube-system pods without a PDB that do not run on every node, pods not managed by a controller, pods with local storage (unless annotated safe to evict), pods whose constraints mean they cannot fit anywhere else, and pods annotated cluster-autoscaler.kubernetes.io/safe-to-evict: "false".
The FAQ is blunt about two anti-patterns: do not use CPU-usage-based node autoscalers with Kubernetes (they add nodes nobody can use or remove nodes running critical pods), and do not run another node-group autoscaler next to CA.
How Karpenter works
Karpenter watches the same unschedulable pods, but instead of choosing among fixed groups it computes which machines would fit them and launches those machines directly; the kube-scheduler still binds the pods once the new nodes register.
| Concept | Meaning |
|---|---|
| NodePool | Allowed shapes (instance families, sizes, architectures, zones, karpenter.sh/capacity-type spot or on-demand), taints, labels, limits, disruption settings |
| NodeClass | Cloud settings (image family, subnets, security groups, block devices, kubelet settings) |
| NodeClaim | Karpenter's record of one machine it launched |
| Bin-packing | Batches pending pods and chooses instance types that fit their combined requests and constraints |
| Consolidation | Deletes or replaces nodes to cut cost: WhenEmpty, Balanced, or WhenEmptyOrUnderutilized (the default, with consolidateAfter: 0s) |
| Drift | Nodes whose spec no longer matches the NodePool or NodeClass (for example a new image) are replaced |
| Expiration | expireAfter (default 720h, 30 days) caps node lifetime |
| Disruption budgets | spec.disruption.budgets, default one budget of nodes: 10% per NodePool; they gate drift, emptiness and consolidation but not expiration |
| Interruption handling | On AWS, an SQS queue fed by EventBridge rules delivers spot interruption and maintenance events; Karpenter cordons, drains and replaces ahead of termination |
Karpenter runs one disruption method at a time (drift first, then consolidation), simulates whether the pods on a candidate node can be rescheduled (possibly onto one cheaper replacement), checks budgets, pre-launches replacements, then drains through the Eviction API. Spot-to-spot consolidation is behind a feature flag and, for single-node replacements, requires at least 15 cheaper instance-type options so it does not race to the most interruptible capacity.
One interplay most teams miss: Karpenter's docs recommend configuring the kube-scheduler with the MostAllocated scoring strategy. The default LeastAllocated spreads pods across nodes, the opposite of bin-packing, which leaves Karpenter-launched nodes under-packed and makes consolidation churn.
Cluster Autoscaler vs Karpenter
| Cluster Autoscaler | Karpenter | |
|---|---|---|
| Unit of scaling | Node group with one machine shape | Individual machines chosen per batch of pods |
| Instance choice | Fixed by you per group; expander chooses among groups | Any type allowed by the NodePool, chosen for fit and price |
| Provisioning path | Cloud group API (ASG desired count) | Direct launch (EC2 Fleet on AWS) |
| Scale-down | Utilization threshold plus unneeded timer, one non-empty node at a time | Consolidation (delete or replace), drift, expiration, budgets |
| Spot handling | Separate spot node groups; interruption handling via other tools | Capacity type in requirements; native interruption queue on AWS |
| Clouds | Many providers (AWS, GCP, Azure, and others) | Provider-specific: AWS (first) and Azure are the mature ones; EKS Auto Mode and AKS node auto-provisioning run it as a managed service |
| Strengths | Predictable, mature, simple mental model, works with existing groups | Fewer groups to manage, better packing and cost, faster replacement |
| Risks | Group sprawl to cover shapes; slower to adapt; waste from fixed shapes | More churn from consolidation; needs good PDBs and requests; broad instance lists need testing |
"""Simulation: node-group autoscaling (Cluster Autoscaler style) vs groupless
provisioning (Karpenter style) for the same burst of pending pods.
Instance shapes and hourly prices are EXAMPLE numbers, not a price list.
CA logic follows its FAQ: a template node per node group, simulate whether the
pending pods fit, then an expander (default least-waste) picks the group.
Karpenter logic: pick instance types from a broad catalog to bin-pack the
pending pods at the lowest price, then consolidate when utilisation drops.
"""
DAEMON = (0.2, 0.5) # per-node DaemonSet overhead: vCPU, GiB
CATALOG = { # name: (vCPU, GiB, $/h) EXAMPLE values
"c-2x4": (2, 4, 0.085), "g-4x16": (4, 16, 0.192), "c-8x16": (8, 16, 0.340),
"g-16x64": (16, 64, 0.768), "m-8x64": (8, 64, 0.504),
}
pending = [(1.5, 3)] * 6 + [(3.5, 12)] * 2 + [(0.5, 1)] * 10 # (vCPU, GiB) requests
def pack(pods, shape):
"""first-fit decreasing onto identical nodes of one shape; returns node loads"""
cpu, mem, _ = CATALOG[shape]
nodes = []
for c, m in sorted(pods, reverse=True):
if c > cpu - DAEMON[0] or m > mem - DAEMON[1]:
return None # pod can never fit this shape
for n in nodes:
if n[0] + c <= cpu - DAEMON[0] and n[1] + m <= mem - DAEMON[1]:
n[0] += c; n[1] += m; break
else:
nodes.append([c, m])
return nodes
def report(label, shape, nodes):
cpu, mem, price = CATALOG[shape]
used = sum(n[0] for n in nodes) + DAEMON[0] * len(nodes)
print(f" {label:<26} {len(nodes)} x {shape:<8} ${price * len(nodes):.3f}/h "
f"cpu util {100 * used / (cpu * len(nodes)):.0f}%")
return price * len(nodes)
req_cpu = sum(c for c, _ in pending); req_mem = sum(m for _, m in pending)
print(f"pending: {len(pending)} pods, {req_cpu} vCPU, {req_mem} GiB requested\n")
print("Cluster Autoscaler: the node groups someone created earlier are c-8x16 and g-16x64; expander=least-waste")
options = {}
for group in ("c-8x16", "g-16x64"):
nodes = pack(pending, group)
idle = CATALOG[group][0] * len(nodes) - sum(n[0] for n in nodes) - DAEMON[0] * len(nodes)
options[group] = (idle, nodes)
print(f" simulate +{len(nodes)} {group}: idle cpu after scale-up {idle:.1f}")
choice = min(options, key=lambda g: options[g][0])
ca_cost = report(f"least-waste picks {choice}", choice, options[choice][1])
def groupless(pods):
"""toy heuristic: open one node at a time, choosing the shape with the lowest
price per packed vCPU for the pods still waiting (mixed shapes allowed)"""
left, fleet = sorted(pods, reverse=True), []
while left:
best = None
for shape, (cpu, mem, price) in CATALOG.items():
load, took = [0.0, 0.0], []
for p in left:
if load[0] + p[0] <= cpu - DAEMON[0] and load[1] + p[1] <= mem - DAEMON[1]:
load[0] += p[0]; load[1] += p[1]; took.append(p)
if took and (best is None or price / load[0] < best[0]):
best = (price / load[0], shape, took)
fleet.append(best[1])
for p in best[2]:
left.remove(p)
return fleet
def fleet_report(label, fleet):
from collections import Counter
cost = sum(CATALOG[s][2] for s in fleet)
print(f" {label:<26} {dict(Counter(fleet))} ${cost:.3f}/h")
return cost
print("\nKarpenter: NodePool allows the whole catalog (groupless), mixed shapes allowed")
k_cost = fleet_report("launch", groupless(pending))
print("\nlater: the 10 small pods finish; 1.5-vCPU and 3.5-vCPU pods remain")
remaining = [(1.5, 3)] * 6 + [(3.5, 12)] * 2
ca_nodes = pack(remaining, choice)
print(f" CA: scale-down removes a node only if its requests are < 50% of allocatable and its pods fit elsewhere,")
print(f" after 10 min unneeded -> at best {len(ca_nodes)} x {choice} ${CATALOG[choice][2] * len(ca_nodes):.3f}/h")
fleet_report("Karpenter consolidation ->", groupless(remaining))
print(f"\nburst cost: CA ${ca_cost:.3f}/h vs Karpenter ${k_cost:.3f}/h (example prices)")Output:
pending: 18 pods, 21.0 vCPU, 52 GiB requested
Cluster Autoscaler: the node groups someone created earlier are c-8x16 and g-16x64; expander=least-waste
simulate +4 c-8x16: idle cpu after scale-up 10.2
simulate +2 g-16x64: idle cpu after scale-up 10.6
least-waste picks c-8x16 4 x c-8x16 $1.360/h cpu util 68%
Karpenter: NodePool allows the whole catalog (groupless), mixed shapes allowed
launch {'g-16x64': 1, 'g-4x16': 1, 'c-2x4': 2} $1.130/h
later: the 10 small pods finish; 1.5-vCPU and 3.5-vCPU pods remain
CA: scale-down removes a node only if its requests are < 50% of allocatable and its pods fit elsewhere,
after 10 min unneeded -> at best 3 x c-8x16 $1.020/h
Karpenter consolidation -> {'g-16x64': 1, 'c-2x4': 1} $0.853/h
burst cost: CA $1.360/h vs Karpenter $1.130/h (example prices)The prices are made up, but the shape of the result is typical. CA can only choose among the groups you created, so the cheapest packing may not be available to it. A groupless provisioner can mix shapes for one burst and then re-pack when load drops. That flexibility is also its cost: nodes change more often, so workloads must tolerate eviction.
Interplay with HPA and VPA (without repeating them)
- HPA creates the pending pods; the node autoscaler reacts to them. If HPA cannot add replicas because of
maxReplicas, the node autoscaler never hears about the demand. - Requests are the contract. Both CA and Karpenter simulate scheduling from requests. Over-requesting wastes nodes; under-requesting packs too tightly and leads to throttling and eviction. VPA recommendations feed directly into node counts.
- Scale-down needs movable pods. Every PDB with
maxUnavailable: 0, every bare pod and everysafe-to-evict: "false"pins a node.
The scale-up latency budget and over-provisioning
// Simulation: the latency budget from "traffic spike" to "new pod serving".
// Stage values marked FAQ come from the Cluster Autoscaler FAQ (HPA ~1 min typical,
// CA reaction < 30 s in most cases, GCE node provisioning usually 3-4 min).
// Values marked EXAMPLE are assumptions for illustration; measure your own.
type Stage = { name: string; seconds: number; source: "FAQ" | "EXAMPLE" };
const hpa: Stage = { name: "HPA notices load and raises replicas", seconds: 60, source: "FAQ" };
const ca: Stage = { name: "autoscaler sees Pending pod, requests node", seconds: 30, source: "FAQ" };
const boot: Stage = { name: "VM boots, kubelet registers, CNI ready", seconds: 210, source: "FAQ" };
const pull: Stage = { name: "image pull on a cold node", seconds: 40, source: "EXAMPLE" };
const pullWarm: Stage = { name: "image already cached on node", seconds: 2, source: "EXAMPLE" };
const ready: Stage = { name: "app start + readiness probe passes", seconds: 20, source: "EXAMPLE" };
const preempt: Stage = { name: "preempt a pause pod (grace 0) and bind", seconds: 2, source: "EXAMPLE" };
const total = (s: Stage[]) => s.reduce((a, b) => a + b.seconds, 0);
function show(title: string, stages: Stage[]) {
console.log(`${title}: ${total(stages)} s`);
for (const s of stages) console.log(` ${String(s.seconds).padStart(4)} s ${s.name} [${s.source}]`);
}
const cold = [hpa, ca, boot, pull, ready];
const buffered = [hpa, preempt, pullWarm, ready];
show("A) no spare capacity", cold);
show("B) over-provisioned with low-priority pause pods", buffered);
console.log(" (the evicted pause pod goes Pending and triggers a node add in the background)");
// What does the buffer cost? EXAMPLE: buffer = 2 nodes at $0.192/h
const bufferNodes = 2, price = 0.192, hoursPerMonth = 730;
console.log(`\nbuffer cost: ${bufferNodes} nodes x $${price}/h x ${hoursPerMonth} h = $${(bufferNodes * price * hoursPerMonth).toFixed(2)}/month (example price)`);
// Where should you spend effort? Show each stage's share of the cold path.
console.log("\nshare of cold path:");
for (const s of cold) console.log(` ${((100 * s.seconds) / total(cold)).toFixed(0).padStart(3)}% ${s.name}`);Output:
A) no spare capacity: 360 s
60 s HPA notices load and raises replicas [FAQ]
30 s autoscaler sees Pending pod, requests node [FAQ]
210 s VM boots, kubelet registers, CNI ready [FAQ]
40 s image pull on a cold node [EXAMPLE]
20 s app start + readiness probe passes [EXAMPLE]
B) over-provisioned with low-priority pause pods: 84 s
60 s HPA notices load and raises replicas [FAQ]
2 s preempt a pause pod (grace 0) and bind [EXAMPLE]
2 s image already cached on node [EXAMPLE]
20 s app start + readiness probe passes [EXAMPLE]
(the evicted pause pod goes Pending and triggers a node add in the background)
buffer cost: 2 nodes x $0.192/h x 730 h = $280.32/month (example price)
share of cold path:
17% HPA notices load and raises replicas
8% autoscaler sees Pending pod, requests node
58% VM boots, kubelet registers, CNI ready
11% image pull on a cold node
6% app start + readiness probe passesExpectedA) no spare capacity: 360 s 60 s HPA notices load and raises replicas [FAQ] 30 s autoscaler sees Pending pod, requests node [FAQ] 210 s VM boots, kubelet registers, CNI ready [FAQ] 40 s image pull on a cold node [EXAMPLE] 20 s app start + readiness probe passes [EXAMPLE] B) over-provisioned with low-priority pause pods: 84 s 60 s HPA notices load and raises replicas [FAQ] 2 s preempt a pause pod (grace 0) and bind [EXAMPLE] 2 s image already cached on node [EXAMPLE] 20 s app start + readiness probe passes [EXAMPLE] (the evicted pause pod goes Pending and triggers a node add in the background) buffer cost: 2 nodes x $0.192/h x 730 h = $280.32/month (example price) share of cold path: 17% HPA notices load and raises replicas 8% autoscaler sees Pending pod, requests node 58% VM boots, kubelet registers, CNI ready 11% image pull on a cold node 6% app start + readiness probe passes
Press Run. Snippets must be self-contained — no network, files, or native modules.
Node boot dominates, so the cheapest latency fix is often to keep spare capacity warm. The classic pattern from the CA FAQ: run a Deployment of pause pods with a very low PriorityClass (below zero but above the -10 expendable cutoff, so they still trigger scale-up) and requests sized like your typical burst. Real pods preempt them instantly. The evicted pause pods go Pending, which makes the autoscaler add a node in the background. Karpenter and CA also implement a CapacityBuffer API (autoscaling.x-k8s.io/v1alpha1, alpha in Karpenter) that expresses the same buffer as virtual pods instead of real ones. Other levers: pre-pulled or smaller images, faster-booting images, and warm pools.
PodDisruptionBudgets and draining
Node removal, node upgrades and consolidation are voluntary disruptions. They go through the Eviction API, which honors PodDisruptionBudgets: an eviction that would violate a PDB gets 429 Too Many Requests, and the drainer retries. kubectl drain does the same thing: cordon, then evict.
| PDB setting | Effect | Pitfall |
|---|---|---|
minAvailable: N or a percentage | At least N pods must stay healthy | minAvailable equal to replicas means no voluntary eviction ever |
maxUnavailable: N | At most N pods down at once | maxUnavailable: 0 blocks drains and scale-down forever |
unhealthyPodEvictionPolicy: IfHealthyBudget (default) | Running but not-ready pods are evictable only while the app is not disrupted | A crash-looping app can block a drain |
unhealthyPodEvictionPolicy: AlwaysAllow | Not-ready running pods are always evictable | Usually the better choice for stateless apps |
PDBs do not protect against node-pressure eviction by the kubelet or a VM that simply dies; they only rate-limit voluntary disruption. Rollout strategies and PDB sizing for deployments are covered in Rolling, Blue-Green & Canary.
Cost levers
| Lever | Saves | Watch out for |
|---|---|---|
| Right-sized requests (VPA recommendations) | The biggest lever: fewer nodes for the same work | Under-requesting causes eviction and throttling |
Bin-packing (MostAllocated, groupless provisioning) | Fewer half-empty nodes | Less headroom per node for spikes |
| Spot or preemptible capacity | Large discounts for interruptible work | Interruptions; needs PDBs, multiple instance types and graceful shutdown |
| Consolidation | Removes waste after load drops | Churn; long-running jobs need do-not-disrupt style protection |
| Scale-down timers | Shorter timers return money sooner | Thrash: scale down, then pay boot latency again |
| Over-provisioning buffer | Buys latency | Costs a fixed amount every hour (the run above: about $280 per month for 2 example nodes) |
| Larger nodes | DaemonSet and system overhead amortized over more pods | Bigger blast radius per node failure; slower drains |
Decision chart: picking a node scaling strategy.
Decisions
- 1
A
- nextUse it: EKS Auto Mode, GKE autoscaling or node auto-provisioning, AKS node auto-provisioning
- nextStep 2: on AWS or Azure and willing to let the autoscaler pick instance types?
- 2
Use it: EKS Auto Mode, GKE autoscaling or node auto-provisioning, AKS node auto-provisioning
- ?
Step 2: on AWS or Azure and willing to let the autoscaler pick instance types?
- nextKarpenter: NodePools with broad instance families, budgets, MostAllocated scoring
- nextCluster Autoscaler with a few node groups, least-waste or priority expander
- 4
Karpenter: NodePools with broad instance families, budgets, MostAllocated scoring
- nextStep 3: is scale-up latency within your SLO?
- 5
Cluster Autoscaler with a few node groups, least-waste or priority expander
- nextStep 3: is scale-up latency within your SLO?
- ?
Step 3: is scale-up latency within your SLO?
- nextAdd a buffer: low-priority pause pods or CapacityBuffer, smaller images, warm pools
- nextStep 4: do drains and scale-down actually happen?
- 7
Add a buffer: low-priority pause pods or CapacityBuffer, smaller images, warm pools
- nextStep 4: do drains and scale-down actually happen?
- ?
Step 4: do drains and scale-down actually happen?
- nextFix PDBs with maxUnavailable 0, bare pods, safe-to-evict false, local storage
- nextTune timers and consolidation policy against churn
- 9
Fix PDBs with maxUnavailable 0, bare pods, safe-to-evict false, local storage
- 10
Tune timers and consolidation policy against churn
What happens if you choose an alternative
| Choice | Instead of | Consequence |
|---|---|---|
| CPU-based ASG scaling policies | A pending-pod-aware autoscaler | Nodes added that pods cannot use, or nodes removed while running critical pods |
| One node group per instance shape, dozens of groups | Karpenter or a few flexible groups | Slow CA simulations, hard capacity planning, waste |
| No PDBs | PDBs on every multi-replica service | Consolidation and upgrades can take down all replicas at once |
PDBs with maxUnavailable: 0 | Realistic budgets | Nodes never drain; upgrades stall; cost leaks |
| Aggressive consolidation on stateful or long jobs | WhenEmpty or do-not-disrupt protection | Repeated restarts of work that is expensive to redo |
| No buffer for spiky traffic | Pause pods or CapacityBuffer | Minutes of degraded service while nodes boot |
Pitfalls
- Pending pods the autoscaler cannot help: a pod larger than any allowed shape, node selectors for labels no group has, taints no group tolerates, or a zone-pinned volume in a zone at max size.
- Cloud quota and capacity: the autoscaler asks, the cloud says no. Watch for insufficient-capacity errors and spread across instance types and zones.
- Zonal volumes plus scale-to-zero groups per zone: CA must know which zone a template node is in, or a pod pinned to zone b may trigger scale-up in zone a.
- DaemonSet overhead on small nodes: every node pays it, so tiny nodes waste a large share.
- Long graceful termination (for example 1 hour) makes every drain and consolidation slow, and CA caps graceful termination at 10 minutes by default anyway.
Interview Q&A
How does Cluster Autoscaler decide to add a node?
Answer
Every scan interval it looks for pods the scheduler marked unschedulable, builds a template node for each node group, simulates whether the pods would fit, and uses an expander (least-waste by default) to choose which group to grow. It does not look at CPU utilization.
When does it remove a node?
Answer
When the sum of the node's pod requests is below the utilization threshold (50% by default), all its pods can be moved elsewhere (PDBs, controllers, local storage and constraints allow it), and this has been true for the unneeded time (10 minutes by default). It then evicts pods through the Eviction API and terminates the instance.
How is Karpenter different?
Answer
It is groupless: it chooses instance types per batch of pending pods from everything a NodePool allows and launches machines directly, then continuously consolidates by deleting or replacing nodes to reduce cost, handles drift and expiration, and on AWS reacts to spot interruption notices via an SQS queue.
Why might scale-up take five minutes, and how do you make it faster?
Answer
Most of it is VM boot and node registration, plus image pulls and readiness. The autoscaler itself reacts in seconds. Keep warm spare capacity with low-priority pause pods or a capacity buffer, shrink and pre-pull images, and use faster-booting node images.
What do PodDisruptionBudgets protect against?
Answer
Voluntary disruptions through the Eviction API: drains, autoscaler scale-down, consolidation and upgrades. They do not protect against node crashes or kubelet node-pressure eviction.
Why should requests be accurate for cost?
Answer
Both the scheduler and node autoscalers pack by requests. Inflated requests leave real capacity idle but "full", which forces extra nodes; deflated requests over-pack nodes and cause throttling and eviction.
Which pods block Cluster Autoscaler scale-down?
Answer
Pods with a restrictive PDB, kube-system pods without a PDB that do not run on every node, bare pods, pods with local storage, pods that cannot fit elsewhere, and pods annotated safe-to-evict false.
What is a CapacityBuffer?
Answer
An autoscaling API, alpha in Karpenter, that expresses spare capacity as virtual pods instead of real pause pods.
Check yourself
In a test cluster, scale a Deployment beyond current capacity and time each phase with kubectl get events and node creation timestamps. Compare your numbers with the latency budget above.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: HPA, VPA & Autoscaling Gotchas — Metrics, Stabilization & Thrash, Requests, Limits & QoS — CPU Throttling, Memory OOM & Scheduling, Rolling, Blue-Green & Canary — Strategies, PDBs & Blast Radius, Private Networking — VPC, NAT, SSM, Tunnels & Ingress, Multi-AZ vs Multi-Region - Blast Radius, Cell Architecture & Static Stability, Load, Chaos & Production Validation — Shadow Traffic, Canaries & Game Days, Region Failover in Practice - Runbooks, DNS/GLB Cutover, Failback & Game Days, Service-to-Service Auth — mTLS, Client Credentials & Workload Identity, Observability Triad — Metrics, Logs & Distributed Tracing.