HPA, VPA & Autoscaling Gotchas — Metrics, Stabilization & Thrash
HPA scales replicas from CPU/memory/custom/external metrics. Stabilization windows stop flap. VPA changes requests/limits and fights HPA if both target the same resource. Queue consumers often need custom metrics, not raw CPU.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
CPU is busy because each Pod is too small
Prefer
Split the actuators
HPA adds Pods from a signal that means user load. VPA, if you run it, sizes requests offline or only at Pod create.
- One controller owns replica count.
- Request changes are reviewed, or applied as Initial.
- Stabilization stops a one-minute spike from doubling the fleet.
Alternative
HPA and VPA Auto both on CPU
VPA raises the request, so utilization falls, so HPA removes Pods, so each Pod gets hotter, so VPA raises the request again.
- Replica count and request size oscillate.
- Recreate mode evicts Pods to apply the new size.
- You page on churn instead of on the original load.
One HPA sync
Resource metrics need metrics-server. Custom and external metrics need an adapter. No samples means no confident scale.
- 1
Read metrics
CPU or memory versus requests, or a custom or external value such as queue lag. - 2
Per-metric replicas
ceil of current replicas times current value over target. Then take the max. - 3
Clamp and behavior
Honor min and max replicas. Apply scale-up and scale-down policies and the stabilization window. - 4
Write the scale
Update the Deployment scale. Scale-down still has to pass the PDB.
Overview
targetCPUUtilizationPercentage: 70 looks finished until production. Replicas flap. Scale-down hits a PDB. VPA resizes the CPU request HPA is watching. This lesson is signals, windows, and conflicts.
The workload spine ends here: controllers, probes, capacity, progressive delivery, then autoscaling. Pair entry traffic with Private Networking. Keep this page on the desired-state loop inside the cluster.
The replica formula
For a resource utilization metric, the simplified ratio is:
desiredReplicas = ceil(currentReplicas * (currentMetric / targetMetric))
CPU utilization is usage divided by the request, not the limit. A tiny request makes a calm Pod look 400 percent busy. That is why requests and HPA are one story.
Several metrics: compute a desired count per metric, then take the maximum. A hot CPU and a cool memory still scales to the CPU answer. Then clamp to minReplicas and maxReplicas.
Example: 4 replicas, CPU at 90 against a target of 60. ceil(4 * 90 / 60) = ceil(6) = 6. A second metric at 40 against 70 wants fewer Pods. The max still wins, so you stay on the scale-up path.
A toy that treats current replicas of 0 as 1 hides a real limit: CPU HPA does not scale from zero. There is no Pod to measure. Waking an idle Deployment needs an external signal that exists without Pods, or a tool built for scale-from-zero. Do not claim the CPU percentage will do it.
If metrics-server is missing, resource HPA stays unknown. Custom metrics (per Pod or per object) and external metrics (a queue, a pub/sub backlog, an adapter) are how you scale on something other than cgroup CPU.
Policies and stabilization
spec.behavior.scaleUp and scaleDown limit how fast the count may change: a percent, a Pod count, and a periodSeconds. The stabilization window looks backward.
- Scale-down window: use the highest recommendation in the window, so a brief quiet period does not delete Pods you will need again.
- Scale-up window: use the lowest recommendation in the window, so one spike does not jump straight to the max.
- Defaults people forget: scale-down stabilization is on the order of five minutes. Scale-up stabilization defaults to zero, so up is fast and down is slow. That is intentional.
Without a window, a spike adds Pods, cold caches make CPU worse, and you add more Pods. Or a dip removes Pods, the remainder overloads, and you add them back. That is thrash.
HPA sync is periodic (classically 15 seconds). A batch spike shorter than the period is over before the controller acts. Do not use HPA as a sub-second load shed. That is a different pattern.
When CPU is the wrong signal
- Queue consumers. CPU can sit low while lag grows, or sit high on a tight loop that is not the backlog. Scale on lag or backlog (external or custom). Kafka consumer lag belongs with that metric, not with a 70 percent CPU target.
- Spikes shorter than the sync period. HPA arrives late. Size for the spike, or shed load in process.
- Memory with a heap cache. After scale-up the heap often does not shrink, so memory utilization stays high and HPA never scales in. Memory HPA is sticky.
- A single-threaded process with a tiny CPU request. Utilization math says you are saturated because the denominator is a lie. Fix the request before you trust the percentage.
PDB and scale-down
Scale-down evicts Pods. Eviction is a voluntary disruption, so the PodDisruptionBudget can refuse it. HPA will want fewer replicas and keep failing to remove them. Scale-up usually proceeds. Watch quota and node capacity: new Pods that stay Pending need a cluster autoscaler or a bigger pool, which is slower and is not an application metric.
VPA modes
| Mode | What it does | Serving-path note |
|---|---|---|
| Off | Recommendations only | Safe default while you learn the real request |
| Initial | Sets requests when the Pod is created | No live resize fight with HPA |
| Recreate (older docs say Auto) | Updates resources, often by evicting the Pod | Conflicts with a CPU HPA on the same resource |
Safe patterns: HPA on CPU with VPA Off, and you apply recommendations in review. Or HPA on a custom or external metric while VPA owns CPU and memory size. Do not run both autonomously on the same resource.
Cluster autoscaler adds nodes when Pods are Pending. It is underneath HPA when the bottleneck is the node, not the replica count. It is slow relative to HPA, and it stops at cloud quota. Mention it when Pods are Pending after a scale-up. Do not substitute it for a lag metric.
Decisions
- 1
1. Adapter scrapes signals
- next2. HPA sync runs
- 2
2. HPA sync runs
- next3. Replicas per metric
- 3
3. Replicas per metric
- next4. Take max, clamp
- 4
4. Take max, clamp
- next5. Apply window
- 5
5. Apply window
- next6. Desired vs current
- ?
6. Desired vs current
- up7. Raise replica count
- down8. Lower if PDB allows
- same9. No change this sync
- 7
7. Raise replica count
- next10. Next sync sees load
- 8
8. Lower if PDB allows
- next10. Next sync sees load
- 9
9. No change this sync
- next10. Next sync sees load
- 10
10. Next sync sees load
Lesson map
HPA, VPA & Autoscaling Gotchas — Metrics, Stabilization & Thrash
HPA scales replicas from CPU/memory/custom/external metrics. Stabilization windows stop flap. VPA changes requests/limits and fights HPA if both target the same resource. Queue consumers often need custom metrics, not raw CPU.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Adapter scrapes signals"] b["2. HPA sync runs"] c["3. Replicas per metric"] d["4. Take max, clamp"] a -->|1. Adapter scrapes signals| b b -->|2. HPA sync runs to 3. Replicas per metric| c c -->|3. Replicas per metric| d
desiredReplicas, in memory
The function takes the max across metric pairs, then clamps. It does not model the stabilization window. A zero replica count stays zero before the clamp, so a minReplicas of 2 still floors the result. That is the clamp, not evidence that CPU metrics exist at zero Pods.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Expect 6, then 6 again (the max across metrics), then 4 (ceil(8 * 0.5) = 4), then 2 because a current of 0 produces no candidate and the clamp lifts 0 to minReplicas. The last line is the floor, not a measured scale-from-zero.
Interview Q&A
Write the basic HPA replica formula.
Answer
ceil(currentReplicas * currentMetric / targetMetric). Do that per metric, take the max, then clamp to min and max replicas. CPU current is utilization against the request.
Why a stabilization window on scale-down?
Answer
Metrics oscillate around the target. The scale-down window keeps the highest recent recommendation so a short lull does not delete Pods. The scale-up window keeps the lowest recent recommendation so a spike does not jump to the ceiling. Default down is patient. Default up is not.
Why can memory HPA be sticky?
Answer
Heaps and caches often stay large after you add Pods. Utilization does not fall, so the ratio never asks for fewer replicas. Prefer a signal that drops when you are over-provisioned, or accept that memory scale-in needs a different policy.
What is the best metric for a queue consumer?
Answer
Lag or backlog, as a custom or external metric. CPU can be low while the queue grows, and high while a consumer spins. The replica count should track the work waiting, not the cgroup.
What does an HPA versus VPA conflict look like?
Answer
Replica counts and request sizes move against each other. Pods get evicted if VPA recreates them to apply a new request. Utilization drops because the denominator grew, HPA scales in, and the cycle repeats. Split the resource or the mode.
HPA wants to scale down and the PDB refuses. Then what?
Answer
Evictions fail. Replicas stay high until the budget allows a deletion or the desired count comes back up. The autoscaler does not override PodDisruptionBudget. Look at events for disruption failures, not only at the HPA object.
Does HPA use limits or requests for CPU utilization?
Answer
Requests. The resource metrics API reports usage as a percentage of the request. A missing or tiny request makes the percentage nonsense. Limits still throttle, which can change the usage you observe, but they are not the denominator.
When is the cluster autoscaler the missing piece?
Answer
When HPA has raised replicas and the new Pods are Pending because no node has room for their requests. Adding Pods cannot help until something adds nodes, and that something is slower than the HPA loop. Quota can stop it anyway.
Can CPU HPA wake a Deployment at zero replicas?
Answer
No. There is no container to report CPU. Keep minReplicas at or above 1 for a CPU target, or scale on an external metric that exists while the Deployment is idle.
What did this cluster deliberately skip?
Answer
Ingress, Gateway API, and north-south weights. Those are Ingress Controllers & North-South. You now have the inside of the loop: Deployment to ReplicaSet to Pod, probes, QoS, rollout strategy, PDB, and autoscaling that does not thrash.
Pitfalls
- Target 70 percent CPU with a request copied from a hello-world manifest.
- No stabilization, then a page every time the cron job runs for a minute.
- VPA Recreate on the same CPU the HPA scales.
- Scale-down "stuck" that is actually a PDB.
- Pending Pods after scale-up, blamed on HPA, when the node pool is full.
Eight replicas, CPU 30 against a target of 60, queue lag 2 times the target. minReplicas is 2 and maxReplicas is 20. What does HPA want before the stabilization window, and why might the Deployment still show 8 after that recommendation?