Requests, Limits & QoS — CPU Throttling, Memory OOM & Scheduling
Requests drive scheduling; limits cap cgroups. Guaranteed / Burstable / BestEffort decide who dies under pressure. CPU over limit throttles; memory over limit OOMs. Overcommit without quotas is how noisy neighbors win.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
A latency-critical service on a shared node
Prefer
Requests set, limits chosen on purpose
The scheduler can place the Pod. Memory has a ceiling at or above peak. CPU limit is a product decision, not a copy of the request.
- Guaranteed if you set CPU and memory request equal to limit on every container.
- Burstable if you want burst headroom above the request.
- You can name who loses under node memory pressure.
Alternative
BestEffort, no requests or limits
The Pod schedules into whatever slack exists. Under pressure it is evicted first. User traffic should not live here.
- No reservation, so neighbors can starve it.
- Pending is less likely until the node is actually full of requests from others.
- Fine for scratch batch. Wrong for a serving path.
From YAML to enforcement
Requests are a scheduling contract. Limits become cgroup caps after the Pod is bound.
- 1
Pod is created
Each container may set CPU and memory requests and limits. - 2
Scheduler filters
A Node must have allocatable room for the sum of requests. - 3
Kubelet cgroup
Limits become the hard cap. CPU quota throttles. Memory overage can kill. - 4
Node pressure
If the node itself is out of memory, eviction order follows QoS and priority.
Overview
Requests and limits are a reliability feature. They decide where Pods land, who gets CPU when the node is busy, and who dies first when memory is scarce. Bad values show up as Pending Pods, p99 latency from throttling, or a sudden OOMKill.
The hub named the loop. Deployments create the Pods. Probes decide Ready. This page is why a Pod stays Pending, why it gets slow without dying, and why the kernel kills it.
Requests versus limits
Requests are what the scheduler adds up. The sum must fit node allocatable (capacity minus system reservations). CPU requests also influence relative weight under contention. A memory request reserves scheduling room and feeds the QoS class. It does not, by itself, cap the process.
Limits are enforced by the kubelet through cgroups.
- CPU over the limit throttles. The process lives. Threads stall. p99 rises. You see it in CPU throttling stats, not as a restart.
- Memory over the limit OOMKills the container. The process is gone. If the leak remains, you get a crash loop.
Practical rule: set the memory request near the expected working set and the memory limit at or above peak, with headroom. Set the CPU request from the latency SLO you need to schedule. A CPU limit is optional and often harmful on a dedicated pool, because throttle looks like a slow app. A missing CPU limit on a shared multi-tenant node lets one Pod steal cycles. Say which cluster you are in before you drop the limit.
QoS classes
Kubernetes assigns one class to the whole Pod.
| Class | Rule | Who dies first under node pressure | Typical use |
|---|---|---|---|
| Guaranteed | Every container has CPU request equal to CPU limit, and memory request equal to memory limit | Last | Latency-critical, stateful primaries |
| Burstable | At least one request or limit, but not fully Guaranteed | Middle | Most services that should burst |
| BestEffort | No container sets a request or a limit | First | Disposable batch only |
Guaranteed is strict. If any container omits CPU, or sets memory request below memory limit, the Pod is Burstable. Init containers and sidecars count. A perfect app container plus a BestEffort sidecar demotes the Pod.
Eviction under node memory pressure is not the same event as a container OOMKill. OOMKill is "this cgroup exceeded its limit." Eviction is "the node is in trouble, pick a Pod." QoS and PriorityClass order the eviction. A Guaranteed Pod can still be OOMKilled if it exceeds its own memory limit.
Throttle versus OOM
CPU throttle is a latency bug you can miss if you only alert on restarts. The process is up, probes may pass, and users see slowness. Memory OOMKill is a restart. If the working set was simply larger than the limit, raising the limit or finding the leak is the fix. Tightening the CPU limit "for safety" can fail probes that were fine, because the probe timeout expires while the process is throttled.
Do not treat "a limit exists" as automatically good. On a node pool that runs one workload, a CPU limit mostly adds throttle. On a shared pool, no limit and a huge request is how you pack badly. Overcommit is normal: the sum of limits may exceed the node while the sum of requests still fits. That is safe until everyone bursts together.
LimitRange, ResourceQuota, and neighbors
LimitRange defaults, and min or max, for requests and limits in a namespace. It stops a Deployment with empty resources from landing as BestEffort by accident. It also stops a single container from requesting the whole node.
ResourceQuota caps the namespace: sum of requests, sum of limits, object counts. One team cannot schedule unbounded replicas. A Pending Pod with "exceeded quota" is not a scheduler puzzle. The quota object is the reason.
Overcommit without quota is how a noisy neighbor wins. Their Burstable Pod bursts into your cycles, or their BestEffort Pod was only fine until the node hit memory pressure and yours was BestEffort too.
VPA changes request and limit sizes. HPA changes replica count. Both on CPU oscillate: VPA raises the request, utilization falls, HPA scales in, load per Pod rises. The full conflict, stabilization, and the formula are HPA, VPA & Autoscaling. A common split is HPA on a custom metric, and VPA in Off or Initial so it recommends or sets the size only at Pod create.
Decisions
- 1
1. Pod sets requests
- next2. Filter by requests
- 2
2. Filter by requests
- next3. Score and bind
- 3
3. Score and bind
- next4. Kubelet sets cgroups
- 4
4. Kubelet sets cgroups
- next5. What exceeded
- ?
5. What exceeded
- cpu6. CFS throttles CPU
- memory7. OOMKill the container
- node8. Evict by QoS class
- 6
6. CFS throttles CPU
- 7
7. OOMKill the container
- 8
8. Evict by QoS class
Lesson map
Requests, Limits & QoS — CPU Throttling, Memory OOM & Scheduling
Requests drive scheduling; limits cap cgroups. Guaranteed / Burstable / BestEffort decide who dies under pressure. CPU over limit throttles; memory over limit OOMs. Overcommit without quotas is how noisy neighbors win.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Pod sets requests"] b["2. Filter by requests"] c["3. Score and bind"] d["4. Kubelet sets cgroups"] a -->|1. Pod sets requests| b b -->|2. Filter by requests| c c -->|3. Score and bind| d
Classify a Pod, in memory
Guaranteed requires every container to set both CPU and memory, with request equal to limit for each. Anything else with at least one number is Burstable. Empty resources are BestEffort.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Expect Guaranteed, Burstable, BestEffort, then Burstable again. The last case is a perfect app container beside an empty sidecar. One empty container demotes the Pod.
Interview Q&A
What does the scheduler look at, requests or limits?
Answer
Requests, against node allocatable. Limits do not reserve space. A Pod with a huge limit and a tiny request can schedule onto a node that cannot survive the burst.
What happens when a container exceeds its CPU limit?
Answer
The kernel throttles it through the CFS quota. The process is not killed. Latency climbs. Probes can time out if the throttle is severe. Look at throttling stats before you restart the Deployment.
What happens when it exceeds its memory limit?
Answer
The container is OOMKilled. That is a cgroup event, not node eviction. If the working set is legitimate, the limit is too low. If it grows without bound, the limit did its job and you still have a leak.
How do you get Guaranteed QoS?
Answer
Every container, including sidecars, sets CPU request equal to CPU limit and memory request equal to memory limit. Both resources, both equal, every container. Miss one and the Pod is Burstable.
Why omit a CPU limit?
Answer
On a dedicated node pool, the limit's main effect is throttle, which looks like a slow service. Keep the request so the scheduler still reserves CPU. On a shared cluster, omitting the limit needs a quota and a neighbor policy, or one Pod can consume the node.
What is ResourceQuota for?
Answer
A namespace ceiling: total requests, total limits, and sometimes object counts. It stops one team from scheduling the whole cluster. LimitRange is the per-container default and min or max. Quota is the sum.
How does VPA interact with HPA?
Answer
VPA edits the request. HPA's CPU percentage is usage divided by that request. If both react to CPU, the request and the replica count chase each other. Split the signal or run VPA in Off or Initial. The next-but-one lesson is the full conflict.
Pods are Pending with insufficient CPU. What do you check?
Answer
Sum of requests versus node allocatable, taints and affinity, ResourceQuota, and whether a higher PriorityClass should preempt. Limits are not the first column. A request of 8 CPU will not place on a node with 4 allocatable, no matter the limit.
What is overcommit?
Answer
Requests fit on the node. Limits sum to more than the node. Everyone can burst until they burst together. Then you get throttle, OOM, or eviction. Quota and honest requests are how you bound that bet.
Why did a Guaranteed Pod still get killed?
Answer
Guaranteed is last in line for node eviction. It is not immune to its own memory limit. Exceed the limit and the cgroup OOM killer runs. QoS does not raise the limit for you.
Pitfalls
- Copying request equal to a tiny tutorial value, then wondering why HPA thinks you are at 400 percent CPU.
- A CPU limit equal to a small request on a latency path. You throttled yourself.
- Memory limit below the real heap peak. Restart loop, probes never settle.
- No LimitRange, so a new service ships BestEffort into a cluster that evicts it on the first pressure event.
- Forgetting the sidecar. The Pod class is only as strong as the weakest container.
A node has 4 allocatable CPUs. Pod A requests 2 and limits at 2. Pod B requests 1 and limits at 3. Pod C sets nothing. Which QoS is each, which ones schedule together, and who is evicted first if the node hits memory pressure?