← blog · September 26, 2026

Kubernetes Resource Requests and Limits: CPU Throttling, OOMKilled and QoS Classes

A request is the scheduler's reservation, a limit is the kernel's quota. Why CPU gets paused while memory gets killed, how QoS classes decide eviction order, and how to derive the right numbers from measurement instead of guesswork.

Filling in the resources block of a pod looks trivial, but it encodes the two decisions that shape cluster health more than anything else: which node the scheduler will place this workload on, and who gets sacrificed first when that node runs short. Requests and limits sit on adjacent lines, yet they operate at different layers. A request is a reservation used only by the scheduler; nothing is physically set aside, but the sum of requests can never exceed a node's allocatable capacity. A limit is enforced by the kernel: a cgroup quota for CPU, a cgroup ceiling for memory. Teams that blur the two end up with either a cluster that looks "full" while idling or services that quietly slow down under load.

CPU and memory are not the same kind of resource

CPU is compressible. A container that exceeds its limit is not killed, it is paused. The kubelet translates a CPU limit into a Linux CFS quota: in every period, 100 ms by default, the container receives CPU time equal to its limit, and once the quota is spent the process is suspended until the period ends. A 500m limit means "50 ms out of every 100 ms." That sounds fine for a single-threaded workload, but a multi-threaded application running on eight cores at once burns through 50 ms of quota in about 6 ms and then sits idle for the remaining 94 ms of the period. Average CPU usage on the dashboard shows 30 percent while request latency multiplies. This is the classic signature of throttling, and most teams misdiagnose it as an application bug.

Memory is not compressible. When a container crosses its limit the kernel's OOM killer steps in, the process dies with exit code 137, and the pod status reads OOMKilled. There is no pause and no warning. The measured number is also not what you expect: the kubelet and kubectl top report the working set, which is RSS plus active page cache. A container that writes a lot of files appears far heavier than what its code actually holds.

QoS classes and eviction order

Kubernetes derives a quality of service class for each pod from the relationship between requests and limits:

  • Guaranteed: every container has requests equal to limits for both CPU and memory.
  • Burstable: at least one request or limit is set, but they are not all equal.
  • BestEffort: nothing is set at all.

When a node comes under memory pressure, the kubelet evicts BestEffort pods first, then Burstable pods that are using more than they requested; Guaranteed pods go last. Leaving a critical service as Burstable because "it has a request anyway" means accepting eviction caused by a neighbour's overflow even while you stay under your own limit. Nor can you use the full node: allocatable capacity is total capacity minus kubelet and system reservations minus the eviction threshold. The Allocatable line in kubectl describe node is the real budget, not the Capacity line.

A defensible starting point

  1. Set memory request equal to memory limit. A limit is mandatory because memory overflow harms neighbours; if the request is lower than the limit, the pod is placed as if there were room and then pushes the node into pressure at its true consumption.
  2. Set a CPU request, and add a CPU limit only when you have a reason. The request already provides fair sharing as a CFS weight: when the node is saturated everyone gets a share proportional to their request, and when it is idle everyone runs freely. A limit only makes sense to hard-cap a noisy neighbour or to qualify for the Guaranteed class.
  3. If you need Guaranteed, set the CPU limit equal to the request and ask for whole cores; with the kubelet's static CPU manager policy, Guaranteed pods requesting integer CPUs get exclusive cores and quota throttling disappears entirely.
resources:
  requests:
    cpu: "500m"
    memory: "512Mi"
  limits:
    memory: "512Mi"

Two safety nets exist at namespace level. A LimitRange assigns defaults to pods that declare nothing, preventing them from falling into BestEffort; a ResourceQuota caps a team's total reservation.

apiVersion: v1
kind: LimitRange
metadata:
  name: defaults
spec:
  limits:
    - type: Container
      defaultRequest:
        cpu: "100m"
        memory: "128Mi"
      default:
        memory: "256Mi"

How to find the numbers

Do not guess, measure. Run under real traffic for a week and look at the peak and 95th percentile of container_memory_working_set_bytes and rate(container_cpu_usage_seconds_total[5m]). Put the memory request near the peak and the CPU request at the 95th percentile. The Vertical Pod Autoscaler's recommendation mode (updateMode: "Off") does this measurement for you and writes the suggested values into its own status without touching the pod; running only in this mode for months before switching on automatic updates is a good habit.

To catch throttling, watch the ratio container_cpu_cfs_throttled_periods_total / container_cpu_cfs_periods_total. A container that stays above 25 percent either does not deserve a limit or has one that is far too low. For OOM, the kube-state-metrics series kube_pod_container_status_last_terminated_reason{reason="OOMKilled"} counts restarts by cause; looking at restart counts alone mixes probe-driven restarts with memory deaths.

Pitfalls

The runtime does not know the limit. The JVM reads the container limit, but by default it turns only a quarter of the memory limit into heap; unless you raise -XX:MaxRAMPercentage, you get a service that is granted 2 GiB and uses 512 MiB. In the opposite direction, the Go runtime before version 1.25 ignored the cgroup CPU quota entirely and picked GOMAXPROCS from the machine's full core count, which is exactly the throttling scenario above. In Go, the GOMEMLIMIT environment variable also lets the garbage collector see the memory limit; without it a service near its limit gets OOM-killed instead of slowing down.

Init containers change the arithmetic. A pod's effective request is the larger of the sum of its regular containers and the largest single init container. An init container that asks for 2 GiB for setup means a 2 GiB reservation on that node for as long as the pod lives, even if your application uses 256 MiB.

High requests, empty cluster. If requests sit far above real usage, the scheduler considers nodes full, new pods wait in Pending, and the cluster autoscaler spins up machines nobody needs. "Our cluster is at 20 percent but pods will not schedule" is almost always this.

A pod without a memory limit takes the node down. An unbounded memory leak first triggers eviction of neighbours and then squeezes the kubelet itself. Do not allow a pod without a memory limit even in a single namespace; that is what LimitRange is for.

Changing a limit means a restart. The resources block is part of the pod template; editing it in a Deployment starts a new rollout. In-place resizing is arriving in recent releases but is not yet mature; do not build your capacity plan around it.

When not to follow this recipe

For workloads with strict latency targets, such as real-time media or a low-latency database, the "no CPU limit" advice backfires: on a saturated node, fair sharing makes latency unpredictable. For that class, Guaranteed plus the static CPU manager is the right path. Batch jobs are the opposite case: keeping the memory request low, the limit high, and accepting eviction is cheaper, because the job simply re-queues and nobody is waiting on it. The universal part of the recipe fits in one sentence: know which resource is compressible, and only relax the limit there.