Kubernetes Resource Requests and Limits: CPU Throttling, OOMKilled and QoS Classes
A request is the scheduler's reservation, a limit is the kernel's quota. Why CPU gets paused while memory gets killed, how QoS classes decide eviction order, and how to derive the right numbers from measurement instead of guesswork.
Filling in the resources block of a pod looks trivial, but it encodes the two decisions that shape cluster health more than anything else: which node the scheduler will place this workload on, and who gets sacrificed first when that node runs short. Requests and limits sit on adjacent lines, yet they operate at different layers. A request is a reservation used only by the scheduler; nothing is physically set aside, but the sum of requests can never exceed a node's allocatable capacity. A limit is enforced by the kernel: a cgroup quota for CPU, a cgroup ceiling for memory. Teams that blur the two end up with either a cluster that looks "full" while idling or services that quietly slow down under load.
CPU and memory are not the same kind of resource
CPU is compressible. A container that exceeds its limit is not killed, it is paused. The kubelet translates a CPU limit into a Linux CFS quota: in every period, 100 ms by default, the container receives CPU time equal to its limit, and once the quota is spent the process is suspended until the period ends. A 500m limit means "50 ms out of every 100 ms." That sounds fine for a single-threaded workload, but a multi-threaded application running on eight cores at once burns through 50 ms of quota in about 6 ms and then sits idle for the remaining 94 ms of the period. Average CPU usage on the dashboard shows 30 percent while request latency multiplies. This is the classic signature of throttling, and most teams misdiagnose it as an application bug.
Memory is not compressible. When a container crosses its limit the kernel's OOM killer steps in, the process dies with exit code 137, and the pod status reads OOMKilled. There is no pause and no warning. The measured number is also not what you expect: the kubelet and kubectl top report the working set, which is RSS plus active page cache. A container that writes a lot of files appears far heavier than what its code actually holds.
QoS classes and eviction order
Kubernetes derives a quality of service class for each pod from the relationship between requests and limits:
Guaranteed: every container has requests equal to limits for both CPU and memory.
Burstable: at least one request or limit is set, but they are not all equal.
BestEffort: nothing is set at all.
When a node comes under memory pressure, the kubelet evicts BestEffort pods first, then Burstable pods that are using more than they requested; Guaranteed pods go last. Leaving a critical service as Burstable because "it has a request anyway" means accepting eviction caused by a neighbour's overflow even while you stay under your own limit. Nor can you use the full node: allocatable capacity is total capacity minus kubelet and system reservations minus the eviction threshold. The Allocatable line in kubectl describe node is the real budget, not the Capacity line.
A defensible starting point
Set memory request equal to memory limit. A limit is mandatory because memory overflow harms neighbours; if the request is lower than the limit, the pod is placed as if there were room and then pushes the node into pressure at its true consumption.
Set a CPU request, and add a CPU limit only when you have a reason. The request already provides fair sharing as a CFS weight: when the node is saturated everyone gets a share proportional to their request, and when it is idle everyone runs freely. A limit only makes sense to hard-cap a noisy neighbour or to qualify for the Guaranteed class.
If you need Guaranteed, set the CPU limit equal to the request and ask for whole cores; with the kubelet's static CPU manager policy, Guaranteed pods requesting integer CPUs get exclusive cores and quota throttling disappears entirely.
Two safety nets exist at namespace level. A LimitRange assigns defaults to pods that declare nothing, preventing them from falling into BestEffort; a ResourceQuota caps a team's total reservation.
Do not guess, measure. Run under real traffic for a week and look at the peak and 95th percentile of container_memory_working_set_bytes and rate(container_cpu_usage_seconds_total[5m]). Put the memory request near the peak and the CPU request at the 95th percentile. The Vertical Pod Autoscaler's recommendation mode (updateMode: "Off") does this measurement for you and writes the suggested values into its own status without touching the pod; running only in this mode for months before switching on automatic updates is a good habit.
To catch throttling, watch the ratio container_cpu_cfs_throttled_periods_total / container_cpu_cfs_periods_total. A container that stays above 25 percent either does not deserve a limit or has one that is far too low. For OOM, the kube-state-metrics series kube_pod_container_status_last_terminated_reason{reason="OOMKilled"} counts restarts by cause; looking at restart counts alone mixes probe-driven restarts with memory deaths.
Pitfalls
The runtime does not know the limit. The JVM reads the container limit, but by default it turns only a quarter of the memory limit into heap; unless you raise -XX:MaxRAMPercentage, you get a service that is granted 2 GiB and uses 512 MiB. In the opposite direction, the Go runtime before version 1.25 ignored the cgroup CPU quota entirely and picked GOMAXPROCS from the machine's full core count, which is exactly the throttling scenario above. In Go, the GOMEMLIMIT environment variable also lets the garbage collector see the memory limit; without it a service near its limit gets OOM-killed instead of slowing down.
Init containers change the arithmetic. A pod's effective request is the larger of the sum of its regular containers and the largest single init container. An init container that asks for 2 GiB for setup means a 2 GiB reservation on that node for as long as the pod lives, even if your application uses 256 MiB.
High requests, empty cluster. If requests sit far above real usage, the scheduler considers nodes full, new pods wait in Pending, and the cluster autoscaler spins up machines nobody needs. "Our cluster is at 20 percent but pods will not schedule" is almost always this.
A pod without a memory limit takes the node down. An unbounded memory leak first triggers eviction of neighbours and then squeezes the kubelet itself. Do not allow a pod without a memory limit even in a single namespace; that is what LimitRange is for.
Changing a limit means a restart. The resources block is part of the pod template; editing it in a Deployment starts a new rollout. In-place resizing is arriving in recent releases but is not yet mature; do not build your capacity plan around it.
When not to follow this recipe
For workloads with strict latency targets, such as real-time media or a low-latency database, the "no CPU limit" advice backfires: on a saturated node, fair sharing makes latency unpredictable. For that class, Guaranteed plus the static CPU manager is the right path. Batch jobs are the opposite case: keeping the memory request low, the limit high, and accepting eviction is cheaper, because the job simply re-queues and nobody is waiting on it. The universal part of the recipe fits in one sentence: know which resource is compressible, and only relax the limit there.
[OK] 25 records · 8 featured · 1 retainer · 16 standard
/02 services
services.tree
> note
· Prices are in USD. Volume discounts of 7% for 6-12 day projects and 14% for 13+ day projects apply. Programs of 100+ person-days are priced individually. Prices are negotiable based on scope, urgency and long-term collaboration.
> ls ~/products/ --open-source
[OK] 2 entries · 1 live · 1 pre-release
/03 products
products.list
liveopen source · mit/prod/01
filex
the self-hosted file manager that embeds anywhere
single go binary · vue/react/web-component embed · 5 storage drivers · realtime collab · rbac · native multi-tenancy · mcp server for ai agents
We build infrastructure, automate everything, and keep systems running.
BRF Tech is a Bursa-based DevSecOps and software consulting firm. We operate across 25 service areas, from Kubernetes cluster management to serverless platform setup, CI/CD pipeline design to AI agent development.
What we do is simple: we set up your systems, automate them, and make sure they won't wake you up at 3 AM. We codify your infrastructure with Terraform, move your deployments to GitOps with ArgoCD, and monitor everything with Prometheus. If something breaks — we intervene before it does.
We're against vendor lock-in. We work with open-source tools, self-hosted solutions, and industry-standard technologies. We build your own serverless platform on Knative, isolate with Kata Containers, manage your secrets with Vault. Every project is delivered with clear scope, clear timeline, clear pricing.
DevSecOps & CI/CD
Kubernetes & Serverless
Infrastructure as Code
AI & Agent (MCP/ACP)
Security & Zero Trust
Full-Stack Development
Self-Hosted Solutions
> grep -i question ~/faq.md
[OK] 3 entries · click to expand
[?]What services are included in BRF Tech's DevOps solutions?▾
We offer CI/CD pipeline setup, Kubernetes cluster management, container migration, infrastructure automation with Terraform, monitoring & alerting, in-house technology installations and DevOps consulting services.
[?]What are your software development services?▾
We offer MVP development, existing software performance optimization, backend API development (Node.js, Go, Python) and full-stack application development. Every project is delivered with minimal technical debt.
[?]Which DevOps tools do you work with?▾
From Kubernetes and Docker to Talos Linux, from Terraform and OpenTofu to Ansible, and across Jenkins, GitLab CI/CD, GitHub Actions, ArgoCD, Flux, Helm and Rancher, we work with industry-standard tooling. For observability and error tracking we use Prometheus, Grafana, Loki, Tempo, OpenTelemetry and GlitchTip; for networking, Cilium and Envoy; for test automation, Playwright, Cypress, k6, Locust, Testcontainers and Trivy. We determine the most suitable toolset together, based on your project’s needs.
Automated build, test and deployment pipelines with GitHub Actions, GitLab CI/CD.
Automated build, test and deploy pipeline setup for existing or new repositories. GitHub Actions, GitLab CI/CD or preferred tool is used. Separate workflows are defined for staging + production environments.
Container architecture migration, cluster setup and orchestration.
Production-ready Kubernetes cluster setup on bare-metal or cloud. Includes Helm charts, Ingress, TLS certificate management, namespace isolation and RBAC.
Migration of existing applications to container architecture.
Migration of existing applications to container architecture. Docker image design, Compose configuration and conversion to Kubernetes manifests. Kata Containers / gVisor can be included for advanced isolation.
In-house tools like GitLab, Nextcloud, VPN, mail server.
Installation of tools like GitLab, Mattermost, Nextcloud, mail server, VPN, internal monitoring on company servers. Deployed on Docker Compose or Kubernetes.
DB + cache + code bottleneck identification and resolution.
Profiling of existing application, identification and resolution of bottlenecks. DB query optimization, caching layer (Redis), service-level improvements.
MCP/ACP agents, RAG pipelines, multi-agent orchestration, model serving.
LLM integration (OpenAI, Claude, Gemini, local models), tool-augmented AI agents with MCP (Model Context Protocol) and ACP (Agent Communication Protocol). RAG pipeline, multi-agent orchestration, model serving (vLLM/Ollama). AI layer for existing business processes.
Event-driven data pipeline with Apache Kafka, RabbitMQ, NATS. ETL/ELT orchestration with Apache Airflow, real-time data streaming, CDC (Change Data Capture). Data lake/warehouse design, schema registry, dead letter queue management.
Packages
Kafka/RabbitMQ setup + basic producer/consumer
$3,500
5–7 person-day
ETL pipeline (Airflow + source → warehouse)
$5,200
8–12 person-day
Full event-driven architecture (CDC + streaming + DLQ)
Hands-on technical training for your teams. Docker & container fundamentals, Kubernetes operations, Terraform IaC, CI/CD best practices, AI/LLM integration workshops. Practical exercises in live lab environments, content customized by skill level.
E2E test automation with Playwright, Cypress, load testing with k6/Gatling. Pre-deploy quality gate integrated into CI/CD pipeline, test coverage reporting, visual regression testing. Reduce test writing time by up to 60% with AI-assisted test generation.
Your Collected Personal Data, Collection Method and Legal Basis
Any information that identifies or makes you identifiable is considered "personal data." When you visit our website or use our services, contact information such as your name, surname, email address, phone number, as well as your IP address and browser cookie data may be collected through automatic or semi-automatic means. This data is processed based on the legal grounds specified in Article 5 of the Personal Data Protection Law No. 6698: "being directly related to the establishment or performance of a contract" and "being mandatory for the legitimate interests of the data controller, provided that it does not harm the fundamental rights and freedoms of the data subject."
Purpose of Processing Your Personal Data
Your collected personal data is processed for the purposes of providing and improving our services, fulfilling customer requests, meeting legal obligations, conducting information security processes and managing communication activities. Your data is processed in a limited and proportionate manner for the stated purposes; when the purpose ceases to exist, data is deleted, destroyed or anonymized.
To Whom and For What Purposes Collected Personal Data May Be Transferred
Your personal data may be transferred to public institutions and organizations as required by legal regulations, to our business partners and technical infrastructure providers for the purpose of delivering services, and to lawyers and consultants in legal disputes. Transfers are carried out in accordance with Articles 8 and 9 of the Law, with necessary technical and administrative measures in place.
Your Rights as a Data Subject
In accordance with Article 11 of Law No. 6698, you have the right to: learn whether your personal data is being processed; request information if it has been processed; learn the purpose of processing and whether it is used in accordance with its purpose; know the third parties to whom it is transferred domestically or abroad; request correction if it has been processed incompletely or incorrectly; request deletion or destruction within the framework of conditions set out in Article 7 of the Law; request that the operations carried out be notified to third parties to whom data has been transferred; object to any adverse result arising from the analysis of data exclusively through automated systems; and claim compensation for damages in case of unlawful processing. You may contact us through the communication channels on our website to exercise these rights.
Terms & Conditions
Terms & Conditions
1. Service Scope and Changes
BRF Tech conducts its software development, DevOps consulting, infrastructure setup and technical support services within the framework of these terms and conditions. The scope of service is determined separately for each project and finalized through mutual agreement. BRF Tech reserves the right to update the scope, pricing and technical details of its services, provided that prior notice is given.
2. Privacy Policy
Personal data collected within the scope of our services is processed in accordance with our Privacy Policy. For detailed information, please review our Privacy Policy page.
3. Service Usage and Responsibilities
Clients agree to use the provided services solely within legal and ethical boundaries. BRF Tech reserves the right to suspend or terminate services in the event of misuse, violation of third-party rights, or use for illegal activities. The client is responsible for the accuracy and currency of the information provided within the project scope.
4. Payment Terms
Payment terms are determined on a project basis and finalized through mutual agreement at the start of the project. Unless otherwise specified, invoices are payable within 15 days of project delivery. A monthly late payment interest of 2% may apply to overdue payments.
5. Cancellation and Refund Policy
Cancellation requests must be submitted in writing at least 7 days before the service start date. For projects already in progress, billing is based on the proportion of completed work. No refunds are issued for software and infrastructure components specifically developed within the project scope.