Autoscaling Pods in Kubernetes: Choosing Between HPA, VPA, and KEDA
HPA scales replica count, VPA scales resource sizing, KEDA scales event-driven load. Using one in place of another is the root cause of most autoscaling complaints.
The problem: a single metric runs out of road fast
Deciding how many pod replicas a service needs by hand works fine as long as traffic is flat. The moment traffic fluctuates, you either pay for idle capacity around the clock or hit a wall during peak hours. Kubernetes autoscaling tools exist to make that decision for you based on live metrics, but "autoscaling" is not one tool. It is three mechanisms that solve three different problems: the Horizontal Pod Autoscaler (HPA) changes the number of replicas, the Vertical Pod Autoscaler (VPA) changes the resources allocated to each replica, and KEDA scales event-driven workloads based on queue depth or other external signals. Trying to use one in place of another is the root cause of most autoscaling complaints.
HPA: horizontal scaling based on CPU and memory
HPA is defined under the autoscaling/v2 API group, and its default behavior is to raise or lower replica count based on CPU utilization. A basic definition looks like this:
This is simple, but it depends on three things: metrics-server running in the cluster, every pod having a correctly filled resources.requests field (HPA computes utilization percentage against that request), and a custom metrics adapter such as Prometheus Adapter if you need anything beyond CPU or memory. When requests is left empty or set far from actual usage, the percentage HPA computes becomes meaningless. This is the most common HPA pitfall in practice, and it is not that HPA "doesn't work," it is that it works correctly on a bad input.
The setting nobody reads: behavior and flapping
The least-known and most trap-prone part of HPA is the behavior field. Left at its defaults, an HPA reacts to a short CPU spike by adding replicas immediately, then removes them just as fast once the spike passes. This is known as flapping, and in services whose requests rely on connection pools it causes connections to be rebuilt on every scaling event, which shows up as a latency spike. behavior.scaleDown.stabilizationWindowSeconds lets you delay scale-down decisions by looking at the highest demand over the last N seconds, while behavior.scaleUp lets you cap how many replicas can be added at once during a sudden surge:
Leaving this untouched is the most common reason behind "HPA is behaving erratically" complaints; the problem is not HPA itself, it is that the default window does not match your traffic pattern. The same section also supports multiple metrics at once, for example CPU alongside per-request latency. In that case HPA applies whichever metric recommends the highest replica count, not the lowest. Adding a latency metric on top of CPU without knowing this produces a system that scales up earlier than expected but never scales up later than expected.
VPA: getting the size right, at a cost
While HPA changes replica count, in some workloads the real problem is not how many replicas run but how much CPU and memory each one was given in the first place. VPA derives that estimate from historical usage and either recommends or automatically applies resources.requests/limits:
The updateMode: "Auto" option is historically the most contentious part of VPA: the traditional VPA implementation changes a resource request by restarting the pod. A pod restarting at an arbitrary moment causes a brief outage for a single-replica workload, and for a multi-replica one, restarting several pods around the same time without a correctly configured PodDisruptionBudget temporarily reduces service capacity. That is why a safer starting point in production is running VPA in "Off" mode first, where it only produces recommendations without applying them, and reviewing those numbers manually before turning automatic updates on. Running VPA and HPA on the same resource metric (CPU) at the same time is also discouraged: both try to react to the same signal by changing either the request or the replica count, and they can end up undoing each other's decisions in a loop.
KEDA: event-driven workloads and scaling to zero
HPA and VPA were designed for services that run continuously and whose load is measurable through resource consumption. A consumer that processes messages off a queue does not burn CPU while the queue is empty, so an HPA watching CPU never scales this workload correctly, because it is looking at the wrong signal entirely. KEDA fills that gap: it drives replica count from external signals such as queue depth, a messaging system's own metric, a cron schedule, or a Prometheus query, and it can scale a workload down to zero replicas when there is no work at all.
Scaling to zero is attractive because you stop paying for idle capacity, but it has a cost: bringing the first pod up from zero (cold start) can take seconds, and longer still if the image is large or startup runs a heavy initialization step. For background jobs that can tolerate a delay, scaling to zero makes sense. For a path where a user is waiting on an immediate response, keeping minReplicaCount at one or higher avoids paying that latency back with interest on what you saved in cost.
Which one, and when
For continuously running, request-driven services whose CPU or memory usage genuinely reflects load, HPA remains the right default. It is simple to set up and it is the best-tested path in the ecosystem. If you are unsure whether resource requests are sized correctly, run VPA in recommendation-only mode first and decide based on real usage data before moving to automatic application in production. For workloads triggered by a queue, a messaging system, a schedule, or any external metric, prefer KEDA; it does not replace HPA so much as add a metric source on top of it, since KEDA creates its own HPA object behind the scenes once installed.
When to skip autoscaling entirely
Installing autoscaling for a service whose load is nearly flat throughout the day, whose traffic spikes are known in advance (a batch job that fires at a fixed hour, for instance), or that a single replica handles comfortably, adds a layer of complexity that looks like it is solving a problem you do not actually have. In those cases, a fixed replicas count and, if needed, a simple scheduled kubectl scale command cost less than debugging the interaction of three separate controllers. The cost of autoscaling is not only CPU and memory, it is a behavioral layer added to the system: if metrics lag, if the stabilization window is misconfigured, or if a replica scales up and back down erratically, someone needs to understand that behavior. For a small, predictable load, that person's time is worth more than the automation.
[OK] 25 records · 8 featured · 1 retainer · 16 standard
/02 services
services.tree
> note
· Prices are in USD. Volume discounts of 7% for 6-12 day projects and 14% for 13+ day projects apply. Programs of 100+ person-days are priced individually. Prices are negotiable based on scope, urgency and long-term collaboration.
> ls ~/products/ --open-source
[OK] 2 entries · 1 live · 1 pre-release
/03 products
products.list
liveopen source · mit/prod/01
filex
the self-hosted file manager that embeds anywhere
single go binary · vue/react/web-component embed · 5 storage drivers · realtime collab · rbac · native multi-tenancy · mcp server for ai agents
We build infrastructure, automate everything, and keep systems running.
BRF Tech is a Bursa-based DevSecOps and software consulting firm. We operate across 25 service areas, from Kubernetes cluster management to serverless platform setup, CI/CD pipeline design to AI agent development.
What we do is simple: we set up your systems, automate them, and make sure they won't wake you up at 3 AM. We codify your infrastructure with Terraform, move your deployments to GitOps with ArgoCD, and monitor everything with Prometheus. If something breaks — we intervene before it does.
We're against vendor lock-in. We work with open-source tools, self-hosted solutions, and industry-standard technologies. We build your own serverless platform on Knative, isolate with Kata Containers, manage your secrets with Vault. Every project is delivered with clear scope, clear timeline, clear pricing.
DevSecOps & CI/CD
Kubernetes & Serverless
Infrastructure as Code
AI & Agent (MCP/ACP)
Security & Zero Trust
Full-Stack Development
Self-Hosted Solutions
> grep -i question ~/faq.md
[OK] 3 entries · click to expand
[?]What services are included in BRF Tech's DevOps solutions?▾
We offer CI/CD pipeline setup, Kubernetes cluster management, container migration, infrastructure automation with Terraform, monitoring & alerting, in-house technology installations and DevOps consulting services.
[?]What are your software development services?▾
We offer MVP development, existing software performance optimization, backend API development (Node.js, Go, Python) and full-stack application development. Every project is delivered with minimal technical debt.
[?]Which DevOps tools do you work with?▾
From Kubernetes and Docker to Talos Linux, from Terraform and OpenTofu to Ansible, and across Jenkins, GitLab CI/CD, GitHub Actions, ArgoCD, Flux, Helm and Rancher, we work with industry-standard tooling. For observability and error tracking we use Prometheus, Grafana, Loki, Tempo, OpenTelemetry and GlitchTip; for networking, Cilium and Envoy; for test automation, Playwright, Cypress, k6, Locust, Testcontainers and Trivy. We determine the most suitable toolset together, based on your project’s needs.
Automated build, test and deployment pipelines with GitHub Actions, GitLab CI/CD.
Automated build, test and deploy pipeline setup for existing or new repositories. GitHub Actions, GitLab CI/CD or preferred tool is used. Separate workflows are defined for staging + production environments.
Container architecture migration, cluster setup and orchestration.
Production-ready Kubernetes cluster setup on bare-metal or cloud. Includes Helm charts, Ingress, TLS certificate management, namespace isolation and RBAC.
Migration of existing applications to container architecture.
Migration of existing applications to container architecture. Docker image design, Compose configuration and conversion to Kubernetes manifests. Kata Containers / gVisor can be included for advanced isolation.
In-house tools like GitLab, Nextcloud, VPN, mail server.
Installation of tools like GitLab, Mattermost, Nextcloud, mail server, VPN, internal monitoring on company servers. Deployed on Docker Compose or Kubernetes.
DB + cache + code bottleneck identification and resolution.
Profiling of existing application, identification and resolution of bottlenecks. DB query optimization, caching layer (Redis), service-level improvements.
MCP/ACP agents, RAG pipelines, multi-agent orchestration, model serving.
LLM integration (OpenAI, Claude, Gemini, local models), tool-augmented AI agents with MCP (Model Context Protocol) and ACP (Agent Communication Protocol). RAG pipeline, multi-agent orchestration, model serving (vLLM/Ollama). AI layer for existing business processes.
Event-driven data pipeline with Apache Kafka, RabbitMQ, NATS. ETL/ELT orchestration with Apache Airflow, real-time data streaming, CDC (Change Data Capture). Data lake/warehouse design, schema registry, dead letter queue management.
Packages
Kafka/RabbitMQ setup + basic producer/consumer
$3,500
5–7 person-day
ETL pipeline (Airflow + source → warehouse)
$5,200
8–12 person-day
Full event-driven architecture (CDC + streaming + DLQ)
Hands-on technical training for your teams. Docker & container fundamentals, Kubernetes operations, Terraform IaC, CI/CD best practices, AI/LLM integration workshops. Practical exercises in live lab environments, content customized by skill level.
E2E test automation with Playwright, Cypress, load testing with k6/Gatling. Pre-deploy quality gate integrated into CI/CD pipeline, test coverage reporting, visual regression testing. Reduce test writing time by up to 60% with AI-assisted test generation.
Your Collected Personal Data, Collection Method and Legal Basis
Any information that identifies or makes you identifiable is considered "personal data." When you visit our website or use our services, contact information such as your name, surname, email address, phone number, as well as your IP address and browser cookie data may be collected through automatic or semi-automatic means. This data is processed based on the legal grounds specified in Article 5 of the Personal Data Protection Law No. 6698: "being directly related to the establishment or performance of a contract" and "being mandatory for the legitimate interests of the data controller, provided that it does not harm the fundamental rights and freedoms of the data subject."
Purpose of Processing Your Personal Data
Your collected personal data is processed for the purposes of providing and improving our services, fulfilling customer requests, meeting legal obligations, conducting information security processes and managing communication activities. Your data is processed in a limited and proportionate manner for the stated purposes; when the purpose ceases to exist, data is deleted, destroyed or anonymized.
To Whom and For What Purposes Collected Personal Data May Be Transferred
Your personal data may be transferred to public institutions and organizations as required by legal regulations, to our business partners and technical infrastructure providers for the purpose of delivering services, and to lawyers and consultants in legal disputes. Transfers are carried out in accordance with Articles 8 and 9 of the Law, with necessary technical and administrative measures in place.
Your Rights as a Data Subject
In accordance with Article 11 of Law No. 6698, you have the right to: learn whether your personal data is being processed; request information if it has been processed; learn the purpose of processing and whether it is used in accordance with its purpose; know the third parties to whom it is transferred domestically or abroad; request correction if it has been processed incompletely or incorrectly; request deletion or destruction within the framework of conditions set out in Article 7 of the Law; request that the operations carried out be notified to third parties to whom data has been transferred; object to any adverse result arising from the analysis of data exclusively through automated systems; and claim compensation for damages in case of unlawful processing. You may contact us through the communication channels on our website to exercise these rights.
Terms & Conditions
Terms & Conditions
1. Service Scope and Changes
BRF Tech conducts its software development, DevOps consulting, infrastructure setup and technical support services within the framework of these terms and conditions. The scope of service is determined separately for each project and finalized through mutual agreement. BRF Tech reserves the right to update the scope, pricing and technical details of its services, provided that prior notice is given.
2. Privacy Policy
Personal data collected within the scope of our services is processed in accordance with our Privacy Policy. For detailed information, please review our Privacy Policy page.
3. Service Usage and Responsibilities
Clients agree to use the provided services solely within legal and ethical boundaries. BRF Tech reserves the right to suspend or terminate services in the event of misuse, violation of third-party rights, or use for illegal activities. The client is responsible for the accuracy and currency of the information provided within the project scope.
4. Payment Terms
Payment terms are determined on a project basis and finalized through mutual agreement at the start of the project. Unless otherwise specified, invoices are payable within 15 days of project delivery. A monthly late payment interest of 2% may apply to overdue payments.
5. Cancellation and Refund Policy
Cancellation requests must be submitted in writing at least 7 days before the service start date. For projects already in progress, billing is based on the proportion of completed work. No refunds are issued for software and infrastructure components specifically developed within the project scope.