← blog · September 15, 2026

Autoscaling Pods in Kubernetes: Choosing Between HPA, VPA, and KEDA

HPA scales replica count, VPA scales resource sizing, KEDA scales event-driven load. Using one in place of another is the root cause of most autoscaling complaints.

The problem: a single metric runs out of road fast

Deciding how many pod replicas a service needs by hand works fine as long as traffic is flat. The moment traffic fluctuates, you either pay for idle capacity around the clock or hit a wall during peak hours. Kubernetes autoscaling tools exist to make that decision for you based on live metrics, but "autoscaling" is not one tool. It is three mechanisms that solve three different problems: the Horizontal Pod Autoscaler (HPA) changes the number of replicas, the Vertical Pod Autoscaler (VPA) changes the resources allocated to each replica, and KEDA scales event-driven workloads based on queue depth or other external signals. Trying to use one in place of another is the root cause of most autoscaling complaints.

HPA: horizontal scaling based on CPU and memory

HPA is defined under the autoscaling/v2 API group, and its default behavior is to raise or lower replica count based on CPU utilization. A basic definition looks like this:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: api-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: api
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70

This is simple, but it depends on three things: metrics-server running in the cluster, every pod having a correctly filled resources.requests field (HPA computes utilization percentage against that request), and a custom metrics adapter such as Prometheus Adapter if you need anything beyond CPU or memory. When requests is left empty or set far from actual usage, the percentage HPA computes becomes meaningless. This is the most common HPA pitfall in practice, and it is not that HPA "doesn't work," it is that it works correctly on a bad input.

The setting nobody reads: behavior and flapping

The least-known and most trap-prone part of HPA is the behavior field. Left at its defaults, an HPA reacts to a short CPU spike by adding replicas immediately, then removes them just as fast once the spike passes. This is known as flapping, and in services whose requests rely on connection pools it causes connections to be rebuilt on every scaling event, which shows up as a latency spike. behavior.scaleDown.stabilizationWindowSeconds lets you delay scale-down decisions by looking at the highest demand over the last N seconds, while behavior.scaleUp lets you cap how many replicas can be added at once during a sudden surge:

  behavior:
    scaleDown:
      stabilizationWindowSeconds: 300
    scaleUp:
      stabilizationWindowSeconds: 0

Leaving this untouched is the most common reason behind "HPA is behaving erratically" complaints; the problem is not HPA itself, it is that the default window does not match your traffic pattern. The same section also supports multiple metrics at once, for example CPU alongside per-request latency. In that case HPA applies whichever metric recommends the highest replica count, not the lowest. Adding a latency metric on top of CPU without knowing this produces a system that scales up earlier than expected but never scales up later than expected.

VPA: getting the size right, at a cost

While HPA changes replica count, in some workloads the real problem is not how many replicas run but how much CPU and memory each one was given in the first place. VPA derives that estimate from historical usage and either recommends or automatically applies resources.requests/limits:

apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: api-vpa
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: api
  updatePolicy:
    updateMode: "Auto"

The updateMode: "Auto" option is historically the most contentious part of VPA: the traditional VPA implementation changes a resource request by restarting the pod. A pod restarting at an arbitrary moment causes a brief outage for a single-replica workload, and for a multi-replica one, restarting several pods around the same time without a correctly configured PodDisruptionBudget temporarily reduces service capacity. That is why a safer starting point in production is running VPA in "Off" mode first, where it only produces recommendations without applying them, and reviewing those numbers manually before turning automatic updates on. Running VPA and HPA on the same resource metric (CPU) at the same time is also discouraged: both try to react to the same signal by changing either the request or the replica count, and they can end up undoing each other's decisions in a loop.

KEDA: event-driven workloads and scaling to zero

HPA and VPA were designed for services that run continuously and whose load is measurable through resource consumption. A consumer that processes messages off a queue does not burn CPU while the queue is empty, so an HPA watching CPU never scales this workload correctly, because it is looking at the wrong signal entirely. KEDA fills that gap: it drives replica count from external signals such as queue depth, a messaging system's own metric, a cron schedule, or a Prometheus query, and it can scale a workload down to zero replicas when there is no work at all.

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: worker-scaler
spec:
  scaleTargetRef:
    name: worker
  minReplicaCount: 0
  maxReplicaCount: 20
  triggers:
    - type: rabbitmq
      metadata:
        queueName: jobs
        mode: QueueLength
        value: "5"

Scaling to zero is attractive because you stop paying for idle capacity, but it has a cost: bringing the first pod up from zero (cold start) can take seconds, and longer still if the image is large or startup runs a heavy initialization step. For background jobs that can tolerate a delay, scaling to zero makes sense. For a path where a user is waiting on an immediate response, keeping minReplicaCount at one or higher avoids paying that latency back with interest on what you saved in cost.

Which one, and when

For continuously running, request-driven services whose CPU or memory usage genuinely reflects load, HPA remains the right default. It is simple to set up and it is the best-tested path in the ecosystem. If you are unsure whether resource requests are sized correctly, run VPA in recommendation-only mode first and decide based on real usage data before moving to automatic application in production. For workloads triggered by a queue, a messaging system, a schedule, or any external metric, prefer KEDA; it does not replace HPA so much as add a metric source on top of it, since KEDA creates its own HPA object behind the scenes once installed.

When to skip autoscaling entirely

Installing autoscaling for a service whose load is nearly flat throughout the day, whose traffic spikes are known in advance (a batch job that fires at a fixed hour, for instance), or that a single replica handles comfortably, adds a layer of complexity that looks like it is solving a problem you do not actually have. In those cases, a fixed replicas count and, if needed, a simple scheduled kubectl scale command cost less than debugging the interaction of three separate controllers. The cost of autoscaling is not only CPU and memory, it is a behavioral layer added to the system: if metrics lag, if the stabilization window is misconfigured, or if a replica scales up and back down erratically, someone needs to understand that behavior. For a small, predictable load, that person's time is worth more than the automation.