← blog · October 4, 2026

PodDisruptionBudgets and Node Drains in Kubernetes: Writing the Budget and Why Maintenance Gets Stuck

A PodDisruptionBudget only limits voluntary disruptions, and a badly written one either protects nothing or blocks node drains forever. Choosing between minAvailable and maxUnavailable, the unhealthy pod eviction policy, and the pitfalls of draining nodes.

Nodes in a Kubernetes cluster get drained all the time: for OS and kubelet upgrades, when the autoscaler removes spare capacity, when hardware is replaced. Every one of these is a voluntary disruption, meaning whoever operates the cluster is deliberately moving a pod. A PodDisruptionBudget (PDB) exists for exactly these moments and states "no more than this many of my pods may be gone at once". It is a tiny object, yet a badly written one produces one of two outcomes: it protects nothing, or it blocks every maintenance operation indefinitely.

What a PDB protects, and what it does not

A PDB only limits requests that go through the Eviction API. kubectl drain, the cluster autoscaler and node pool upgrades in managed services all use that API. When the budget does not allow an eviction, the API answers with HTTP 429 and the caller retries later.

What it does not cover matters just as much:

  • Involuntary disruptions such as a node crash, a kernel panic or a process killed for running out of memory. A PDB cannot prevent them; they are simply counted against the budget, which makes subsequent voluntary evictions more cautious.
  • Deleting a pod directly with kubectl delete pod. Deletion does not go through the Eviction API.
  • A Deployment's own rolling update. How many pods go away during a rollout is decided by maxUnavailable and maxSurge on the Deployment, not by the PDB.

Teams often assume "we added a PDB, so we have no downtime". All a PDB really does is make the people operating the cluster respect your application's limits.

minAvailable or maxUnavailable

There are two fields and a single PDB may use only one of them. Both accept an integer or a percentage.

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: api
spec:
  maxUnavailable: 1
  selector:
    matchLabels:
      app: api

In most cases I prefer maxUnavailable. The reason: minAvailable: 2 is independent of the replica count. If the application is later scaled down to two replicas, the budget silently becomes zero allowed disruptions and drains start hanging. maxUnavailable: 1 keeps allowing one pod to leave no matter how the replica count changes, so it stays correct as you scale.

minAvailable is the right choice when there is a hard floor. A three member quorum system needs at least two members up, and minAvailable: 2 expresses that intent most clearly. Even so, maxUnavailable: 1 gives the same result for three members and stays sensible when you grow to five.

With percentages, rounding on small replica counts can surprise you. Rather than doing the math in your head, apply the PDB and look at the ALLOWED DISRUPTIONS column of kubectl get pdb. If it reads zero, a drain will wait on that application.

The drain sequence

Manual node maintenance usually looks like this:

kubectl cordon node-a
kubectl drain node-a --ignore-daemonsets --delete-emptydir-data --timeout=15m
# maintenance
kubectl uncordon node-a

cordon stops new pods from being scheduled on the node. drain does the same and then evicts pods. DaemonSet pods are skipped because they run on every node anyway, and for pods using emptyDir you have to explicitly accept that the data is lost. If there are bare pods not owned by any controller, drain stops, because nothing would recreate them after eviction. --force overrides that check, and the pod is then really gone.

Without --timeout, drain waits for the budget forever. In automation always set a limit and treat hitting it as a failure; a script that quietly moves on to the next node accumulates half finished maintenance.

Pitfalls

A PDB on a single replica. A one replica Deployment guarded by minAvailable: 1 or maxUnavailable: 0 can never be evicted, and the drain waits on it forever. Managed services differ here: some ignore the budget after a grace period and delete the pod anyway, others mark the upgrade as failed. Neither is what you wanted. A single replica application already accepts downtime; instead of hiding that behind a PDB, either add replicas or skip the PDB and accept the brief outage.

Unhealthy pods lock the budget. A PDB counts pods that are Ready. If one pod of an application is crash looping, the budget already looks exhausted and the healthy pods cannot be evicted. Worse, the broken pod cannot be evicted either. unhealthyPodEvictionPolicy: AlwaysAllow, enabled by default since Kubernetes 1.27, lets pods that are not ready be evicted regardless of the budget. For most stateless services this is the right setting, since protecting a broken pod helps nobody.

spec:
  maxUnavailable: 1
  unhealthyPodEvictionPolicy: AlwaysAllow
  selector:
    matchLabels:
      app: api

Readiness time slows every drain. The budget does not open again until the replacement pod is Ready. For applications that take minutes to start, warm a cache or load a large dataset, upgrading a twenty node pool can take hours. Make sure the readiness probe measures actual readiness to serve traffic and is not padded with needless delays.

The selector does not match the real pods. A PDB whose selector matches nothing raises no error; it just protects nothing. This happens quietly when labels are renamed. The opposite is also a problem: if a pod falls under more than one PDB, the Eviction API refuses to evict it and the drain gets stuck. Keep one PDB per application with a selector identical to the Deployment's selector.

All replicas on one node. A PDB says how many pods may leave at once, not where they run. If all three replicas sit on the same node and that node dies, the PDB can do nothing. Spread replicas across nodes, and across zones where possible, with topologySpreadConstraints. Budget and spreading work together; either one alone is incomplete.

No spare capacity, no open budget. If there is nowhere for the evicted pods to go, the replacements stay Pending, never become Ready, and the budget stays closed. For node upgrades, add the new node first and drain the old one afterwards. Most managed services expose this as a surge setting.

When not to use a PDB

Single replica workloads that tolerate downtime, batch Jobs and development environments usually gain nothing from a PDB and pay for it in stuck maintenance. For batch work, eviction restarts the job, but protecting it with a PDB ties node maintenance to the job's runtime; making the job safely restartable is the sturdier fix.

Likewise, in a cluster where node pool upgrades are automated, letting every team write arbitrary PDBs means a single bad budget can halt upgrades for the whole cluster. The practical guard is an admission controller or policy engine that rejects locking values such as maxUnavailable: 0 and refuses PDBs on single replica workloads.

A short checklist

For each stateless service, a reasonable baseline is at least two replicas, maxUnavailable: 1, unhealthyPodEvictionPolicy: AlwaysAllow and a placement rule that spreads replicas across nodes. For stateful services that need a quorum, derive the budget from the quorum math and define readiness so that it proves the member has actually rejoined the cluster. Finally, test all of this by draining a node on an ordinary day, not on maintenance day. Every application showing zero in ALLOWED DISRUPTIONS is the one that will keep you waiting when it counts.