← blog · September 22, 2026

Zero-Downtime Deployments in Kubernetes: Probes, preStop and Getting Graceful Shutdown in the Right Order

A rolling update alone does not give you zero downtime. Unless readiness and startup probes, a preStop delay, terminationGracePeriodSeconds, maxUnavailable and a PodDisruptionBudget are tuned together, every deploy leaves a short error spike. How to align the four pieces, and the pitfalls people hit.

When you update a Kubernetes Deployment, the documentation tells you a rolling update gives you zero downtime. In practice, every deploy produces a few seconds of 502s, connection resets or timeouts; a small error spike shows up in the load balancer logs and nobody can fully explain it. The problem is not the promise itself but the fact that it depends on four separate pieces being configured correctly: when the application says it is ready, when traffic stops being sent to it, how long it is allowed to shut down, and how many replicas may be gone at once. If any one of them is missing, deploys go cleanly "most of the time", which means they never go cleanly.

Where the errors come from

Two separate windows produce errors.

The startup window. A new pod receives traffic as soon as it reaches Running. If your application is still building its connection pool, warming a cache or JIT-compiling, the first requests are either rejected or answered very slowly. Without a readiness probe, the container starting and the pod being considered ready are the same instant.

The shutdown window. When an old pod starts being deleted, two things run in parallel: the kubelet sends SIGTERM to the container, and the endpoint controller removes the pod's address from the Service. Neither waits for the other. Propagating the endpoint list to kube-proxy, the Ingress controller and any external load balancer takes seconds. During those seconds the pod no longer accepts requests, but traffic is still routed to it. Most of the short error spike you see after a deploy comes from here.

This asynchronous behaviour is a design decision, not a bug. That is why the fix is not in the application alone but in aligning both sides with timing.

Describing startup correctly: three probes

Kubernetes offers three probes and they answer three different questions.

readinessProbe asks "can I take traffic right now". When it fails the pod is not restarted; it is only removed from the Service endpoints. That is where dependency checks belong: if you cannot reach the database, cutting traffic is right, killing the container is not.

livenessProbe asks "am I deadlocked", and failure restarts the container. The most common mistake is putting dependency checks into the liveness probe: the database goes down for five minutes, every pod enters a restart loop, and when the database comes back it is greeted with a connection storm. Liveness should measure only the process's own internal state and never ask the outside world.

startupProbe exists for slow-starting applications. When defined, the other two probes do not run until it succeeds. For a JVM service or one that loads a large model, giving the startup probe a long total budget is better than loosening the liveness threshold, because you keep catching post-startup deadlocks quickly.

containers:
  - name: api
    ports:
      - containerPort: 8080
    startupProbe:
      httpGet:
        path: /healthz
        port: 8080
      periodSeconds: 5
      failureThreshold: 36
    readinessProbe:
      httpGet:
        path: /ready
        port: 8080
      periodSeconds: 5
      failureThreshold: 2
    livenessProbe:
      httpGet:
        path: /healthz
        port: 8080
      periodSeconds: 10
      failureThreshold: 3

Here the startup probe tolerates 180 seconds in total; readiness cuts traffic after two consecutive failures; liveness only fires after 30 seconds of silence. Using two separate endpoints is deliberate: /healthz checks whether the process is alive, /ready asks the dependencies.

Ordering shutdown correctly

The goal during shutdown is this: SIGTERM should arrive after traffic has stopped, and after SIGTERM there should be enough time for in-flight requests to finish.

Step one: wait in preStop. When pod deletion begins, you insert an artificial delay between endpoint removal and SIGTERM. SIGTERM is not sent until the preStop hook completes; meanwhile the endpoint change propagates through the network layer, and the pod keeps answering the traffic it is still receiving.

lifecycle:
  preStop:
    exec:
      command: ["sh", "-c", "sleep 10"]
terminationGracePeriodSeconds: 45

If the image has no shell, the sleep action (preStop.sleep.seconds), beta in Kubernetes 1.30 and stable in later releases, does the same job without one; do not rely on it without checking your cluster version.

Step two: handle SIGTERM in the application. After preStop finishes, SIGTERM arrives and the application must stop accepting new connections, finish open requests, then exit. http.Server.Shutdown in Go, server.close in Node and Spring Boot's server.shutdown=graceful setting do this. If the application does not catch SIGTERM, the time preStop bought is wasted: the process dies immediately and open requests are cut off.

Step three: budget the total correctly. The terminationGracePeriodSeconds clock starts when preStop begins, not when it ends. So if preStop takes 10 seconds and your longest request takes 30, the total must be at least 40. The default of 30 seconds is often not enough for preStop plus the application's own shutdown; when it runs out, SIGKILL arrives and you see half-finished requests again.

How many replicas may be gone at once

The rolling update parameters decide how wide the disruption window is. The default maxUnavailable: 25% rounds down: with four replicas one pod goes away and a quarter of your capacity with it, and you only notice by doing the arithmetic. Writing an absolute number instead of a percentage makes the intent readable.

strategy:
  type: RollingUpdate
  rollingUpdate:
    maxSurge: 1
    maxUnavailable: 0

maxUnavailable: 0 means "bring up the new one first, then remove the old one", and it does nothing without a readiness probe, because that probe is where Kubernetes gets its "ready" decision. The minReadySeconds field additionally requires a new pod to stay ready for a period after first appearing ready; it catches applications that report ready and crash on the first request.

Disruptions outside a Deployment rollout need a PodDisruptionBudget. Node drains, cluster upgrades and autoscaler scale-downs do not look at the Deployment strategy; they look at the PDB.

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: api
spec:
  minAvailable: 2
  selector:
    matchLabels:
      app: api

Pitfalls

There is no zero downtime with a single replica. replicas: 1 with maxUnavailable: 0 technically works, but the old pod carries traffic until the new one is ready, and during the shutdown window nobody carries it. Nothing in this article is sufficient without at least two replicas.

HTTP keep-alive stretches shutdown. If the load balancer has long-lived connections open to the pod, endpoint removal does not close them. While the application stops accepting new requests on SIGTERM, it should send Connection: close on open connections or close idle ones. Otherwise the load balancer keeps sending requests down the same connection and gets resets.

Do not do heavy work in preStop. The hook exists only to buy time. If you put database cleanup or cache flushing in it, the hook gets longer and nobody sees it fail; a non-zero preStop exit code does not stop pod deletion, it only lands in the event log.

The readiness check must be cheap. If an endpoint called every five seconds runs a heavy query, the probe itself times out under load, the pod drops out of the endpoints, the remaining pods take more load and drop out too. This chain is known as a healthy application being taken down by its own probe, and it happens at every scale.

The Ingress controller has its own propagation delay. Even if the endpoint change takes seconds inside Kubernetes, some Ingress controllers reload configuration or update an external cloud load balancer through an API. That can exceed 10 seconds. Set the preStop duration by measuring, not guessing: watch the error count during a deploy and increase the delay until it reaches zero.

When it is not worth the effort

For batch jobs, queue consumers, or components called only by internal services with retry logic, most of these settings are unnecessary. For a queue consumer the only thing that matters is finishing and acknowledging the current message on SIGTERM; endpoint propagation delay is irrelevant to it. For user-facing HTTP services, however, "zero downtime" without all four pieces is not a measurement, it is a hope.