← blog · September 16, 2026

The PID 1 Problem in Containers: Zombie Processes and When You Need an Init System

The first process in a container is not an ordinary process as far as the kernel is concerned. Here is how that difference leads to zombie process buildup, how it eventually hits the cgroup pid limit, and how to choose between tini, dumb-init, and a full process manager.

PID 1 means something different to the kernel

When a container starts, the first process inside it becomes PID 1 as far as the Linux kernel is concerned, and PID 1 carries two responsibilities an ordinary process does not have. The first is signal handling: when the kernel delivers a signal like SIGTERM or SIGINT to a process that has not installed its own handler, it normally applies a default handler, but that rule does not apply to PID 1. An ordinary process that receives SIGTERM terminates by default; PID 1's default behavior is to do nothing at all. The second is reaping: when a child process exits, the kernel keeps it around as a zombie (defunct) until its parent calls wait() on it and collects the exit code. In a normal process tree, orphaned children get reparented to PID 1, and init systems such as systemd or SysV init exist precisely to reap that inheritance. PID 1 inside a container is usually a web server, an application process, or a shell script, none of which were written for that job.

Both responsibilities look ignorable because most containers run a single process and never fork children. The problem shows up exactly where a child process does get forked but is never properly reaped.

How zombies accumulate

The typical accumulation pattern looks like this: the application shells out to run a periodic task in the background, the shell backgrounds the command with & and exits immediately. When the shell exits, the process it left behind is orphaned and reparented to PID 1. If PID 1 never calls wait() on it, the process sits there as a zombie once it finishes. A one-off background task will not make this visible; a scheduled job that runs every minute or dozens of times an hour accumulates them fast.

Zombies consume neither memory nor CPU, they show up as <defunct> in ps output, and they look harmless. The actual limit is Linux's cgroup pids controller: every zombie still occupies a PID slot, and once you approach the pids.max limit, new process creation (fork()) starts failing. The failure then surfaces somewhere completely unrelated to the zombies themselves: a new command cannot run inside the container, a new request handler cannot be forked, a scheduler silently stops. A liveness probe usually checks whether the main process is still up, and it stays green as long as that process is alive; the only signs of trouble are a rising pids.current in cgroup metrics or a single fork: Resource temporarily unavailable line buried in the logs.

Three approaches

There are three ways to deal with this, and the right one depends on what the container actually runs.

The first is using a lightweight init process. tini and dumb-init were written exactly for this: a few hundred lines of code whose only job is to become PID 1, forward signals to the real application, and reap orphaned children. Docker's docker run --init flag makes Docker's own bundled tini the container's PID 1 and starts your actual command as its child; Compose gets the same behavior by adding init: true to a service definition. Bundling tini into your own image and writing ENTRYPOINT ["/usr/bin/tini", "--", "/app"] achieves the same result without depending on Docker's built-in init support.

The second is using a full process manager: tools like s6-overlay or supervisord take on PID 1 responsibilities while also managing several processes in the same container and restarting one if it crashes. This makes sense when you genuinely need more than one long-lived process in a single container, for example an application server alongside a log shipper; putting a full process manager in a single-process container just to reap zombies is unnecessary complexity.

The third option is adding nothing, and making sure that is a deliberate decision. If the application never forks children (most modern web servers and language runtimes do not), there is no zombie accumulation risk to begin with, and adding an init process only adds another layer to signal propagation for no benefit.

Kubernetes is different

There is no direct equivalent of Docker's --init flag in Kubernetes; the pod spec has no field that wraps a container's PID 1 in an init process. Images meant to run on Kubernetes therefore need to bundle the init process themselves: the ENTRYPOINT ["tini", "--", ...] pattern works regardless of the container runtime, because it is the image's own ENTRYPOINT that determines PID 1, not containerd or CRI-O. Setting shareProcessNamespace: true so containers in a pod share a process namespace solves a different problem entirely, letting a sidecar see the main container's processes, and does not solve zombie reaping on its own; PID 1 in that shared namespace is still either the pause container or the first container itself.

Common traps

Writing CMD or ENTRYPOINT in shell form in a Dockerfile (CMD myapp --flag) makes Docker implicitly run it as /bin/sh -c "myapp --flag". In that case PID 1 is not your application but sh, signals go to sh, and sh does not forward most signals to its child; when the container is stopped with docker stop, the graceful shutdown window goes unused and the process is killed with SIGKILL instead. The fix is exec form: CMD ["myapp", "--flag"]. The same mistake can happen when adding tini or dumb-init: invoking the wrapper in shell form hides it behind a sh layer as well.

A second trap is getting the order wrong when combining an init process with a privilege-dropping tool like gosu or su-exec. The correct order has tini as PID 1, tini invoking the privilege-dropping tool, and that tool exec-ing the actual application. In the reverse order, tini ends up sitting above an intermediate shell instead of above the application, and the reaping chain breaks again.

A third trap is trying to track zombie accumulation purely from memory and CPU graphs. Zombies do not show up in either. The correct metric is either counting zombies directly with ps -eo stat,ppid | awk '$1 ~ /^Z/' | wc -l, or watching how close /sys/fs/cgroup/pids.current gets to pids.max under cgroup v2. Without watching one of these two, the problem only surfaces once pids.max is actually hit, usually at the worst possible moment.

When it is not worth bothering

In a container running a single long-lived process that never forks and has no cron-like scheduler inside it, adding an init process is an unnecessary layer. The practical way to assess the risk is to shell into the image, look at what is running with ps -ef, and search the application's codebase for calls like exec, fork, subprocess, or Process.Start. If none of those show up, PID 1's special behavior never comes into play. If they do, the cost of bundling a binary that is a few megabytes in size is far lower than the cost of hitting a sudden "cannot create new process" error in production one day.