The PID 1 Problem in Containers: Zombie Processes and When You Need an Init System
The first process in a container is not an ordinary process as far as the kernel is concerned. Here is how that difference leads to zombie process buildup, how it eventually hits the cgroup pid limit, and how to choose between tini, dumb-init, and a full process manager.
When a container starts, the first process inside it becomes PID 1 as far as the Linux kernel is concerned, and PID 1 carries two responsibilities an ordinary process does not have. The first is signal handling: when the kernel delivers a signal like SIGTERM or SIGINT to a process that has not installed its own handler, it normally applies a default handler, but that rule does not apply to PID 1. An ordinary process that receives SIGTERM terminates by default; PID 1's default behavior is to do nothing at all. The second is reaping: when a child process exits, the kernel keeps it around as a zombie (defunct) until its parent calls wait() on it and collects the exit code. In a normal process tree, orphaned children get reparented to PID 1, and init systems such as systemd or SysV init exist precisely to reap that inheritance. PID 1 inside a container is usually a web server, an application process, or a shell script, none of which were written for that job.
Both responsibilities look ignorable because most containers run a single process and never fork children. The problem shows up exactly where a child process does get forked but is never properly reaped.
How zombies accumulate
The typical accumulation pattern looks like this: the application shells out to run a periodic task in the background, the shell backgrounds the command with & and exits immediately. When the shell exits, the process it left behind is orphaned and reparented to PID 1. If PID 1 never calls wait() on it, the process sits there as a zombie once it finishes. A one-off background task will not make this visible; a scheduled job that runs every minute or dozens of times an hour accumulates them fast.
Zombies consume neither memory nor CPU, they show up as <defunct> in ps output, and they look harmless. The actual limit is Linux's cgroup pids controller: every zombie still occupies a PID slot, and once you approach the pids.max limit, new process creation (fork()) starts failing. The failure then surfaces somewhere completely unrelated to the zombies themselves: a new command cannot run inside the container, a new request handler cannot be forked, a scheduler silently stops. A liveness probe usually checks whether the main process is still up, and it stays green as long as that process is alive; the only signs of trouble are a rising pids.current in cgroup metrics or a single fork: Resource temporarily unavailable line buried in the logs.
Three approaches
There are three ways to deal with this, and the right one depends on what the container actually runs.
The first is using a lightweight init process. tini and dumb-init were written exactly for this: a few hundred lines of code whose only job is to become PID 1, forward signals to the real application, and reap orphaned children. Docker's docker run --init flag makes Docker's own bundled tini the container's PID 1 and starts your actual command as its child; Compose gets the same behavior by adding init: true to a service definition. Bundling tini into your own image and writing ENTRYPOINT ["/usr/bin/tini", "--", "/app"] achieves the same result without depending on Docker's built-in init support.
The second is using a full process manager: tools like s6-overlay or supervisord take on PID 1 responsibilities while also managing several processes in the same container and restarting one if it crashes. This makes sense when you genuinely need more than one long-lived process in a single container, for example an application server alongside a log shipper; putting a full process manager in a single-process container just to reap zombies is unnecessary complexity.
The third option is adding nothing, and making sure that is a deliberate decision. If the application never forks children (most modern web servers and language runtimes do not), there is no zombie accumulation risk to begin with, and adding an init process only adds another layer to signal propagation for no benefit.
Kubernetes is different
There is no direct equivalent of Docker's --init flag in Kubernetes; the pod spec has no field that wraps a container's PID 1 in an init process. Images meant to run on Kubernetes therefore need to bundle the init process themselves: the ENTRYPOINT ["tini", "--", ...] pattern works regardless of the container runtime, because it is the image's own ENTRYPOINT that determines PID 1, not containerd or CRI-O. Setting shareProcessNamespace: true so containers in a pod share a process namespace solves a different problem entirely, letting a sidecar see the main container's processes, and does not solve zombie reaping on its own; PID 1 in that shared namespace is still either the pause container or the first container itself.
Common traps
Writing CMD or ENTRYPOINT in shell form in a Dockerfile (CMD myapp --flag) makes Docker implicitly run it as /bin/sh -c "myapp --flag". In that case PID 1 is not your application but sh, signals go to sh, and sh does not forward most signals to its child; when the container is stopped with docker stop, the graceful shutdown window goes unused and the process is killed with SIGKILL instead. The fix is exec form: CMD ["myapp", "--flag"]. The same mistake can happen when adding tini or dumb-init: invoking the wrapper in shell form hides it behind a sh layer as well.
A second trap is getting the order wrong when combining an init process with a privilege-dropping tool like gosu or su-exec. The correct order has tini as PID 1, tini invoking the privilege-dropping tool, and that tool exec-ing the actual application. In the reverse order, tini ends up sitting above an intermediate shell instead of above the application, and the reaping chain breaks again.
A third trap is trying to track zombie accumulation purely from memory and CPU graphs. Zombies do not show up in either. The correct metric is either counting zombies directly with ps -eo stat,ppid | awk '$1 ~ /^Z/' | wc -l, or watching how close /sys/fs/cgroup/pids.current gets to pids.max under cgroup v2. Without watching one of these two, the problem only surfaces once pids.max is actually hit, usually at the worst possible moment.
When it is not worth bothering
In a container running a single long-lived process that never forks and has no cron-like scheduler inside it, adding an init process is an unnecessary layer. The practical way to assess the risk is to shell into the image, look at what is running with ps -ef, and search the application's codebase for calls like exec, fork, subprocess, or Process.Start. If none of those show up, PID 1's special behavior never comes into play. If they do, the cost of bundling a binary that is a few megabytes in size is far lower than the cost of hitting a sudden "cannot create new process" error in production one day.
[OK] 25 records · 8 featured · 1 retainer · 16 standard
/02 services
services.tree
> note
· Prices are in USD. Volume discounts of 7% for 6-12 day projects and 14% for 13+ day projects apply. Programs of 100+ person-days are priced individually. Prices are negotiable based on scope, urgency and long-term collaboration.
> ls ~/products/ --open-source
[OK] 2 entries · 1 live · 1 pre-release
/03 products
products.list
liveopen source · mit/prod/01
filex
the self-hosted file manager that embeds anywhere
single go binary · vue/react/web-component embed · 5 storage drivers · realtime collab · rbac · native multi-tenancy · mcp server for ai agents
We build infrastructure, automate everything, and keep systems running.
BRF Tech is a Bursa-based DevSecOps and software consulting firm. We operate across 25 service areas, from Kubernetes cluster management to serverless platform setup, CI/CD pipeline design to AI agent development.
What we do is simple: we set up your systems, automate them, and make sure they won't wake you up at 3 AM. We codify your infrastructure with Terraform, move your deployments to GitOps with ArgoCD, and monitor everything with Prometheus. If something breaks — we intervene before it does.
We're against vendor lock-in. We work with open-source tools, self-hosted solutions, and industry-standard technologies. We build your own serverless platform on Knative, isolate with Kata Containers, manage your secrets with Vault. Every project is delivered with clear scope, clear timeline, clear pricing.
DevSecOps & CI/CD
Kubernetes & Serverless
Infrastructure as Code
AI & Agent (MCP/ACP)
Security & Zero Trust
Full-Stack Development
Self-Hosted Solutions
> grep -i question ~/faq.md
[OK] 3 entries · click to expand
[?]What services are included in BRF Tech's DevOps solutions?▾
We offer CI/CD pipeline setup, Kubernetes cluster management, container migration, infrastructure automation with Terraform, monitoring & alerting, in-house technology installations and DevOps consulting services.
[?]What are your software development services?▾
We offer MVP development, existing software performance optimization, backend API development (Node.js, Go, Python) and full-stack application development. Every project is delivered with minimal technical debt.
[?]Which DevOps tools do you work with?▾
From Kubernetes and Docker to Talos Linux, from Terraform and OpenTofu to Ansible, and across Jenkins, GitLab CI/CD, GitHub Actions, ArgoCD, Flux, Helm and Rancher, we work with industry-standard tooling. For observability and error tracking we use Prometheus, Grafana, Loki, Tempo, OpenTelemetry and GlitchTip; for networking, Cilium and Envoy; for test automation, Playwright, Cypress, k6, Locust, Testcontainers and Trivy. We determine the most suitable toolset together, based on your project’s needs.
Automated build, test and deployment pipelines with GitHub Actions, GitLab CI/CD.
Automated build, test and deploy pipeline setup for existing or new repositories. GitHub Actions, GitLab CI/CD or preferred tool is used. Separate workflows are defined for staging + production environments.
Container architecture migration, cluster setup and orchestration.
Production-ready Kubernetes cluster setup on bare-metal or cloud. Includes Helm charts, Ingress, TLS certificate management, namespace isolation and RBAC.
Migration of existing applications to container architecture.
Migration of existing applications to container architecture. Docker image design, Compose configuration and conversion to Kubernetes manifests. Kata Containers / gVisor can be included for advanced isolation.
In-house tools like GitLab, Nextcloud, VPN, mail server.
Installation of tools like GitLab, Mattermost, Nextcloud, mail server, VPN, internal monitoring on company servers. Deployed on Docker Compose or Kubernetes.
DB + cache + code bottleneck identification and resolution.
Profiling of existing application, identification and resolution of bottlenecks. DB query optimization, caching layer (Redis), service-level improvements.
MCP/ACP agents, RAG pipelines, multi-agent orchestration, model serving.
LLM integration (OpenAI, Claude, Gemini, local models), tool-augmented AI agents with MCP (Model Context Protocol) and ACP (Agent Communication Protocol). RAG pipeline, multi-agent orchestration, model serving (vLLM/Ollama). AI layer for existing business processes.
Event-driven data pipeline with Apache Kafka, RabbitMQ, NATS. ETL/ELT orchestration with Apache Airflow, real-time data streaming, CDC (Change Data Capture). Data lake/warehouse design, schema registry, dead letter queue management.
Packages
Kafka/RabbitMQ setup + basic producer/consumer
$3,500
5–7 person-day
ETL pipeline (Airflow + source → warehouse)
$5,200
8–12 person-day
Full event-driven architecture (CDC + streaming + DLQ)
Hands-on technical training for your teams. Docker & container fundamentals, Kubernetes operations, Terraform IaC, CI/CD best practices, AI/LLM integration workshops. Practical exercises in live lab environments, content customized by skill level.
E2E test automation with Playwright, Cypress, load testing with k6/Gatling. Pre-deploy quality gate integrated into CI/CD pipeline, test coverage reporting, visual regression testing. Reduce test writing time by up to 60% with AI-assisted test generation.
Your Collected Personal Data, Collection Method and Legal Basis
Any information that identifies or makes you identifiable is considered "personal data." When you visit our website or use our services, contact information such as your name, surname, email address, phone number, as well as your IP address and browser cookie data may be collected through automatic or semi-automatic means. This data is processed based on the legal grounds specified in Article 5 of the Personal Data Protection Law No. 6698: "being directly related to the establishment or performance of a contract" and "being mandatory for the legitimate interests of the data controller, provided that it does not harm the fundamental rights and freedoms of the data subject."
Purpose of Processing Your Personal Data
Your collected personal data is processed for the purposes of providing and improving our services, fulfilling customer requests, meeting legal obligations, conducting information security processes and managing communication activities. Your data is processed in a limited and proportionate manner for the stated purposes; when the purpose ceases to exist, data is deleted, destroyed or anonymized.
To Whom and For What Purposes Collected Personal Data May Be Transferred
Your personal data may be transferred to public institutions and organizations as required by legal regulations, to our business partners and technical infrastructure providers for the purpose of delivering services, and to lawyers and consultants in legal disputes. Transfers are carried out in accordance with Articles 8 and 9 of the Law, with necessary technical and administrative measures in place.
Your Rights as a Data Subject
In accordance with Article 11 of Law No. 6698, you have the right to: learn whether your personal data is being processed; request information if it has been processed; learn the purpose of processing and whether it is used in accordance with its purpose; know the third parties to whom it is transferred domestically or abroad; request correction if it has been processed incompletely or incorrectly; request deletion or destruction within the framework of conditions set out in Article 7 of the Law; request that the operations carried out be notified to third parties to whom data has been transferred; object to any adverse result arising from the analysis of data exclusively through automated systems; and claim compensation for damages in case of unlawful processing. You may contact us through the communication channels on our website to exercise these rights.
Terms & Conditions
Terms & Conditions
1. Service Scope and Changes
BRF Tech conducts its software development, DevOps consulting, infrastructure setup and technical support services within the framework of these terms and conditions. The scope of service is determined separately for each project and finalized through mutual agreement. BRF Tech reserves the right to update the scope, pricing and technical details of its services, provided that prior notice is given.
2. Privacy Policy
Personal data collected within the scope of our services is processed in accordance with our Privacy Policy. For detailed information, please review our Privacy Policy page.
3. Service Usage and Responsibilities
Clients agree to use the provided services solely within legal and ethical boundaries. BRF Tech reserves the right to suspend or terminate services in the event of misuse, violation of third-party rights, or use for illegal activities. The client is responsible for the accuracy and currency of the information provided within the project scope.
4. Payment Terms
Payment terms are determined on a project basis and finalized through mutual agreement at the start of the project. Unless otherwise specified, invoices are payable within 15 days of project delivery. A monthly late payment interest of 2% may apply to overdue payments.
5. Cancellation and Refund Policy
Cancellation requests must be submitted in writing at least 7 days before the service start date. For projects already in progress, billing is based on the proportion of completed work. No refunds are issued for software and infrastructure components specifically developed within the project scope.