Some time ago I wrote here about a Node service that ignored SIGTERM because its Dockerfile started it through a shell, and the fix: exec form, so the application is PID 1 and receives signals directly. That fix was right. Eighteen months later it produced a different incident in another service, and I owe the sequel.
Our document service renders invoices to PDF by starting a converter through a small wrapper script. If a render takes longer than twenty seconds, our code kills the wrapper. The converter the wrapper had started is then an orphan, and the kernel reparents orphans to PID 1. On an ordinary Linux machine PID 1 is an init system, and one of its duties is to wait on those orphans when they exit so the kernel can release their entries in the process table. In our container PID 1 was Node, which does no such thing for processes it never spawned. Each orphaned converter finished or crashed and then stayed as a zombie: no memory, no CPU, one process id.
The platform sets a limit of one thousand and twenty four processes per pod. At peak we hit the render timeout around thirty times an hour, so a busy pod reached the limit in about a day and a half. After that every attempt to start a converter failed with resource temporarily unavailable, every render returned an error, and the HTTP server carried on answering health checks. One Monday two in five invoice renders failed across three of five pods. A process listing inside one of them showed one thousand and nineteen entries in state Z.
We kept exec form and put a real init in front of it: tini as the entrypoint, which forwards signals to the application and reaps whatever gets orphaned. The wrapper script is gone; we start the converter directly in its own process group and kill the whole group on timeout, so nothing is orphaned in the first place. Each container now exports its process count from the cgroup, with an alert at half the limit. The base chart now adds an init process to every service by default.
Whatever sits at PID 1 inherits the job description of init, whether or not it has read it. We moved our application into the role to fix signals and never noticed the role had a second duty.
– Sergey Shinder
Top comments (1)