This is the reason
docker stop takes ten seconds on a container that is doing nothing at all, and it is not Docker being slow. It is documented kernel behaviour, in pidnamespaces(7), and it surprised me for longer than I would like to admit.Try it:
docker run -d --name t alpine sleep 1000
time docker stop t
Ten seconds, near enough exactly. The container is running `sleep`. There is nothing to flush, nothing to close, no shutdown work of any kind. It sits there for the full grace period and then gets killed.
The reason is that the first process in a PID namespace is treated as an init process, and the kernel protects it from itself. From pidnamespaces(7):
> a process in an ancestor namespace can send signals to the "init" process of a child PID namespace only if the "init" process has established a handler for that signal. SIGKILL or SIGSTOP are treated exceptionally: these signals are forcibly delivered when sent from an ancestor PID namespace.
So SIGTERM to a PID 1 that has not installed a handler is not ignored by the process. It is discarded by the kernel, before the process ever sees it.
sleep has no SIGTERM handler, so the signal goes nowhere, and the ten seconds are the runtime correctly waiting out a grace period for a shutdown that can never begin. Then SIGKILL, which is the one signal that gets through regardless.You can see whether a given process will do this before you try to stop it.
SigCgt in /proc/PID/status is a hex bitmask of the signals that process has installed handlers for:grep SigCgt /proc/$(pgrep -n sleep)/status
Bit 15 is SIGTERM, so the mask is significant against
0x4000. All zeroes means nothing is caught, which means SIGTERM will be discarded, which means you are going to wait the full grace period.Three ways out, and they are not equivalent.
--init puts a tiny init as PID 1 which does have handlers and forwards them, which fixes it for anything. A trap in your entrypoint fixes it for your own script. Changing --stop-signal only helps if the process actually handles the signal you switch to.The part I find worth knowing regardless of containers: this is why an entrypoint shell script that ends with a command instead of
exec behaves differently from one that does. Without exec your shell is PID 1, and shells do not install a default SIGTERM handler for non-interactive use either.https://redd.it/1vcfsm9
@r_linux