Skip to the content.

SF-014 · Zombie process looks alive

Status: stable | Family: E · Process & Environment | Confidence: high | as of 2026-09-19

One line. A hung process reports no error, holds no CPU and never exits — so every liveness signal says it is fine.

Symptom

A long-running job stops making progress. It does not crash, does not log an error, and does not exit. Every check built on “is the process there?” answers yes.

A representative observation:

pid 41208   age 19h   CPU time 0   RSS 0.25 MB   last output: 18h ago
status: running

Four independent signals all point the wrong way:

Signal Reading Naive conclusion
process exists yes it is working
exit code none it has not failed
stderr empty no errors
CPU time 0 idle, presumably waiting

The actual state is stuck. The job will never finish, and nothing will report that.

Compounding variant: the orchestration layer accumulates process groups it never reaps. Spawning the same job five times leaves five live groups — each with its own set of child processes — and the old ones never exit. A fixed 15 processes per group becomes 75 resident processes, most of them zombies. Now even the resource accounting is misleading, because the zombies hold memory and PIDs while contributing nothing.

Why it is silent

A signal that can prove the claim is missing. Liveness is being inferred from existence, and existence is the one property a hung process keeps.

This is Family E: the logic that asked “is it running?” is correct, and the environment gives an answer that is technically true and operationally useless.

Two design factors make it durable:

  1. Absence-based status. “No error” and “no exit” are read as success. There is no positive progress signal to contradict them.
  2. No timeout. A job with no deadline can be in its valid states forever, so “still running” and “stuck” are indistinguishable by construction.

Minimal reproduction

import subprocess, time

p = subprocess.Popen(["python", "worker.py"])       # worker blocks on a dead socket
time.sleep(1)
print("alive:", p.poll() is None)                   # → True
time.sleep(60 * 60)
print("alive:", p.poll() is None)                   # → still True, still no output
# and nothing anywhere has reported a problem

Observed: alive: True indefinitely, no error, no exit — Expected after fix: a timeout kills it and reports stalled: no output for 600s

Self-check

Three questions, all cheap:

  1. When did it last make progress? Not “is it alive” — “when was the last output, the last heartbeat, the last state change?”
  2. Is there a deadline? If any job can run forever without complaint, it has no failure mode.
  3. Are child processes reaped? Compare the expected process count against the observed count. A number that only grows is a leak.

A sharper form of (1): define progress in the code, not in the observer’s head. “Last output timestamp” is a progress signal; “process is running” is not.

Fix

Add positive progress signals, timeouts and reaping — and make the absence of progress a failure.

DEADLINE      = 600        # seconds of silence before declaring a stall
HEARTBEAT     = 30         # expected interval

last_progress = time.monotonic()

with subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
                      start_new_session=True) as p:      # own process group → reaping
    try:
        for line in p.stdout:
            print(line.decode(errors="replace"), end="")
            last_progress = time.monotonic()             # positive progress signal
            if time.monotonic() - last_progress > DEADLINE:
                os.killpg(os.getpgid(p.pid), signal.SIGKILL)
                raise RuntimeError(f"stalled: no output for {DEADLINE}s")
    finally:
        p.wait(timeout=10)                                # never leave it behind

Principles, in order of importance:

Negative control Expected result
Make the worker hang on purpose watchdog kills it and reports stalled, with the silence duration
Run the worker 5 times in sequence resident process count returns to baseline — no accumulation
Run a healthy worker completes normally, no false stall

中文要点