Skip to content

[Bug] SIGTERM during INIT_FAILED publication can escape forked child main #1948

Description

@doraemonmj

Platform

All / Unknown

Runtime Variant

All / Unknown

Description

The device-free Ubuntu UT

tests/ut/py/test_worker/test_startup_readiness.py::TestNextLevelStartupFailure::test_second_child_failure_reaps_first

has intermittently failed twice while reclaiming a next-level child after another child reports an injected init failure. The test body observes the expected injected inner init failure, but its finally: w4.close() then raises:

RuntimeError: child next pid <pid> still alive; shm not freed

There is a signal race in _forked_child_main consistent with this failure. During setup() failure it does:

except BaseException as e:
    _write_error(...)
    _mailbox_store_i32(state_addr, _INIT_FAILED)
    os._exit(1)

The parent can observe _INIT_FAILED immediately and start rollback by sending SIGTERM. The child's SIGTERM handler raises _StartupCancelled. If SIGTERM arrives after the _INIT_FAILED store but before os._exit(1), _StartupCancelled is raised from inside the except BaseException suite. The sibling except _StartupCancelled does not catch an exception raised inside another except suite, so it can escape _forked_child_main and unwind into the forked copy of the caller's startup/test frames.

That violates _forked_child_main's documented load-bearing invariant: a forked child must never unwind into the copied parent startup frames and must always terminate through os._exit.

The CI symptom is consistent with the escaping child not becoming reapable before the shared startup-abort deadline expires. It is then placed in the cleanup journal, and the subsequent nonblocking waitpid(..., WNOHANG) in close() can still observe it alive. The controlled repro below proves the exception escape; the exact final still alive outcome remains scheduler-dependent.

Observed CI failures on two consecutive PR heads whose diff does not touch python/simpler/worker.py, this test, or the UT workflow:

A third rerun passed, which matches an intermittent timing race:

Steps to Reproduce

The unmodified test is low probability: it passed 20/20 consecutive iterations on a local Linux aarch64 host.

A controlled device-free repro widens only the interval after publishing _INIT_FAILED. It reliably shows _StartupCancelled escaping _forked_child_main:

PYTHONPATH=tests/ut/py .venv/bin/python - <<'PY'
import time
import simpler.worker as worker
from test_worker.test_startup_readiness import TestNextLevelStartupFailure

original_store = worker._mailbox_store_i32


def delayed_store(addr, value):
    original_store(addr, value)
    if value == worker._INIT_FAILED:
        time.sleep(0.25)


worker._mailbox_store_i32 = delayed_store
TestNextLevelStartupFailure().test_second_child_failure_reaps_first()
PY

The relevant output is:

During handling of the above exception, another exception occurred:

  File "python/simpler/worker.py", line 4371, in _forked_child_main
    _mailbox_store_i32(state_addr, _INIT_FAILED)
  File "<stdin>", line 10, in delayed_store
  File "python/simpler/worker.py", line 4359, in _on_cancel
    raise _StartupCancelled()
simpler.worker._StartupCancelled

The traceback continues into the forked copy of test_second_child_failure_reaps_first, demonstrating the forbidden unwind. On a fast local host the parent usually still reaps the process in time, so the controlled command may exit successfully despite printing the child traceback.

Expected Behavior

  • A setup failure publishes its error and the forked child always terminates through os._exit.
  • SIGTERM during any point of child initialization or failure publication cannot unwind past _forked_child_main.
  • Startup rollback followed by close() is convergent and does not fail because a just-terminated child has not yet become reapable.

Actual Behavior

In CI, rollback leaves the next-level child journaled as a survivor:

[worker pid=...] WARN: startup child reclaim failed (continuing best-effort):
child process(es) [...] did not exit within the close budget

Then w4.close() replays the journal cleanup and fails:

python/simpler/worker.py:3650: RuntimeError
RuntimeError: child next pid <pid> still alive; shm not freed

The child stderr contains the expected injected init failure immediately before the reclaim warning.

Git Commit ID

3b578e30a2a9e2859d19908b2647393ffdecc543 (main base containing the affected Worker code); also observed at the PR heads listed above, where the affected files are unchanged from this base.

CANN Version

N/A — device-free UT

Driver Version

N/A — device-free UT

Host Platform

Linux (x86_64) in CI; controlled exception-escape repro also confirmed on Linux (aarch64).

Additional Context

Related but not duplicates:

A fix should preserve the invariant at the outermost child boundary even if SIGTERM is delivered while handling another exception. Independently, the journal replay path may need a bounded reap retry after termination rather than a single immediate WNOHANG check, but that is secondary to preventing the child entry point from unwinding.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions