Skip to content

[Bug] A5 2-card: intermittent GM-read stall wedges AICore under repeated dispatch (CSA kernels + cross-card MoE traffic); HEAP_RING_DEADLOCK at <=1GiB ring heap #1844

Description

@lwDavid

Background

In pypto-lib, models/deepseek_v4_pro/prefill_layer.py on a5 (2 cards) has a runtime problem: sustained benchmark re-dispatch of the same prepared graph intermittently wedges the device — ~84 AIC/AIV cores hit the AICore watchdog at once (507015, subErrType 0x4, "timeout or trap error") at a random round (observed rounds 5–62 of a 100-round bench across many runs), and the device pair frequently stays unusable afterwards (the next run fails startup with chip worker(s) [N] did not become ready within 300.0s). This is the residual failure keeping prefill_layer red in the A5 nightly after pypto-lib#960 fixed the MoE window-epoch protocol — first-dispatch validation is stably green; only repeated dispatch dies.

Reproduce with (2 A5 cards):

PYPTO_BENCH=1 PYPTO_BENCH_ROUNDS=200 \
PTO2_RING_DEP_POOL=16384 PTO2_RING_TASK_WINDOW=16384 PTO2_RING_HEAP=1073741824 \
python models/deepseek_v4_pro/prefill_layer.py -p a5 -d <c0>,<c1> --layer-id 2

Repro case: models/deepseek_v4_pro/prefill_layer.py at pypto-lib main (any commit containing pypto-lib#960, e.g. 7a57215 or later; stock file, no local edits).

Built-in negative controls (same command, byte-identical otherwise): --layer-id 0 (SWA layer kind) and --layer-id 3 (HCA layer kind) both complete 200 rounds cleanly under identical conditions. --layer-id 2 selects the CSA layer kind and dies (observed at round 8 in the paired control run).

Reproduction environment:

Component Version
pypto-lib main with #960 (repro verified at the #960 merge commit content)
pypto ed2aaa18 (branch: main; also reproduces on the nightly CI's newer pypto heads, 2026-08-13..08-17)
simpler 3165cc89 (detached; == pypto runtime gitlink)
ptoas v0.57 (matches toolchain/versions.env)
pto-isa 83d01313 (matches runtime/pto_isa.pin)
CANN 9.2.0

Diagnosis: simpler — the terminal state is a device-runtime/memory-path wedge: cores parked at ptoas-inserted intra-core pipe wait_flags whose producing MTE2 TLOAD never completes (details below); with a small ring heap the same stall surfaces instead as orch_error_code=2 HEAP_RING_DEADLOCK. The involved kernels are provably victims (they run clean for 200+ rounds when the CSA kernel set is absent).

Description

Two-card repeated dispatch of the composed prefill layer (CSA attention kind + MoE cross-card dispatch/combine) intermittently stalls GM reads on AIC cores; the AICore watchdog then sweeps the device. Evidence chain, all from host-side observation plus build artifacts:

  1. Trap dump decoding. Every dump's pc start values (0x…000278 for AIC cores, 0x…002f78 for AIV cores; low bits shift slightly across builds) are the persistent dispatcher entry points, not kernels — the ~75-84 cores parked there (current pc ≈ entry+0x3d8..0x43c) are idle collateral. The truly stuck cores' current PCs fall inside the uploaded ChipCallable image (single image, 92 children in the failing build).

  2. Victim identification. Mapping those PCs through the image layout (code_addr = image_base + 0x38 + child_offset(i), offsets read via ChipCallable bindings; image base from the upload_chip_callable_buffer DEBUG log line) lands overwhelmingly in aic_kv_proj_matmul (a plain bf16 [T,7168]x[7168,512] split-K matmul with atomic Add GM accumulation), offsets +0x728..0x9f4, plus occasional aic_prefill_idx_qr_proj / aiv_prefill_idx_qr_rope / aiv_rope_cs.

  3. Stuck instruction. Resolving those offsets through the kernel .o's .debug_line shows every parked core sits at a compiler-inserted intra-core pipe sync — wait_flag(PIPE_MTE2→PIPE_MTE1) / wait_flag(PIPE_M→PIPE_MTE1) immediately guarding a TLOAD consumer — i.e. the MTE2 pipe's GM read never completed. The dumps' mte error info / l1 error info fields are non-zero on these cores; vec/cube error info are clean.

  4. Trigger isolation (layer-kind bisection, 200 rounds each, same devices/env):

    Runtime kernel set Result
    SWA attention + MoE (--layer-id 0) PASS (206 dispatches)
    HCA attention + MoE (--layer-id 3) PASS
    CSA attention + MoE (--layer-id 2) dies round ~8
    MoE standalone (l3_moe, nightly bench) PASS (105 dispatches)
    CSA attention standalone (nightly bench) PASS

    So the CSA-exclusive kernels (idx qr/hadamard/score, c4 softmax/pool, csa cache_write, compressors, inner-state) must run concurrently with cross-card MoE window traffic (remote_store + AtomicAdd notifies) to fire. aic_kv_proj_matmul itself is healthy without the CSA set present.

  5. Ring-heap interaction (amplifier, not cause). scope_stats shows one dispatch flows ~1.2 GB through the ring heap, with the reclaim tail pinned at 0 for the whole first request (top reaches 696 MB before the first reclaim) — at the CI's PTO2_RING_HEAP=1GiB the ring must wrap intra-dispatch with minimal slack. Consequences, all verified:

    • 256 MiB heap → deterministic first-dispatch death with orch_error_code=2 HEAP_RING_DEADLOCK ("ran out of both task slots and heap bytes");
    • 1 GiB → dies rounds 5-14; at least one death recorded orch_error_code=2 before the core traps;
    • 2 GiB / 4 GiB (effective values confirmed via the resolve_arena_sizing "Ring buffer sizes" log) → still dies (rounds 62 / 7 / 35 observed) with pure 507015 and no orchestration error — the heap margin only changes how the underlying stall is reported.
  6. Post-mortem device state. After the trap ("force reset will follow in finalize"), the pair often rejects the next run (chip worker(s) did not become ready within 300.0s) or produces corrupt validation output, until some later recovery — a secondary robustness gap worth its own attention.

Related: pypto-lib#858 (VEC UB out-of-bounds in the same CSA/sparse-attn family, surfaces as S1 stall), pypto-lib#696, pypto-lib#960 (protocol fix that exposed this residual), simpler#1816 (different signature: 507018 paged-attention stalls).

Steps to Reproduce

  1. 2-card A5 host, environment pinned as in the Background table (stock pypto-lib main, pypto pins for simpler/ptoas/pto-isa).
  2. Run the command from the Background (CSA layer kind, 200 bench rounds, CI-default ring env).
  3. Typically dies within rounds 5–62; repeat once if a run survives. --layer-id 0 / 3 under the same command are the negative controls.
  4. Optional: set PTO2_RING_HEAP=268435456 for a deterministic first-dispatch HEAP_RING_DEADLOCK, or PYPTO_RUNTIME_LOG=debug on a validate-only run to log the image base for PC mapping.

Expected Behavior

The CSA-layer benchmark completes all rounds like the SWA/HCA layer kinds do (first-dispatch validation already passes in all cases), and a failed run leaves the device pair usable.

Actual Behavior

Random-round device wedge: benchmark rounds interrupted: RuntimeError: ... chip_process dev=N: RuntimeError: run failed with code 507015 (or code -2 with orch_error_code=2 HEAP_RING_DEADLOCK at 1 GiB heap and below); trap dump shows dozens of cores on both dies with subErrType 0x4, the stuck ones at MTE2-guarding wait_flags inside aic_kv_proj_matmul et al.; the device pair often needs manual recovery afterwards.


Platform: a5 (Ascend 950 hardware)
Runtime Variant: tensormap_and_ringbuffer ("device orchestration mode")
Git Commit ID: 3165cc8
CANN Version: 9.2.0
Host Platform: Linux (x86_64)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions