Background
In pypto-lib, models/deepseek_v4_pro/prefill_layer.py on a5 (2 cards) has a runtime problem: sustained benchmark re-dispatch of the same prepared graph intermittently wedges the device — ~84 AIC/AIV cores hit the AICore watchdog at once (507015, subErrType 0x4, "timeout or trap error") at a random round (observed rounds 5–62 of a 100-round bench across many runs), and the device pair frequently stays unusable afterwards (the next run fails startup with chip worker(s) [N] did not become ready within 300.0s). This is the residual failure keeping prefill_layer red in the A5 nightly after pypto-lib#960 fixed the MoE window-epoch protocol — first-dispatch validation is stably green; only repeated dispatch dies.
Reproduce with (2 A5 cards):
PYPTO_BENCH=1 PYPTO_BENCH_ROUNDS=200 \
PTO2_RING_DEP_POOL=16384 PTO2_RING_TASK_WINDOW=16384 PTO2_RING_HEAP=1073741824 \
python models/deepseek_v4_pro/prefill_layer.py -p a5 -d <c0>,<c1> --layer-id 2
Repro case: models/deepseek_v4_pro/prefill_layer.py at pypto-lib main (any commit containing pypto-lib#960, e.g. 7a57215 or later; stock file, no local edits).
Built-in negative controls (same command, byte-identical otherwise): --layer-id 0 (SWA layer kind) and --layer-id 3 (HCA layer kind) both complete 200 rounds cleanly under identical conditions. --layer-id 2 selects the CSA layer kind and dies (observed at round 8 in the paired control run).
Reproduction environment:
| Component |
Version |
| pypto-lib |
main with #960 (repro verified at the #960 merge commit content) |
| pypto |
ed2aaa18 (branch: main; also reproduces on the nightly CI's newer pypto heads, 2026-08-13..08-17) |
| simpler |
3165cc89 (detached; == pypto runtime gitlink) |
| ptoas |
v0.57 (matches toolchain/versions.env) |
| pto-isa |
83d01313 (matches runtime/pto_isa.pin) |
| CANN |
9.2.0 |
Diagnosis: simpler — the terminal state is a device-runtime/memory-path wedge: cores parked at ptoas-inserted intra-core pipe wait_flags whose producing MTE2 TLOAD never completes (details below); with a small ring heap the same stall surfaces instead as orch_error_code=2 HEAP_RING_DEADLOCK. The involved kernels are provably victims (they run clean for 200+ rounds when the CSA kernel set is absent).
Description
Two-card repeated dispatch of the composed prefill layer (CSA attention kind + MoE cross-card dispatch/combine) intermittently stalls GM reads on AIC cores; the AICore watchdog then sweeps the device. Evidence chain, all from host-side observation plus build artifacts:
-
Trap dump decoding. Every dump's pc start values (0x…000278 for AIC cores, 0x…002f78 for AIV cores; low bits shift slightly across builds) are the persistent dispatcher entry points, not kernels — the ~75-84 cores parked there (current pc ≈ entry+0x3d8..0x43c) are idle collateral. The truly stuck cores' current PCs fall inside the uploaded ChipCallable image (single image, 92 children in the failing build).
-
Victim identification. Mapping those PCs through the image layout (code_addr = image_base + 0x38 + child_offset(i), offsets read via ChipCallable bindings; image base from the upload_chip_callable_buffer DEBUG log line) lands overwhelmingly in aic_kv_proj_matmul (a plain bf16 [T,7168]x[7168,512] split-K matmul with atomic Add GM accumulation), offsets +0x728..0x9f4, plus occasional aic_prefill_idx_qr_proj / aiv_prefill_idx_qr_rope / aiv_rope_cs.
-
Stuck instruction. Resolving those offsets through the kernel .o's .debug_line shows every parked core sits at a compiler-inserted intra-core pipe sync — wait_flag(PIPE_MTE2→PIPE_MTE1) / wait_flag(PIPE_M→PIPE_MTE1) immediately guarding a TLOAD consumer — i.e. the MTE2 pipe's GM read never completed. The dumps' mte error info / l1 error info fields are non-zero on these cores; vec/cube error info are clean.
-
Trigger isolation (layer-kind bisection, 200 rounds each, same devices/env):
| Runtime kernel set |
Result |
SWA attention + MoE (--layer-id 0) |
PASS (206 dispatches) |
HCA attention + MoE (--layer-id 3) |
PASS |
CSA attention + MoE (--layer-id 2) |
dies round ~8 |
MoE standalone (l3_moe, nightly bench) |
PASS (105 dispatches) |
| CSA attention standalone (nightly bench) |
PASS |
So the CSA-exclusive kernels (idx qr/hadamard/score, c4 softmax/pool, csa cache_write, compressors, inner-state) must run concurrently with cross-card MoE window traffic (remote_store + AtomicAdd notifies) to fire. aic_kv_proj_matmul itself is healthy without the CSA set present.
-
Ring-heap interaction (amplifier, not cause). scope_stats shows one dispatch flows ~1.2 GB through the ring heap, with the reclaim tail pinned at 0 for the whole first request (top reaches 696 MB before the first reclaim) — at the CI's PTO2_RING_HEAP=1GiB the ring must wrap intra-dispatch with minimal slack. Consequences, all verified:
- 256 MiB heap → deterministic first-dispatch death with
orch_error_code=2 HEAP_RING_DEADLOCK ("ran out of both task slots and heap bytes");
- 1 GiB → dies rounds 5-14; at least one death recorded
orch_error_code=2 before the core traps;
- 2 GiB / 4 GiB (effective values confirmed via the
resolve_arena_sizing "Ring buffer sizes" log) → still dies (rounds 62 / 7 / 35 observed) with pure 507015 and no orchestration error — the heap margin only changes how the underlying stall is reported.
-
Post-mortem device state. After the trap ("force reset will follow in finalize"), the pair often rejects the next run (chip worker(s) did not become ready within 300.0s) or produces corrupt validation output, until some later recovery — a secondary robustness gap worth its own attention.
Related: pypto-lib#858 (VEC UB out-of-bounds in the same CSA/sparse-attn family, surfaces as S1 stall), pypto-lib#696, pypto-lib#960 (protocol fix that exposed this residual), simpler#1816 (different signature: 507018 paged-attention stalls).
Steps to Reproduce
- 2-card A5 host, environment pinned as in the Background table (stock pypto-lib
main, pypto pins for simpler/ptoas/pto-isa).
- Run the command from the Background (CSA layer kind, 200 bench rounds, CI-default ring env).
- Typically dies within rounds 5–62; repeat once if a run survives.
--layer-id 0 / 3 under the same command are the negative controls.
- Optional: set
PTO2_RING_HEAP=268435456 for a deterministic first-dispatch HEAP_RING_DEADLOCK, or PYPTO_RUNTIME_LOG=debug on a validate-only run to log the image base for PC mapping.
Expected Behavior
The CSA-layer benchmark completes all rounds like the SWA/HCA layer kinds do (first-dispatch validation already passes in all cases), and a failed run leaves the device pair usable.
Actual Behavior
Random-round device wedge: benchmark rounds interrupted: RuntimeError: ... chip_process dev=N: RuntimeError: run failed with code 507015 (or code -2 with orch_error_code=2 HEAP_RING_DEADLOCK at 1 GiB heap and below); trap dump shows dozens of cores on both dies with subErrType 0x4, the stuck ones at MTE2-guarding wait_flags inside aic_kv_proj_matmul et al.; the device pair often needs manual recovery afterwards.
Platform: a5 (Ascend 950 hardware)
Runtime Variant: tensormap_and_ringbuffer ("device orchestration mode")
Git Commit ID: 3165cc8
CANN Version: 9.2.0
Host Platform: Linux (x86_64)
Background
In pypto-lib,
models/deepseek_v4_pro/prefill_layer.pyona5(2 cards) has a runtime problem: sustained benchmark re-dispatch of the same prepared graph intermittently wedges the device — ~84 AIC/AIV cores hit the AICore watchdog at once (507015,subErrType 0x4, "timeout or trap error") at a random round (observed rounds 5–62 of a 100-round bench across many runs), and the device pair frequently stays unusable afterwards (the next run fails startup withchip worker(s) [N] did not become ready within 300.0s). This is the residual failure keepingprefill_layerred in the A5 nightly after pypto-lib#960 fixed the MoE window-epoch protocol — first-dispatch validation is stably green; only repeated dispatch dies.Reproduce with (2 A5 cards):
Repro case:
models/deepseek_v4_pro/prefill_layer.pyat pypto-libmain(any commit containing pypto-lib#960, e.g.7a57215or later; stock file, no local edits).Built-in negative controls (same command, byte-identical otherwise):
--layer-id 0(SWA layer kind) and--layer-id 3(HCA layer kind) both complete 200 rounds cleanly under identical conditions.--layer-id 2selects the CSA layer kind and dies (observed at round 8 in the paired control run).Reproduction environment:
mainwith #960 (repro verified at the #960 merge commit content)ed2aaa18(branch:main; also reproduces on the nightly CI's newer pypto heads, 2026-08-13..08-17)3165cc89(detached; == pyptoruntimegitlink)v0.57(matchestoolchain/versions.env)83d01313(matchesruntime/pto_isa.pin)Diagnosis: simpler — the terminal state is a device-runtime/memory-path wedge: cores parked at ptoas-inserted intra-core pipe
wait_flags whose producing MTE2TLOADnever completes (details below); with a small ring heap the same stall surfaces instead asorch_error_code=2 HEAP_RING_DEADLOCK. The involved kernels are provably victims (they run clean for 200+ rounds when the CSA kernel set is absent).Description
Two-card repeated dispatch of the composed prefill layer (CSA attention kind + MoE cross-card dispatch/combine) intermittently stalls GM reads on AIC cores; the AICore watchdog then sweeps the device. Evidence chain, all from host-side observation plus build artifacts:
Trap dump decoding. Every dump's
pc startvalues (0x…000278for AIC cores,0x…002f78for AIV cores; low bits shift slightly across builds) are the persistent dispatcher entry points, not kernels — the ~75-84 cores parked there (current pc ≈ entry+0x3d8..0x43c) are idle collateral. The truly stuck cores'currentPCs fall inside the uploaded ChipCallable image (single image, 92 children in the failing build).Victim identification. Mapping those PCs through the image layout (
code_addr = image_base + 0x38 + child_offset(i), offsets read viaChipCallablebindings; image base from theupload_chip_callable_bufferDEBUG log line) lands overwhelmingly inaic_kv_proj_matmul(a plain bf16 [T,7168]x[7168,512] split-K matmul withatomic AddGM accumulation), offsets +0x728..0x9f4, plus occasionalaic_prefill_idx_qr_proj/aiv_prefill_idx_qr_rope/aiv_rope_cs.Stuck instruction. Resolving those offsets through the kernel
.o's.debug_lineshows every parked core sits at a compiler-inserted intra-core pipe sync —wait_flag(PIPE_MTE2→PIPE_MTE1)/wait_flag(PIPE_M→PIPE_MTE1)immediately guarding aTLOADconsumer — i.e. the MTE2 pipe's GM read never completed. The dumps'mte error info/l1 error infofields are non-zero on these cores;vec/cube error infoare clean.Trigger isolation (layer-kind bisection, 200 rounds each, same devices/env):
--layer-id 0)--layer-id 3)--layer-id 2)l3_moe, nightly bench)So the CSA-exclusive kernels (idx qr/hadamard/score, c4 softmax/pool, csa cache_write, compressors, inner-state) must run concurrently with cross-card MoE window traffic (
remote_store+AtomicAddnotifies) to fire.aic_kv_proj_matmulitself is healthy without the CSA set present.Ring-heap interaction (amplifier, not cause). scope_stats shows one dispatch flows ~1.2 GB through the ring heap, with the reclaim tail pinned at 0 for the whole first request (top reaches 696 MB before the first reclaim) — at the CI's
PTO2_RING_HEAP=1GiBthe ring must wrap intra-dispatch with minimal slack. Consequences, all verified:orch_error_code=2 HEAP_RING_DEADLOCK("ran out of both task slots and heap bytes");orch_error_code=2before the core traps;resolve_arena_sizing"Ring buffer sizes" log) → still dies (rounds 62 / 7 / 35 observed) with pure 507015 and no orchestration error — the heap margin only changes how the underlying stall is reported.Post-mortem device state. After the trap ("force reset will follow in finalize"), the pair often rejects the next run (
chip worker(s) did not become ready within 300.0s) or produces corrupt validation output, until some later recovery — a secondary robustness gap worth its own attention.Related: pypto-lib#858 (VEC UB out-of-bounds in the same CSA/sparse-attn family, surfaces as S1 stall), pypto-lib#696, pypto-lib#960 (protocol fix that exposed this residual), simpler#1816 (different signature: 507018 paged-attention stalls).
Steps to Reproduce
main, pypto pins for simpler/ptoas/pto-isa).--layer-id 0/3under the same command are the negative controls.PTO2_RING_HEAP=268435456for a deterministic first-dispatchHEAP_RING_DEADLOCK, orPYPTO_RUNTIME_LOG=debugon a validate-only run to log the image base for PC mapping.Expected Behavior
The CSA-layer benchmark completes all rounds like the SWA/HCA layer kinds do (first-dispatch validation already passes in all cases), and a failed run leaves the device pair usable.
Actual Behavior
Random-round device wedge:
benchmark rounds interrupted: RuntimeError: ... chip_process dev=N: RuntimeError: run failed with code 507015(orcode -2withorch_error_code=2 HEAP_RING_DEADLOCKat 1 GiB heap and below); trap dump shows dozens of cores on both dies withsubErrType 0x4, the stuck ones at MTE2-guardingwait_flags insideaic_kv_proj_matmulet al.; the device pair often needs manual recovery afterwards.Platform: a5 (Ascend 950 hardware)
Runtime Variant: tensormap_and_ringbuffer ("device orchestration mode")
Git Commit ID: 3165cc8
CANN Version: 9.2.0
Host Platform: Linux (x86_64)