Skip to content

fix(gpu): preserve 40-bit replay cycles and skip empty arenas - #1404

Open
hero78119 wants to merge 34 commits into
masterfrom
feat/multi_gpu
Open

hero78119 wants to merge 34 commits into
masterfrom
feat/multi_gpu

Conversation

@hero78119

@hero78119 hero78119 commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

GPU compact replay truncates predecessor cycles above 32 bits. Fused dispatch also attempts to process allocated but empty instruction-family arenas, and adaptive sumcheck groups can be fragmented into excessive small launches.

Design Rationale

The replay ABI stores each access cycle as a 32-bit low word plus its high byte in the existing flags word, expanding only the compact record stride. The decoder reconstructs the original cycle without changing proof semantics. Empty arenas are excluded before capacity and registration lookup, while group ordering preserves height partitions and keeps compatible evaluator kinds contiguous. Scheduler state is logged after the owning stream is bound.

Change Highlights

  • Preserve and test predecessor cycles through the 40-bit range in AOT and GPU typed replay.
  • Increase compact strides and retain the future-access mask in the widened flags word.
  • Skip empty fused arenas and report observed versus allocated emission capacities.
  • Order common sumcheck groups by adaptive residual width without changing transcript data.
  • Move GPU task-state logging after stream ownership is established.

Benchmark / Performance Impact

This is a correctness and scheduling fix; no aggregate multi-GPU performance claim is made. The wider compact record adds three bytes per record family. A local two-shard CUDA proof with WITGEN enabled completed and verified; the full-width decoder cost is tracked separately from this correctness change.

Benchmark command(s):

cargo run --release --features 'jemalloc,gpu,aot,parallel' --bin ceno-reth-benchmark-bin -- --block-number 23587691 --chain-id 1 --cache-dir block_data --mode prove-stark

Environment: RTX 5070 Ti local sanity, Rust nightly-2025-11-20, and the head commit of this PR.

Testing

cargo fmt --all --check
cargo make clippy

The AOT replay tests cover cycles beyond 2^32 and all compact layouts. GPU dispatch tests cover empty allocated families in both compact and field-SoA modes. The local CUDA proof completed with canonical verification and two shards.

Risks and Rollout

Compact replay uses more input bytes and decoder work. Incorrect flag packing would corrupt predecessor ordering, so the layout tests reconstruct every supported access shape. The change is limited to GPU replay and scheduling paths; rollback is by reverting this branch.

Follow-ups

Measure the widened decoder with device counters before making further kernel or record-layout changes.

Copilot Reviewer Directive (keep this section)

When Copilot reviews this PR, apply .github/copilot-instructions.md strictly.

@hero78119 hero78119 changed the title Feat/multi gpu feat(gpu): add multi-GPU proving and preserve full replay timestamps Sep 22, 2026
@hero78119 hero78119 changed the title feat(gpu): add multi-GPU proving and preserve full replay timestamps fix(gpu): preserve 40-bit replay cycles and skip empty arenas Sep 23, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant