Repository navigation
Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
GPU compact replay truncates predecessor cycles above 32 bits. Fused dispatch also attempts to process allocated but empty instruction-family arenas, and adaptive sumcheck groups can be fragmented into excessive small launches.
Design Rationale
The replay ABI stores each access cycle as a 32-bit low word plus its high byte in the existing flags word, expanding only the compact record stride. The decoder reconstructs the original cycle without changing proof semantics. Empty arenas are excluded before capacity and registration lookup, while group ordering preserves height partitions and keeps compatible evaluator kinds contiguous. Scheduler state is logged after the owning stream is bound.
Change Highlights
Benchmark / Performance Impact
This is a correctness and scheduling fix; no aggregate multi-GPU performance claim is made. The wider compact record adds three bytes per record family. A local two-shard CUDA proof with WITGEN enabled completed and verified; the full-width decoder cost is tracked separately from this correctness change.
Benchmark command(s):
cargo run --release --features 'jemalloc,gpu,aot,parallel' --bin ceno-reth-benchmark-bin -- --block-number 23587691 --chain-id 1 --cache-dir block_data --mode prove-starkEnvironment: RTX 5070 Ti local sanity, Rust nightly-2025-11-20, and the head commit of this PR.
Testing
The AOT replay tests cover cycles beyond 2^32 and all compact layouts. GPU dispatch tests cover empty allocated families in both compact and field-SoA modes. The local CUDA proof completed with canonical verification and two shards.
Risks and Rollout
Compact replay uses more input bytes and decoder work. Incorrect flag packing would corrupt predecessor ordering, so the layout tests reconstruct every supported access shape. The change is limited to GPU replay and scheduling paths; rollback is by reverting this branch.
Follow-ups
Measure the widened decoder with device counters before making further kernel or record-layout changes.
Copilot Reviewer Directive (keep this section)
When Copilot reviews this PR, apply
.github/copilot-instructions.mdstrictly.