Repository navigation
Conversation
… and consumer Block producers publish each block's lifecycle (open context, per-tx records, sealed header) to the sequence store as it happens; RPC nodes follow the stream, re-execute it deterministically, and hold preconfirmation receipts for the block being built. Design doc: docs/sequencer-bor.md. Producer (eth/sequencer.Publisher, miner hooks): read-before-write follow model — a foreign unsealed window on our tip is followed, not superseded; the only supersede is the seal flush that makes the store match sealed truth. Pre-seal barrier awaits sequencing; the post-seal gate turns store acks into broadcast verdicts (foreign seal refuses, budget expiry broadcasts for liveness after a recheck). Recovery is the reconcile position ladder (anchor, block-anchor probe, floor read) with delta-only re-anchors; producer rotation adopts the dangling window instead of revoking it. Transport hardening: bounded in-flight sends with ack refill, ack-stall watchdog, and a self-heal redial after prolonged channel silence. Consensus-side: a signer outside the active producer set no longer builds at all, so its sequence can never reach the store. Consumer (eth/sequencer.Consumer): follows the gateway stream with warm/cold resume, verifies the commitment chain per entry, re-executes on canonical or parked speculative state (author-nil EVM context, speculative BLOCKHASH, EIP-2935), cross-checks seals (context, gas, receipts root, state root), voids-and-skips on divergence, and fills a capped receipt index evicted on canonical import. The RPC read path that serves these receipts ships separately. Everything is gated behind the [sequencer] config section; the role derives from the sealer flag (mining node publishes, non-mining node consumes). With the section unset there is no behavior change. Validated on kurtosis devnets: a 12-phase chaos campaign (store component restarts, pauses, 200s outages, flapping, partitions, producer and heimdall kills) ended with zero store gaps, zero revoked or reordered preconfirmations, and zero absent heights across 859k entries; preconfirmation receipts measured at p50 ~100ms against ~2.5-2.9s canonical inclusion at 4s blocks, byte-consistent with canonical receipts after import. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ed functions Raise the patch coverage of the sequence-store integration from ~81% to ~91% (local measurement). The consumer gains an execution-level harness: a real imported chain re-executes speculative blocks through the session (canonical and parked parents, every void-and-skip path, the checkSeal divergence matrix, speculative BLOCKHASH resolution), and a stream-level suite follows a live devstore end to end — receipts served pre-seal, evicted on canonical import, and an exact warm resume across a store restart. The worker's barrier-refusal and fill-halt cycles and the backend's role dispatch are covered directly. Decompose the functions the complexity gate flagged — ConfirmSeal, backfillLocked, reader.walk, restoreAbandonedDebtLocked, floorRead, runStream — and split AdoptWindow and SealBlock into their orchestration and resolution halves. The backfill debt machinery moves to debt.go, the build-start read and sealed-height recovery to buildstart.go, and the worker's sequencer integration to miner/sequencer.go, bringing classify.go, adoption.go, and worker.go back under the size gates. No behavior changes. Mutation tiers hold at T1 84.9%, T2 61.7%. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… edge rows Drive the build-start read through every store shape it classifies — unreadable, empty, behind (live window above, seal edge below), sealed with the chain's block, sealed inside and past the recovery grace — and pin the recovered window's exact content. Cover the gate's last-look recheck against sealed generations and live windows, the refusal streak cap and flush unwind, the backfill drain's pruned-jump, byte-budget, and undrainable-debt rows, prime-and-merge debt bounds, the reoffer and regrown-window paths, walk absorb hooks, probe edges, and the worker's store-owned-height discard mid-fill. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Re-arm the finality grace when sync completes: commitWork returns on syncing before it consults the gate, so a resync outliving the grace consumed it silently and opened the gate the moment sync ended — the restart window the gate exists to cover. The anchor is now re-armed in the miner's DoneEvent/FailedEvent handling, before builds unblock. Give the consumer's cold resume ladder distinct rungs (block anchor, then earliest) instead of retrying the identical block anchor before falling back. Fail the build-start read loudly when probeDown violates its contract instead of misclassifying the height as sealed past. Register sequencer flags against defaults when an HCL/JSON config has no sequencer block, instead of dereferencing nil at startup. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Heights at or below a canonical whitelisted milestone are final: immutable and permanently served by the canonical chain, so their store copies serve no consumer. The backfill drain now starts past the milestone — a 1-hour store outage owes seconds of blocks, not the hour — with the skipped prefix crossed by the jump open as a counted forward jump. Without a usable milestone (Heimdall down alongside the store, or a milestone naming a chain we don't hold) a 40-block depth cap below the tip bounds the drain instead. The floor is a fact about our own chain, never a belief about the store: storeSealedTip remains untrusted, so the devnet failure class that removed the earlier freshness bound (skipping heights the store had actually shed) does not reopen — below finality the hole is now intentional. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…arm wiring The drain with a floor inside the pending range (finalized prefix dropped, remainder rebuilt) and the miner's sync-completion re-arm of the finality grace were the two untested branches of the previous commits; codecov's patch gate flagged both. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The two adoption-floor tests swapped a live worker's chainConfig to vary the bor config, racing the worker's newWorkLoop, which re-reads chainConfig.Bor on every veblop tick — CI segfaulted in CalculatePeriod when a tick landed inside the nil-Bor swap window. The floor computation moves to a pure helper (adoptionMinTime) the tests exercise directly; the end-to-end applyAdoption coverage keeps the worker's real config. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…nd pending state Addresses three review findings on the sequence-store publisher: - eth/sequencer: a live store window on a displaced parent (canonical reorg) now mutes the build instead of arming a sticky hold that could refuse to seal forever. The pre-seal barrier lets a muted build seal without a store mirror, and the seal flush supersedes the dead-parent window with sealed truth. - eth/sequencer: a partially adopted window whose transaction proves unexecutable on canonical state now seals on the executable prefix and lets the seal flush supersede the remainder, instead of the barrier refusing the block and re-adopting the same window forever. Adopted window timestamps are also bounded above (mirroring the consensus future-block limit) so a far-future timestamp cannot stall the build. - consensus/bor, miner: a signer outside the active producer set keeps its pending snapshot fresh again, so eth_call/eth_estimateGas against the pending block work on non-producing nodes. The build is instead kept out of the sequence store at the miner via a muted-build flag gated on IsAuthorizedSigner, so it publishes nothing and never contends the store's per-height election. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The test drove a store-owned height with only the fill-loop resync abort (resyncN) guarding the seal, while AwaitSequenced still reported the height uncontested — so under constrained CI scheduling the build could seal before the abort fired (or the started worker's background loop could seal on the shared recorder), failing "must never reach the seal hook". Marking the height contested makes the seal barrier refuse unconditionally, which is what a store-owned height means, so the invariant no longer races the scheduler. Reproduced the old flake at GOMAXPROCS=1 (120/120 fail); green 200/200 after. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…used tasks Two fixes from a devnet load-test freeze: a producer rotated out mid-span kept sealing, its network-rejected block was flushed to the store as sealed truth, and the rotated-in producer then discarded its own valid block against that seal while the chain sat frozen a height below. - eth/sequencer: the broadcast gate now consensus-verifies a foreign store seal before honoring it as ownership of a height, on all three refusal paths (the timeout recheck, a tail-read loss carrying the decoded header, and a header-less STALE loss, which fetches the standing seal once). A seal whose signer the engine rejects can never become canonical, so it is noise, not ownership — the block broadcasts. A consensus-valid seal (an in-flight winner, a twin) keeps its refusal, as does one that cannot be inspected, bounded by the existing refusal cap. The verifier is the engine's VerifySeal, wired through Publisher.SetSealVerifier; a node without one behaves exactly as before. - miner: a gate-refused block now clears its pendingTasks entry. Nothing else ever could — clearPending needs chain progress a refused height never makes — and the leaked entry read as sealing-in-flight to the veblop stall fallback (decideVeblopFallback), disabling the only recovery path while the chain was stalled at that exact height. One refusal became a permanent production stop. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A producer taking over a slot wedged for seventeen seconds on a devnet: its publisher's anchor was stale from idling muted through the previous producer's stint, and the pre-seal barrier read that staleness as "the store holds content this block does not cover", refusing every rebuild. Each refusal stacked another held open in the journal, which is exactly what forbids the anchor rebase (the boundary read only rebases a clean tail), and the transport's classifier was meanwhile holding those entries awaiting the very seal flush the barrier was refusing — a deadlock broken only by heimdall rotating production away as ineffective. The stall also tripped the milestone-lag rotation threshold, and the rotation's floored backfill jumps are what left the dangling windows and store gaps the auditor has been flagging as dropped transactions. The store itself can arbitrate the shape: one walk anchored at the parent proves whether anything stands at or past the new height. Nothing there means a seal can revoke nothing — the barrier passes (both in the coverage fall-through and under a sticky hold, where a takeover's own stale-prefix STALEs look identical to a competitor), the seal flush's STALE rebases the anchor right after, and the store converges. Any entry the walk returns, and anything unreadable, keeps the refusal: a live window at the height still mirrors, a competitor still refuses, and the contested-unreadable protection is untouched. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two liveness guards from an adversarial audit of "the store must never block production", both bounding paths that relied on the store or on heimdall behaving: - The pre-seal barrier had no analog of the gate's refusal cap. Every refusal shape converges when the store is consistent (adopt, mute, backfill, the clean-boundary proof), but a store answering inconsistently across reads — Byzantine, or a load balancer over replicas at different lags — can refuse every rebuild with no benign state ever showing, halting production at that height until heimdall rotates the producer away as ineffective. Funnel every pass/refuse through a per-height streak: past the cap, liveness wins and the seal proceeds, sized past legitimate convergence and below the rotation threshold so the node heals itself before consensus gives up on it. - The gate's consensus check on a foreign store seal runs on the miner's result loop, and the engine's span resolution behind it waits on heimdall availability with no deadline of its own. Heimdall being down already halts production by design in Prepare and Seal, so this adds no new halt — but the result loop freezing over a verdict is needless coupling. Bound the verification; a timeout is unverifiable and keeps whatever verdict stood without it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A loaded producer's pre-seal tail read can land while acks lag the store's committed tail: the anchor sits inside the open window, so the walk returns bare records without the window's open entry. The summary then carries no window and a head past the anchor, and sealMirror's last branch read that as foreign content — refusing the producer's own block, adopting its own window minus the in-flight tail, and costing the height one or two rebuild cycles (2-3s blocks under sustained load, each shaving the last in-flight transaction). The anchor rung's matcher already proves the walked suffix identical to the unacked journal in lockstep; carry that verdict out of the walk (tailInfo.suffixOurs) and let sealMirror compare the block against the journal when it holds — store tail = acked prefix + verified journal suffix means the store's window at this height is the journal's. Any mismatch or extension turns the matcher off and keeps the old adopt path; the mirror rule itself is unchanged. Verified on a 7-validator devnet under 60 tx/s of 32KB-calldata load: 58 refusals across prior runs at this shape, zero after; the new middrainmirror counter caught eight organic occurrences, none costing block time. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Extract sealMirror's journal comparison — at-head or verified mid-drain suffix — into journalMirror, mirroring how sealedThroughParent names its rung. sealMirror had grown past the size gate; the rung reads better with its own contract stated once. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
…g on idle Two root causes behind the store-level reorg John's devnet reported at block 51933 (auditor reorg-class supersession with no chain reorg): a publisher stream that breaks on a healthy low-load stream, then resends and STALEs, then backs off long enough for finality to strand the flush behind a milestone-floor jump — leaving a store hole the auditor reads as a reorg. Ack-stall watchdog: stallTracker.sent() incremented inflight but never touched lastAck, so the deadline measured time since the last ack ever, not how long the current entry had been outstanding. The first send after an idle stretch longer than the deadline read as stalled the instant it went out, and the watchdog reset a healthy stream. sent() now restarts the clock when it opens a fresh wait (inflight 0->1); pipelined sends leave it, so a genuinely hung store still trips. Contention streak: the streak was only ever cleared inside the stale path, so a healthy session ending any other way left it intact — it survived across quiet stretches and each later blip paid the whole 8s ladder. Consolidated into a pure advanceContention that clears the streak whenever the session made progress, whatever ended it. Reproduced on a kurtosis devnet (8s blocks for the idle gap, 1s netem on the redpanda brokers for the slow ack): 43 false ack-stalls in four minutes on the old code, zero after. Both paths keep genuine-stall detection. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
…ync waits (#2385) The call did real work for microseconds - decode, submit to the txpool, submit for preconf - and then held its rpc.SafePool slot while idling in the receipt wait for up to rpc.txsync.defaulttimeout. Since every RPC call is bounded to ep-size (40 per transport, with ep-requesttimeout at 0 disabling the unbounded fallback), that capped concurrent sync calls at the pool size and let a few long-held calls monopolise a transport for a whole wait window. rpc.Slot is a releasable hold on the pool, handed to a task by SafePool.SubmitWithSlot and published on the call context. awaitReceipt releases it on entry and reacquires on exit, so decode/submit and the response tail stay bounded and only the idling is exempt. Reacquisition is best effort: parking here would stop the caller draining the channels it waits on. Waiters are bounded instead by rpc.txsync.maxconcurrent (default 4096), checked before submission so a refusal means the transaction was never accepted. This also caps live chain-event subscriptions. The 100ms tick called the full GetTransactionReceipt, two DB point-misses before reaching the preconf index. Canonical arrival already comes in as a chain event, so the tick now reads only the pending view, with a 1s canonical backstop for an event that carried no matching receipts. The backstop also closes a gap: with preconfs off there was no poll at all, so a missed or receipt-less event meant waiting out the whole timeout. The duplicated blob-sidecar upgrade in SendRawTransaction and SendRawTransactionSync is factored into decodeRawTransaction; behaviour is unchanged. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…itnesses within a signed size band (#2416) * eth/protocols/wit, eth: WIT2 size oracle — accept non-deterministic witnesses within a signed size band Witnesses are not deterministic across nodes: BlockSTM speculative reads make honest nodes collect different-but-valid trie-node sets, so a valid witness routinely hashes differently from the BP-signed WitnessHash. WIT2's fetch-time verifyAgainstSignedHash treated any hash divergence as a fault (reject the bytes, strike the serving peer, fall back to WIT1 after two distinct servers), so a valid witness that hashes differently is rejected and its serving peer struck. Replace exact-hash equality with a size oracle: - Sign the witness size. SignedWitnessAnnouncement gains a WitnessSize field and the announce signing pre-image commits to it (keccak(domain || blockHash || blockNumber || witnessHash || witnessSize)). Producers set it from their own witness length. - Accept within a band for import. verifyAgainstSignedHash accepts for import any witness whose encoded size is <= min(3*signedSize, gas-derived absolute ceiling), regardless of hash; content-correctness is still arbitrated by import-time state-root execution. A within-band witness is re-served/relayed only when byte-identical to the BP's (hash match); a valid non-deterministic variant imports locally but is not re-served, so the signed hash stays a faithful identifier of the bytes on the serving/relay fast-path. Blame for content rests with the producer that signed the announcement, not a relaying or serving peer — preserving WIT2's property of relaying a trusted witness before self-validating. - Bound the size. A witness beyond the band is rejected and the serving peer struck (first occurrence per (peer, block), reusing the distinct-server bookkeeping); distinct servers exceeding the band for the same block fall back to WIT1. The retained gas-derived ceiling keeps the accepted size bounded even when the signed size is implausibly large. No WIT2 is deployed yet, so the signed-announce format changes in place. Scope: WIT2 signed-path only. The WIT1 page-count cross-peer verification is a separate change, handled in a follow-up. * eth/fetcher: fix stale byte-correctness comment on witness verification The comment above verifyAgainstSignedHash still described the pre-oracle byte-correctness invariant (hash must match). This PR's size-oracle rework now accepts hash divergence within the signed-size band, so the comment misdescribed the actual trust boundary an incident responder would rely on. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * eth, eth/fetcher: charge size-oracle import failures, bound the signed size, relax the broadcast gates Review follow-ups for the WIT2 size oracle. Import-time consequence. A witness accepted on the size oracle alone (hash differs from the BP-signed one) that then fails import was logged at debug and forgotten, so a peer relaying the valid BP announce could serve up to the size band in unusable bytes per block at no cost — the WIT1-class weakness WIT2 was closing. Now the witness manager records the serving peer, the divergence and the fetch closure on the import op (verifyAgainstSignedHash returns diverged; enqueueOp keeps the op instead of rebuilding it from parts), and on insertChain failure importBlocks strikes the server, excludes it as a witness source for that block (new SetWitnessSourceExcluder hook → handler witnessSourceExclusions, honoured by resolveWitnessFetchPeer at every tier and released on import), and hands the block back to the witness manager for a re-fetch from another peer (new witnessRetry loop case → retryAfterImportFailure), bounded by maxWitnessImportRetries. A BP-identical witness that fails import is still the BP's fault and is handled as before. Signed size sanity. signedSize*3 could wrap for a hostile size and a zero WitnessSize yielded a zero ceiling; both struck honest servers. The band now saturates and a zero size falls back to the absolute cap, and acceptSignedAnnouncement refuses (and strikes the sender for) a WitnessSize of zero or above the gas-derived absolute cap before deferral, so the announcer is the one charged. Broadcast gates. acceptSignedBroadcast and acceptDeferredBroadcast now apply the same size oracle as the fetch path: a within-band body is accepted for import (sender marked as body-holder) regardless of hash, so a pusher's own post-import witness (flushWitnessWaitersForImported) is no longer rejected downstream; only byte-identical bytes are cached for pre-import serving, and an oversized body is dropped. fetchAndVerifyWitness (relay fetch, serve-only) keeps requiring the BP's bytes by design and now documents that limitation. Also refreshes every remaining "byte-correctness" comment to the size-oracle semantics (fetcher, handler, peerset, wit protocol), and renames the broadcast byte-mismatch meter to broadcast_oversize with new hash_divergence, implausible_size, import_failure and import_retry meters. Tests: end-to-end import-failure re-fetch through the real fetcher loop (TestImportFailureWithDivergedWitnessRefetchesFromAnotherPeer — caught the provenance loss in enqueue), charge/retry unit tests, degenerate-size and saturating-multiply cases, implausible announce size (0 and cap+1 struck, cap accepted), divergent/oversized broadcast on both the signed and deferred paths, source exclusion in resolveWitnessFetchPeer, and the exclusion set lifecycle. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Brme9KQBd7fZBMnMVEhAZU * eth: keep the WIT2 fetcher wiring comment off the lines #2417 touches #2417 changes the NewBlockFetcher call directly above the WIT2 striker wiring; editing the adjacent comment here made the two branches conflict when merged together even though each merges cleanly into the candidate. Leave the striker comment as it was and document the size-oracle extension (import-failure strike + source exclusion) in its own block below. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Brme9KQBd7fZBMnMVEhAZU * eth/fetcher: charge a size-oracle import failure only when the witness could have caused it chargeDivergedWitnessImportFailure struck and excluded the serving peer on any insertChain error. "Diverged" is the normal case on a stateless node (every node persists its own generated witness), so every import failure — including a contract bytecode missing from local disk, which the downloader heals and the witness never carried, or an interrupted insert — would have struck two honest witness sources per block until the node had none left. Gate the charge on a positive allowlist of witness-attributable failures: a missing trie node (incomplete witness), and execution against the witness disagreeing with the header (ErrStatelessStateRootMismatch, ErrGasUsedMismatch, ErrReceiptRootMismatch, ErrBloomMismatch, ErrRequestsHashMismatch, and the fmt-built "invalid merkle root" / stateless self-validation mismatch messages). Everything else — missing code, interrupted or stopped chain, whitelist mismatch, header or DB errors, unknown errors — is not charged. Also: drop the unreachable absBytes == 0 guard in acceptableWitnessSizeCeiling (the page threshold floors at one page); make the deferred-broadcast hash-match test use a size band that would reject the body, so only the hash match admits it; cover the handler size-oracle helpers without a fetcher and the deferred size-band binding's boundary and TTL. Tests: TestChargeDivergedWitnessImportFailureIgnoresNonWitnessErrors, TestIsWitnessAttributableImportError, TestSizeOracleHelpersWithoutFetcher, TestDeferredAnnounceCacheHasWitnessSizeWithin; existing charge/re-fetch tests now use attributable errors. * eth/fetcher: never charge a witness server for a contract bytecode missing from local disk #2401 surfaces a missing contract bytecode on the stateless path as core.ErrStatelessIncompleteState wrapping a *state.MissingCodeError, the same sentinel that wraps a missing trie node. Witnesses carry no code, so the server could not have supplied it and the downloader's self-heal fetches the blob; charging the server would strike and exclude honest witness sources on every block that touches the contract until the node had none left. Check for *state.MissingCodeError first in isWitnessAttributableImportError and return false explicitly, so the wrapped cause — not the shared sentinel — decides attribution: missing node → incomplete witness → charged; missing code → local gap → not charged; the bare sentinel → cause unknown → not charged. Tests: the wrapped MissingCodeError case in TestChargeDivergedWitnessImportFailureIgnoresNonWitnessErrors (no strike, no exclusion, no re-fetch, no budget consumed) and both wrappings of ErrStatelessIncompleteState in TestIsWitnessAttributableImportError. * eth/fetcher: cover the exported size ceiling and the re-fetch guard branches Exercise AcceptableWitnessSizeCeiling from its own package (the handler is its only production caller) and the two retryAfterImportFailure early exits — block already known locally, witness marked unavailable — so the new code is fully covered where it lives. * eth/fetcher, eth: split importBlocks under the complexity gate; pin the sampled mutation survivors Diffguard flagged importBlocks at complexity 21 (threshold 15) after the witness-retry plumbing. Move the goroutine body into runBlockImport, which returns whether the failed import should be retried with a witness from another peer, and the block-tracker log into logTrackedImport; importBlocks now only routes the result to witnessRetry or done. Behaviour is unchanged. Pin the five sampled survivors: assert the diverged flag on both the exact-match and divergent returns of verifyAgainstSignedHash and on the WIT1-only (nil lookup) return; add the exact-boundary case to the saturating multiply (MaxUint64/2 * 2 must multiply, not saturate); and pin that the fetcher's striker and source-excluder callbacks are wired into the handler (StrikeWitnessServer / ExcludeWitnessSource are exported for that test). * eth, eth/fetcher: charge pushed witnesses on import failure; stricter quarantine clear and announce conflict Review follow-ups (pratikspatil024, 2026-09-21). The broadcast path accepted a within-band divergent body for import but could not charge it: handleBroadcast set only op.witness, and both cached-witness attach sites rebuilt the op without provenance, so chargeDivergedWitnessImportFailure returned on its first guard and a NewWitness push was the free way to deliver unusable bytes — and the easier one, since the first witness to arrive is the one attached. acceptSignedBroadcast/acceptDeferredBroadcast now report the divergence, InjectWitness carries it with the pusher, handleBroadcast records witnessPeer/witnessDiverged and the announce's fetch closure, and the before-the-block cache keeps peer+diverged so both attach sites build the op via cachedWitness.injectFor with full provenance. An import failure of pushed bytes now strikes and excludes the pusher and re-fetches from another peer, as on the fetch path; a pusher excluded for a block has its further pushes for that block dropped (broadcast_excluded_source_drop) so it cannot beat every honest re-fetch with the same body. verifyAgainstSignedHash cleared the oversize/quarantine state on every in-band acceptance; a divergent in-band body proves nothing about the signed size, so a server alternating oversized and in-band bodies could keep a block out of quarantine indefinitely. Only BP-identical bytes clear it now (the pending-removal exits and TTL sweep still bound the maps). signedWitnessCache.putIfNewer keyed conflicts on WitnessHash alone; the signed WitnessSize decides the accept band, so a same-hash announce with another size is now a conflict too (first commitment wins) instead of replacing the band once the relay window lapses. Tests: push-path import-failure e2e for both attach sites through the real fetcher loop, provenance bit from both accept functions, excluded pusher dropped, quarantine kept across a divergent acceptance, size conflict rejected before and after the relay window. * eth/fetcher: move the import-failure charge next to its error predicate Diffguard's file-size gate allows a file already over 800 lines to grow by 10% of its base size; block_fetcher.go had reached +138 (10.4%) on the candidate base. chargeDivergedWitnessImportFailure and maxWitnessImportRetries move to witness_import_errors.go, which already holds isWitnessAttributableImportError, the predicate the charge is gated on. No behaviour change. --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
) (cherry picked from commit 4a582dc)
…the whitelisted entry (#2432) * eth/downloader/whitelist: accept re-import of canonical blocks below the whitelisted entry isValidChain rejected any segment that lies entirely below the whitelisted checkpoint/milestone once the local head is past it, without checking whether the segment is already canonical locally. core.insertChain runs ValidateReorg before the known-block trim, so a late re-import of blocks we already have (a downloader cycle whose bodies arrived after the block fetcher imported the same blocks and after the milestone moved past them) surfaced as whitelist.ErrMismatch. The downloader wrapped it as errInvalidChain and struck the sync master, which had only served correct headers. On the block-producer mesh this drove the INC-192 peer drops. Give the finality services a canonical-hash oracle (wired from Service.SetBlockchain for both milestone and checkpoint) and accept a below-whitelist segment only when every header equals the local canonical hash at its number. Forks below the whitelisted entry, partially canonical segments and unknown blocks are still rejected, and a nil oracle keeps the strict behaviour. Tests: unit repro of the false positive (fails on the previous code), fork and partially-canonical rejection, and the public Service.IsValidChain path with checkpoint + milestone whitelisted above the re-imported segment. * eth/downloader, eth/downloader/whitelist: observe stale canonical re-imports and pin the downloader path Make the new accept path observable and cheaper, and cover it at the downloader boundary: - Resolve the canonical hash through ChainReader.GetHeaderByNumber instead of loading the full block; *core.BlockChain already provides it. - Count accepted below-whitelist canonical segments per service (chain/checkpoint/stalecanonical, chain/milestone/stalecanonical) and log them at Info. The event is rare (a handful of times a day on a busy node) and replaces the mismatch warning and peer strike it used to produce, so the log is the signal to watch after rollout. - Add TestImportBlockResultsStaleCanonicalReimport: a downloadTester whose blockchain runs the real whitelist service as its chain validator (the production wiring, which newTester omits). After the fetcher imported the chain to block 64 and a milestone was whitelisted at 60, importBlockResults of the already-canonical blocks 30-31 succeeds and leaves the head alone, while a fork below the milestone is still errInvalidChain wrapping whitelist.ErrMismatch and still classified as whitelist-mismatch. The test fails on the previous whitelist code with "retrieved hash chain is invalid: mismatch error". * .github/workflows: update Kurtosis to fix Foundry installation * eth/downloader/whitelist: split the canonical-segment scan out and assert the stale-canonical meters CI quality metrics flagged isValidChain at cognitive complexity 19 and a surviving mutant on the reportStaleCanonicalReimport call. - Move the per-header canonical check into isCanonicalSegment so isValidChain keeps a flat below-whitelist branch. - Give the mock whitelist service the production service names so the checkpoint and milestone meters are exercised, assert both stale-canonical meters advance by exactly one on an accepted re-import (kills the mutant), and run the meter-touching tests sequentially. - Cover the nil chain-reader guard: without a reader the segment below the whitelisted entry keeps the strict rejection. * eth/downloader/whitelist: exempt canonical re-imports from the locked milestone gate milestone.IsValidChain applies a second gate after the finality check: while this node has a milestone vote in flight (GetVoteOnHash locks the service until the candidate finalizes, the normal state on a validator), IsReorgAllowed refuses every segment ending at or below the locked candidate with no canonicity check. A late re-import of already-canonical blocks was therefore still reported as a mismatch on validators, both below the whitelisted milestone, which nullified the previous commit there, and in the gap between the whitelisted milestone and the locked candidate, which the first fix never covered. Route the gate through contradictsLockedMilestone: a segment proven canonical by the oracle is a no-op re-import and cannot contradict the vote, so it is let through, counted in chain/milestone/lockedcanonical and logged at Info. Forks below the candidate and segments carrying the locked block with a different hash are still rejected. Tests: Service-level coverage of the locked state (both windows accepted, fork and contradicting segment rejected, meters advance once each) and a locked phase in the downloader-boundary test. Both fail without the exemption. * eth/downloader/whitelist: count the locked-candidate meter only above the whitelisted entry A canonical segment lying below the whitelisted milestone is already reported by isValidChain; when the service is also locked the lock gate reported the same event a second time, so the two meters double counted the common validator case. Report from the lock gate only for segments ending at or above the whitelisted entry, keeping the two meters distinct populations. --------- Co-authored-by: Vikram Bhattacharjee <vbhattacharjee@polygon.technology> (cherry picked from commit 16e609e)
A speculative block that opens on its predecessor's parked post-state reuses that StateDB, and the StateDB counter that stamps log.Index never resets, so served preconfirmation receipts and pending logs carried the previous blocks' log counts. It showed intermittently -- whenever the open outran the parent's canonical import -- and accumulated across a run of speculative blocks. Canonical import renumbers on its own path, so only the served view was wrong, and the invalidation ledger could not see it because the receipt comparison does not cover Index. Stamp Index from a per-block counter on the serving path, the way canonical import already does, and seed a prefix continuation from the prefix's log count. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…d-port #2429 and #2432 from v2.10.1-candidate (#2435) * eth/downloader: restrict sync penalty exemptions to trusted peers (#2429) (cherry picked from commit 4a582dc) * eth/downloader/whitelist: accept re-import of canonical blocks below the whitelisted entry (#2432) * eth/downloader/whitelist: accept re-import of canonical blocks below the whitelisted entry isValidChain rejected any segment that lies entirely below the whitelisted checkpoint/milestone once the local head is past it, without checking whether the segment is already canonical locally. core.insertChain runs ValidateReorg before the known-block trim, so a late re-import of blocks we already have (a downloader cycle whose bodies arrived after the block fetcher imported the same blocks and after the milestone moved past them) surfaced as whitelist.ErrMismatch. The downloader wrapped it as errInvalidChain and struck the sync master, which had only served correct headers. On the block-producer mesh this drove the INC-192 peer drops. Give the finality services a canonical-hash oracle (wired from Service.SetBlockchain for both milestone and checkpoint) and accept a below-whitelist segment only when every header equals the local canonical hash at its number. Forks below the whitelisted entry, partially canonical segments and unknown blocks are still rejected, and a nil oracle keeps the strict behaviour. Tests: unit repro of the false positive (fails on the previous code), fork and partially-canonical rejection, and the public Service.IsValidChain path with checkpoint + milestone whitelisted above the re-imported segment. * eth/downloader, eth/downloader/whitelist: observe stale canonical re-imports and pin the downloader path Make the new accept path observable and cheaper, and cover it at the downloader boundary: - Resolve the canonical hash through ChainReader.GetHeaderByNumber instead of loading the full block; *core.BlockChain already provides it. - Count accepted below-whitelist canonical segments per service (chain/checkpoint/stalecanonical, chain/milestone/stalecanonical) and log them at Info. The event is rare (a handful of times a day on a busy node) and replaces the mismatch warning and peer strike it used to produce, so the log is the signal to watch after rollout. - Add TestImportBlockResultsStaleCanonicalReimport: a downloadTester whose blockchain runs the real whitelist service as its chain validator (the production wiring, which newTester omits). After the fetcher imported the chain to block 64 and a milestone was whitelisted at 60, importBlockResults of the already-canonical blocks 30-31 succeeds and leaves the head alone, while a fork below the milestone is still errInvalidChain wrapping whitelist.ErrMismatch and still classified as whitelist-mismatch. The test fails on the previous whitelist code with "retrieved hash chain is invalid: mismatch error". * .github/workflows: update Kurtosis to fix Foundry installation * eth/downloader/whitelist: split the canonical-segment scan out and assert the stale-canonical meters CI quality metrics flagged isValidChain at cognitive complexity 19 and a surviving mutant on the reportStaleCanonicalReimport call. - Move the per-header canonical check into isCanonicalSegment so isValidChain keeps a flat below-whitelist branch. - Give the mock whitelist service the production service names so the checkpoint and milestone meters are exercised, assert both stale-canonical meters advance by exactly one on an accepted re-import (kills the mutant), and run the meter-touching tests sequentially. - Cover the nil chain-reader guard: without a reader the segment below the whitelisted entry keeps the strict rejection. * eth/downloader/whitelist: exempt canonical re-imports from the locked milestone gate milestone.IsValidChain applies a second gate after the finality check: while this node has a milestone vote in flight (GetVoteOnHash locks the service until the candidate finalizes, the normal state on a validator), IsReorgAllowed refuses every segment ending at or below the locked candidate with no canonicity check. A late re-import of already-canonical blocks was therefore still reported as a mismatch on validators, both below the whitelisted milestone, which nullified the previous commit there, and in the gap between the whitelisted milestone and the locked candidate, which the first fix never covered. Route the gate through contradictsLockedMilestone: a segment proven canonical by the oracle is a no-op re-import and cannot contradict the vote, so it is let through, counted in chain/milestone/lockedcanonical and logged at Info. Forks below the candidate and segments carrying the locked block with a different hash are still rejected. Tests: Service-level coverage of the locked state (both windows accepted, fork and contradicting segment rejected, meters advance once each) and a locked phase in the downloader-boundary test. Both fail without the exemption. * eth/downloader/whitelist: count the locked-candidate meter only above the whitelisted entry A canonical segment lying below the whitelisted milestone is already reported by isValidChain; when the service is also locked the lock gate reported the same event a second time, so the two meters double counted the common validator case. Report from the lock gate only for segments ending at or above the whitelisted entry, keeping the two meters distinct populations. --------- Co-authored-by: Vikram Bhattacharjee <vbhattacharjee@polygon.technology> (cherry picked from commit 16e609e) --------- Co-authored-by: Vikram Bhattacharjee <53374818+vbhattaccmu@users.noreply.github.com> Co-authored-by: Vikram Bhattacharjee <vbhattacharjee@polygon.technology>
…lready passed (#2434) * eth/sequencer: judge the served commitment on a matching store seal too The audit only reconciled the served preconf commitment when the store could not confirm a height (NOT_FOUND, unsealed, unusable seal); on a plain match it cleared the commitment unjudged. A matching seal only says what the store holds now. If the generation this node followed was displaced — a backup producer republished the height and its block became canonical — and the node crashed before the live path judged it, the store's current seal matches canonical while what this node served does not. That broken promise was dropped, unrecorded. Judge the commitment on every verdict except a mismatch, which already recorded the height. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * eth/sequencer: judge served commitments the audit watermark already passed The audit never revisits a height below its watermark, so a served preconf commitment still sitting there is never judged again. Two ways one gets there: judgeServed leaves a commitment in place when the canonical body is not local yet ("so a later pass can judge it"), but the pass then persists the mark past it and no later pass ever looks; and after a rewind moves the head below the mark, the node re-serves those heights and persists fresh commitments that the live path never clears (markCanonicalHeadAudited returns early at or below the mark). A contradicted promise in either case went unrecorded. Sweep the served commitments below the range at the start of every pass and judge them with the existing reconcileSkippedServed. One iterator seek per pass; only the entries present are visited, and there are normally none. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * eth/sequencer: bound the audit sweep by finality and cancellation The sweep below the watermark judged [0, watermark] on every pass. The watermark is seeded at the head on a first run, so that included heights above finality, which the walk deliberately never judges. Bound the sweep by the same ceiling. Check the context before and during the sweep so a cancelled pass returns instead of judging every commitment first, and decide the store dial from a side-effect-free read so runAuditPass no longer seeds the watermark on the way to that decision. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * eth/sequencer: seal the canonical block in the matching-seal audit test The fixture sealed one header at 7 and made canonical another, so the seal did not actually match. Derive the seal and canonical hash from the reordered block. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * eth/sequencer: single exit for the served commitment after a verdict Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * eth/sequencer: resolve the audit bounds once per pass Fold rangeToAudit into run: read the head, finality and watermark once, seed on a first run, sweep the served commitments up to min(watermark, finality), then walk. The consumer always builds the lazy store client instead of predicting whether the walk will need it, which removes the read-only twin of rangeToAudit and the nil-fetch special case. Log when the audit writes an invalidation record; a sweep-only pass has no walk report, so this was otherwise silent. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * eth/sequencer: retain served evidence when audit reads or verdict writes fail * eth/sequencer: serialize served audits and bound local sweeps * eth/sequencer: retain unreadable evidence on store mismatches --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
PendingLogRange was the only pending accessor left out of the head fallback added in #2393. On a non-mining sequencer RPC node, in the gap between importing block N and receiving the producer's open for N+1 the consumer has no active pending view, so PendingLogRange fell back to miner.Pending() (nil on such a node) and returned nothing. The filter layer turned that into "-32000 pending logs are not supported", failing eth_getLogs(fromBlock: N, toBlock: "pending") ~16% of the time on a 1s chain, while the equivalent {pending: true} form (which already had the fallback) never failed. Serve the same single empty head+1 view the block and state accessors already return, so the documented fromBlock..pending indexing loop stops erroring in the gap. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…s canonical import (#2439) * core, eth/sequencer: keep preconf receipts readable across the head write CompletePreconf evicted the matched block's preconf receipts and moved the reconciled marker to it before writeHeadBlockWithPreconf made the canonical receipts readable. A receipt read in that window missed the index and failed the read anchor (reconciled ahead of the head), so eth_getTransactionReceipt flipped receipt -> null -> receipt on every import. Matched receipts now stay indexed until a new PreconfHeadObserver hook runs after the head write, and a landing marker lets reads anchor on the matched block while its head write is in flight. Mismatched receipts are still evicted immediately. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * eth/filters: stop a stalled subscriber from freezing the event loop The event system delivers every subscription from one goroutine with blocking sends, and each subscription goroutine wrote to its socket inline. A websocket client that stopped reading blocked its write for the 10s default write timeout, stalling pending logs, logs, newHeads, pending transactions, receipts and new subscriptions for every client. Subscriptions now write from their own goroutine behind a bounded backlog, and end on a failed write or overflow instead of skipping events. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * eth/sequencer: withdraw entries built on a rejected block at completion When canonical import committed a block that did not match its preconfirmation, CompletePreconf dropped only that height. Entries above it, built on the rejected parent, stayed in the pending store and receipt index and were served until the post-import reconcile, because the read anchor still matched the unchanged head until the head write. Completion now runs the future-entry reconcile against the committed block, withdrawing and invalidating every entry that does not extend it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * rpc, eth/filters: close the connection of an abandoned subscription A subscription that overflowed its backlog or failed a write ended silently, leaving a slow but live client with an open connection and no events. JSON-RPC cannot fail a single subscription, so the server now closes the connection (new Notifier.CloseConn), which the client sees as an error on every subscription and can resubscribe from. The backlog limit rises from 1024 to 4096, the event system's own burst allowance, so a burst on a healthy client no longer reaches it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * eth/sequencer, eth/filters, rpc: tidy the preconf and subscription fixes - CompletePreconf takes one matched/unmatched branch, resolves the store once and writes its invalidations in one call; the index clear-from scan is shared with reconcileCanonicalHeadLocked instead of relying on withdrawOffCanonical returning heights in order. - withdrawOffCanonical returns at once on an empty store, so imports on a node with no preconf activity skip the scan and sort. - deliver double-buffers its backlog (no allocation per event for a subscriber that keeps up), shares one teardown path, and sizes the limit from txChanSize. - jsonWriter declares close, so Notifier.CloseConn cannot silently no-op. - The mismatch test reuses the pending RPC coverage fixture. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * eth/filters, eth/sequencer, internal/ethapi: address review on the preconf RPC fixes - Unsubscribe drains a state-sync subscription's channel like the others, so ending one can no longer deadlock against the event loop. - CompletePreconf publishes landing before reconciled, so a reader that sees the new marker also sees the in-flight head write. - eth_getTransactionReceipt retries the canonical lookup once after a preconf miss when the head moved since its first lookup: the preconf copy is evicted only after the canonical write, so the block may have landed in between. An unchanged head skips the retry, so unknown hashes cost no extra lookup. - The delivery tests no longer depend on how much the kernel buffers before a stalled write blocks, or on how the writer batches the backlog. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * wip: review round 2 * eth/sequencer, eth/filters, internal/ethapi: verify preconf review fixes * eth/sequencer: preserve matched landing during delayed head reconciliation * eth/sequencer: verify current and descendant invalidation persistence * eth/filters, rpc: move subscriber isolation into a separate PR --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
SetHead recovers each parent's state before moving the head, all under chainmu. The chain syncer could see the head in that gap, treat the chain as stateless and switch to snap sync, which disables the trie database and makes the rewind exit with log.Crit. Only report a missing head state when it is confirmed under a free chainmu, and take the TD from the head that was checked.
…2458) A build-start read served by a replica that has not yet applied our parent's seal shows the parent as a dangling window, which classifies as buildBehind. That read decodes no seal, so no backfill is primed, and the holdBuild it installs waits on a seal ack that already happened: nothing lifts it, and the whole window batches until its own seal. When the read's head is our own lineage at or behind our store-acked seal of the parent, the store owes nothing, so open as at a boundary instead. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* core, miner, consensus/bor, eth, triedb/pathdb: pipelined state root computation for block import (#2180)
* Initial delayed SRC PoC
* miner, core, consensus/bor: pipelined state root computation (PoC)
* miner: run speculative fillTransactions concurrently with SRC and removed the post tx execution buffer time
* miner: async DB write, concurrent fill, and interrupt timer improvements
* llint fix
* addressed comments and fix test, lint
* core/stateless: (fix unit test) fix NewWitness zeroing breaking witness manager hash matching
* core, consensus/bor, eth, triedb: pipelined state root computation for block import
Overlap SRC(N) with execution of block N+1 on importing/RPC nodes.
After executing block N, defer IntermediateRoot + CommitWithUpdate to a
background SRC goroutine and immediately proceed to block N+1 using a
FlatDiff overlay for state reads. Cross-call persistence allows the SRC
to run across insertChain boundaries.
Key changes:
- Pipeline path in insertChainWithWitnesses with ValidateStateCheap
- FlatDiff overlay in StateAt, StateAtWithReaders, PostExecutionStateAt
- Path DB reader chained fallback for concurrent layer flattening
- Trie-only reader for SRC witness generation (no flat reader bypass)
- WIT handler waits for pipelined witness before returning empty
- WitnessReadyEvent for announcing witnesses to stateless peers
- PropagateReadsTo in checkAndCommitSpan for witness completeness
- Feature gated: --pipeline.enable-import-src
* tests/bor: add pipelined import SRC self-destruct integration test
Adds TestPipelinedImportSRC_SelfDestruct to verify that the FlatDiff
Destructs check in getStateObject correctly handles self-destructed
contracts during pipelined import.
* core/state, triedb/pathdb: fix prefetcher race during pipelined SRC
Two fixes for prefetcher errors during pipelined state root computation:
1. Storage root mismatch: FlatDiff accounts had storage roots from block
N's post-state, but the prefetcher's NodeReader was at the committed
parent root (grandparent). Add prefetchRoot field to stateObject that
stores the grandparent's storage root, read from the flat state reader
when loading from FlatDiff. Use it consistently across all prefetcher
interactions.
2. Layer stale during trie node resolution: SRC's cap() flattens diff
layers concurrently with prefetcher trie walks. Add nodeFallback to
reader.Node(), mirroring the existing accountFallback/storageFallback
pattern — retries via the current base disk layer on errSnapshotStale.
* miner, consensus/bor, core, eth: harden pipelined SRC abort handling
A series of fixes for pipelined SRC under EIP-2935/BLOCKHASH aborts and
abort-heavy devnet load:
1. Skip pipeline pre-Rio.
Pre-Rio speculative Prepare walks unsigned speculative headers and can hit
ecrecover failures on zero-seal Extra data. Disable pipelined SRC before
Rio so the miner stays on the safe sequential path there.
2. Move slot waiting fully to Seal and keep abort rebuilds in-slot.
The miner now always builds block bodies early and uses the slot for tx
selection, while Bor holds propagation until the target time in Seal().
Abort-recovery headers carry a miner-local AbortRecovery flag so late
speculative rebuilds stay in-slot instead of getting pushed to the next
slot by minBlockBuildTime.
3. Isolate block-build timeout state per build environment.
Sequential builds and speculative fills previously shared a worker-global
timeout flag, so one build's timer could interrupt another build's tx
selection. Move timeout state onto each environment and make timer cancel
stop the timer without poisoning the build as timed out.
4. Improve speculative fill behavior and fix DAG metadata on refill.
Speculative blocks now take a late refill pass when they are still under
about 75% full by gas and there is at least 300ms left before the slot,
not only when fully empty. Keep tx dependency DAG state on the block
environment across refill passes so multi-pass speculative fills do not
restart dependency indices from zero and drop metadata with
non-sequential transaction index errors.
5. Harden abort recovery and mined-block propagation.
After speculative aborts, requeue normal work through the standard worker
path instead of re-entering commitWork recursively. On the networking
side, mined inline blocks now still announce correctly when witness data
is already cached but the async block write is not yet visible in the DB.
6. Add regression coverage and clean up logs.
Add tests for Bor timing behavior, speculative refill decisions,
per-build interrupt isolation, DAG metadata persistence across refill
passes, cached-witness announcement, and BLOCKHASH(N) abort-flag
behavior. Also remove duplicate EIP-2935 abort logs and fix negative
seal-delay logging so slightly-late blocks no longer print huge wrapped
unsigned delays.
* core, miner, core/state: added metrics for pipelined SRC
Wires a complete metrics suite for A/B comparing pipelined vs non-pipelined
import and block production on mainnet.
New pipelined metrics (import):
- chain/imports/pipelined/{hit,miss,root_mismatch,enabled}
- chain/imports/witness_ready_end_to_end — apples-to-apples end-to-end timer,
fires in both modes (primary A/B KPI)
New pipelined metrics (build):
- worker/pipelineSpeculativeCommitted, pipelineSRCWait, pipelineSealDuration
- worker/pipelineAnnounceEarlinessMs (signed ms — PIP-66 earliness signal)
- worker/pipelineSpeculativeAborts/{blockhash,src_failed,fallback}
- worker/build_to_announce — producer-side end-to-end, both modes
- worker/pipeline/enabled
Parity wiring for legacy metrics so dashboards work in both modes:
- chain/inserts, account/storage read + hash + update + commit timers,
snapshot/triedb commits, stateCommitTimer, blockBatchWriteTimer,
witnessEncode/DbWrite — emitted from the pipelined branch (main statedb or
SRC goroutine's tmpDB as appropriate)
- worker/writeBlockAndSetHead — emitted from inlineSealAndBroadcast's async
write goroutine
- pipelineAnnounceEarlinessMs and pipelineSpeculativeCommittedCounter also
emitted from resultLoop for the sealBlockViaTaskCh path
Throughput and overlay observability:
- chain/{gas_used_per_block,txs_per_block,mgasps} + chain/witness/size_bytes
- worker/chain/{gas_used_per_block,txs_per_block}
- state/flatdiff/{account_hits,storage_hits} — FlatDiff overlay effectiveness
Metrics that have no clean pipelined semantic (chain/validation, chain/write,
worker/commit, worker/finalizeAndAssemble, worker/intermediateRoot) are left
unemitted in pipelined mode with inline comments documenting the reason and
pointing to the closest pipeline equivalent.
* core, core/state, eth, tests/bor, miner: refactor pipelined-src functions for diffguard compliance
Decompose large pipelined-src-authored functions into focused helpers so
every function owned by this branch sits under diffguard's 50-line /
complexity-10 limits. Pure structural refactor — no behavior change.
miner/pipeline.go:
- commitSpeculativeWork (599) → orchestrator (35) + specSession struct
with ~18 methods (setupInitial, waitForSRCAndSealBlockN, runOneIteration,
prepareNextIteration, sealCurrentAndAdvance, shiftToNext, etc.)
- inlineSealAndBroadcast (100) → 35 + sealViaPrivateChannel,
rebindReceiptsToSealedBlock, announceInlineSealedBlock
- commitPipelined (59) → 37 + buildSpeculativeReq, spawnSRCForFinalBlock
- sealBlockViaTaskCh (52) → 48 (reuses spawnSRCForFinalBlock)
miner/worker.go:
- fillTransactions (59) → 47 + commitTxMaps
- makeEnv (51) → 38 + resolveStateFor
- updateTxDependencyMetadata (68) → 32 + buildTxDependencyArray
Pre-existing develop functions where pipelined-src had grown the body
are reduced back close to or below their develop size by extracting the
added branches:
- commitWork (67 → 36) via clearPendingWorkOnExit + maybeStartPrefetch
- resultLoop (191 → 124; develop was 123) via emitExecutionMetrics,
emitCommitMetrics, writeTaskBlock, announceTaskBlock
- mainLoop (135 → 120; develop was 116) via handleSpeculativeWork
- buildAndCommitBlock (93 → 83; develop was 80) via submitForSealing
core/state/statedb.go:
- CommitSnapshot (95, complexity 40) → 30 + captureMutation,
captureObjectStorage, captureReadOnlyAccount, captureNonExistentRead
- ApplyFlatDiffForCommit (49, complexity 20) → 16 + applyFlatMutation
- ApplyFlatDiff (36, complexity 11) → 13 + applyFlatAccountOverlay
- TouchAllAddresses (25, complexity 11) → 12 + touchAddressAndStorage,
mutatedStorageKeys
core/blockchain.go:
- SpawnSRCGoroutine (127, complexity 35) → 13 + runSRCCompute,
openSRCStateDB, preloadFlatDiffReads, emitSRCStateDBMetrics,
encodeAndCachePendingWitness
- writeBlockAndSetHeadPipelined (108, complexity 29) → 16 +
writePipelinedBlockBatch, writeBorStateSyncLogs, resolveWriteStatus,
emitPipelinedWriteEvents
- handleImportTrieGC (52, complexity 16) → 21 + capTrieIfDirty,
maybeFlushChosen, dereferenceUpTo
- waitForPipelinedWitness (complexity 11) → 9 + waitForPendingSRCWitness,
pollWitnessCache
core/evm.go:
- SpeculativeGetHashFn (complexity 12) → 17 + newPendingBlockNResolver
core/blockchain.go insertChainWithWitnesses pipelined branch (had grown
+222 lines on top of develop's 452) → +42 via buildPipelineImportOpts,
persistPipelinedImport, collectPrevImportSRCIfAny, emitStateSyncFeed,
runImportAutoCollection, verifyImportSRCRoot, publishImportWitness,
emitPipelinedImportParityMetrics.
core/blockchain.go ProcessBlock pipelined branches (+22 lines) → +4
via pipelineReaderRoot, applyFlatDiffOverlayToAll, validateStateForPipeline.
eth/peer.go:
- doWitnessRequest (pipelined-src pushed from 38 → 65) → 32 +
awaitWitnessResponse extracting the goroutine body
eth/handler_wit.go:
- handleGetWitness (pipelined-src pushed from 70 → 91) → 66 +
resolveWitnessSizes consolidating per-hash size resolution (rawdb +
header-existence DoS guard + SRC cache fallback)
tests/bor/helper.go:
- InitMinerWithPipelinedSRC (65) → 32 + newPipelineTestNode (17),
importValidatorKey (11)
- InitImporterWithPipelinedSRC (64) → 31 (same helpers)
Mutation coverage. Ran diffguard in diff-scoped mode (-base develop
-include-paths <module>) across every module pipelined-src touches and
filled the gaps it surfaced:
- core/state: adds core/state/statedb_pipeline_mutations_test.go with 41
targeted tests that kill 24 of 28 mutation survivors in pipelined-src
FlatDiff code (statedb.go lines 2031-2330, 2492-2499). The 4 remaining
are equivalent mutants — Finalise removes destructed addrs before the
guarded branches can fire (2114, 2163), a zero-length loop produces
the same output with or without the guard (2141), and an empty-slice
map entry is observationally equivalent to a missing entry (2150).
Covers CommitSnapshot and its capture helpers, ApplyFlatDiff +
applyFlatAccountOverlay, ApplyFlatDiffForCommit + applyFlatMutation,
NewWithFlatBase, TouchAllAddresses + touchAddressAndStorage +
mutatedStorageKeys, WasStorageSlotRead, and PropagateReadsTo — 14 of
15 functions at 100% line coverage (captureReadOnlyAccount at 90.9%).
- core/stateless: extends witness_test.go with 3 tests targeting
ValidateWitnessPreState's expectedBlock guard (ParentHash and Number
checks that defend against a malicious peer substituting a witness
for a different block / fork). Previous tests all passed nil for
expectedBlock, leaving the entire anti-forgery branch uncovered.
- eth/filters: adds TestResolveBlockNumForRangeCheck and
TestCheckBlockRangeLimit (16 subcases) to api_test.go covering the
RPC range-limit DoS guard at the unit level (sentinel resolution,
span-at-limit boundary, sum-vs-span distinction). Extends
TestInvalidGetRangeLogsRequest in filter_system_test.go to also
exercise GetBorBlockLogs with an inverted range — previously only
GetLogs was covered.
Per-module mutation scores after this coverage: miner 96%, consensus/bor
100% (41/41), core 100% (447/447), core/state 86% (24/28 equivalent),
core/stateless 100% (8/8), core/txpool 100% (12/12), tests/bor 100%
(43/43), triedb/pathdb 100% (20/20). eth at 53% — remaining survivors
are in auto-generated gen_config.go boilerplate (36), P2P dispatcher
cancel-channel plumbing awaitWitnessResponse goroutine cleanup
(3); documented as accepted gaps requiring complex mock infrastructure
for diminishing security return.
Remaining diffguard violations in miner and core are pre-existing
develop functions (commitTransactions, insertChainWithWitnesses,
newWorkLoop, NewBlockChain, ProcessBlock, writeBlockWithState, etc.)
that were over threshold on develop before pipelined-src. Their
pipelined-src deltas are now small (+1 to +42 lines) and out of scope
for this PR.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* core, core/txpool, miner: PR review fixes + pipelined import hardening
Renames (reviewer nits):
- PostExecutionStateAt → PostExecState (BlockChain + txpool/legacypool/
blobpool interfaces + test mocks).
- ResetSpeculativeState → SetSpeculativeState, SpeculativeResetter →
SpeculativeSetter (the method overwrites, it doesn't revert).
Dedup between writeBlockAndSetHead and writeBlockAndSetHeadPipelined.
Both paths now share resolvePostWriteStatus(block, stateless) for
fork-choice + reorg (stateless flag preserves the errInvalidNewChain
escape for fast-forward sync), emitPostWriteEvents for the feed sends,
and writeBorStateSyncLogs for the pre-Madhugiri bor receipt. ~80 lines
of duplicated fork-choice + event logic removed; writeBlockAndSetHead
drops from ~70 to ~11 lines. Batch bodies intentionally not merged —
witness source (statedb.Witness() vs pre-encoded bytes) and trie-commit
timing genuinely differ.
Miner coinbase unification: extracted resolveCoinbase(blockNumber,
fallback) used by makeHeader (fallback=genParams.coinbase) and the
speculative header builders (fallback=etherbase()). Divergence between
the speculative and real header would cause a state root mismatch, so
single-sourcing this is security-meaningful. Rest of buildInitialSpecHeader
kept separate from makeHeader (placeholder parent, deterministic bor
period timestamp, static GasCeil, no engine.Prepare); comment documents
why unifying further would hurt readability.
Pipelined import correctness fixes (core/blockchain.go):
1. persistPipelinedImport now runs the Heimdall milestone/checkpoint
ValidateReorg guard before writeBlockAndSetHeadPipelined — mirrors
the non-pipelined path. Without it, a milestone whitelisted during
block execution could be bypassed.
2. verifyImportSRCRoot wraps its writeHeadBlock revert in chainmu
TryLock/Unlock. The call ran in the auto-collection goroutine
without the mutex, racing any concurrent InsertChain on head state.
Skips + warns if chainmu is closed (shutdown).
3. flushPendingImportSRC error in insertChain's ProcessBlock error
path no longer discarded with `_`; logged like the other two call
sites.
Linter: dropped two `tc := tc` loop-var copies in eth/filters/api_test.go
(copyloopvar, redundant since Go 1.22).
* core: added metrics for preloadFlatDiffReads in pipelined SRC
* core: added metrics - cheap exec and auto-collection phases for pipelined import
* core: stop execution prefetcher in pipelined import path
The pipelined import path replaces IntermediateRoot with CommitSnapshot,
which never terminates the StateDB's trie prefetcher. The deferred
StopPrefetcher in insertChainWithWitnesses only fires on the final
activeState of the batch, so for a batch of N blocks, N-1 prefetchers
leak — each pinning its state/storage tries, decoded trie nodes, and
per-trie hash-set caches.
* core, miner: honour producewitnesses in pipelined import SRC path
The pipelined import path was unconditionally running witness construction
and FlatDiff preload, even when the operator passed makeWitness=false to
InsertChain (i.e. --witness.producewitnesses=false). That made the flag a
no-op for non-witness RPC operators running with --pipeline.enable-import-src,
who paid the full witness + preload cost for nothing.
* core: use multi-reader StateDB on pipelined SRC witness-off path
When the pipelined import SRC goroutine runs without producing a witness,
open the temporary StateDB with state.New instead of state.NewTrieOnly so
pre-state reads in ApplyFlatDiffForCommit, SelfDestruct, and
getOrNewStateObject can hit a flat reader (pathdb StateReader in path
mode, snapshot in hash mode) instead of forcing every read through the
MPT. CommitWithUpdate's MPT hashing walk is unaffected; only the
pre-commit reads avoid cold trie traversals.
The witness-producing path is unchanged (NewTrieOnly remains required so
the witness captures proof-path nodes during the trie walk).
* core, core/stateless, miner: share execution witness with pipelined SRC
The pipelined import path created two witness objects per block: one
attached to the StateDB during ProcessBlock (where EVM execution called
AddCode and AddBlockHash) and a second one inside the SRC goroutine
(where ApplyFlatDiffForCommit + CommitWithUpdate trie walks called
AddState). Only the SRC witness was encoded and published, so
execution-time AddBlockHash entries were dropped. BorWitness encoding
intentionally excludes Codes (verifiers source bytecode from local
storage via CodeRoutingDB), so the wire-observable hole was specifically
the Headers slice: stateless verifiers consuming a published witness
for a BLOCKHASH-using block could not reconstruct the referenced
ancestor hashes.
* core, core/state, eth, internal/cli, miner: add warm-snapshot handoff for pipelined SRC
Pipelined SRC opens a NewTrieOnly StateDB and walks the MPT for
witness-producing pre-state reads. On cold start, this can repeat trie
node loads that the main execution prefetcher already performed for the
same block. This adds an opt-in handoff that lets SRC reuse those
already-loaded trie nodes.
After block execution, StopAndSnapshotPrefetcher drains the subfetcher
goroutines (writers-exited via terminate(false)) and copies their loaded
nodes into an immutable WarmSnapshot keyed by (owner, path, hash). The
snapshot is handed to SRC, which wraps the StateDB's NodeDatabase so trie
reads consult the snapshot before falling through to pathdb. The hash
component prevents a stale entry from satisfying a different state's
read; entries with distinct hashes at the same (owner, path) are both
retained.
The wrapper also forwards Preimage / InsertPreimage / PreimageEnabled
when the inner database supports them, so trie.NewStateTrie still detects
a preimage store on wrapped tries.
Operator-gated by --pipeline.warm-snapshot (default off). Snapshot
capture is further gated on the witness-on SRC path: the witness-off path
uses the multi-reader with flat readers and does not consult the snapshot,
so capturing one would only add copy + Keccak cost.
* core/state: use warm snapshot for SRC commit trie opens
Pipelined SRC's warm snapshot was only installed on the initial trie-only
reader. CommitWithUpdate later reopens account and storage tries through
StateDB.db.OpenTrie/OpenStorageTrie, which meant the commit and witness
collection paths still bypassed the warm snapshot and fell back to pathdb.
Install a snapshot-aware database wrapper on the short-lived SRC StateDB so
Reader, OpenTrie and OpenStorageTrie all share the same hash-verified warm
NodeDatabase. This preserves NewTrieOnly semantics and witness collection:
tries still walk the MPT and resolveAndTrack still records proof nodes; only
the underlying node fetch is short-circuited on snapshot hits.
Also remove the now-dead TrieOnlyReaderWithSnapshot helper and reuse the same
snapshot NodeDatabase for the main trie reader and commit-time trie opens.
* core, core/state: add import phase observability
Add detailed import timing metrics for both normal and pipelined block
import so chain segment elapsed time can be broken down by phase.
For pipelined SRC import, emit timers for cheap execution, cheap
validation, synchronous post-exec work, prefetch stop/snapshot,
CommitSnapshot, previous-SRC collection, state-sync feed, reorg check,
FlatDiff install, metadata write, and SRC spawn. Add slow-block logs that
include the same phase breakdown plus StateDB read/update/hash/witness
sub-timers.
For normal import, emit comparable total/process/validation/reorg/write
timers and matching slow-block logs.
Also add segment-level throughput metrics for the existing "Imported new
chain segment" reporting path, and warm-snapshot hit/miss plus payload
metrics so SRC snapshot effectiveness can be measured directly.
* core, core/state: split pipelined prefetch-stop observability
Add phase-level observability around the pipelined SRC warm-snapshot
handoff. StopAndSnapshotPrefetcher now reports how much time is spent
draining the prefetcher, collecting warm trie-node maps, building the
immutable WarmSnapshot, and reporting prefetch metrics.
Thread those stats into the pipelined import metrics and slow-block log,
including subfetcher counts and account/storage warm-node byte/node
breakdowns. This keeps the shutdown lifecycle unchanged, but makes it
clear whether prefetchStop tails come from drain wait, snapshot
construction, or report overhead.
* core, core/state: build warm snapshot inside SRC goroutine
Move the expensive WarmSnapshot construction out of the synchronous
pipelined import post-exec path.
The import thread now stops the execution-side prefetcher, collects the
quiesced warm-node maps into a WarmSnapshotInput, and passes that handoff
to SRC. The SRC goroutine builds the final immutable WarmSnapshot before
opening the snapshot-aware trie reader.
This keeps the same safety boundary: subfetchers have exited before their
trie witness maps are read, and the final WarmSnapshot still owns copied
node blobs. The difference is that copy/hash/index work no longer inflates
prefetchStop/postExec on the import thread.
The warm_snapshot/build metric now measures SRC-side build time. Slow
pipelined import logs no longer report warmBuild as a synchronous phase.
* core/state: make pipelined SRC prefetch stop snapshot-fast
Pipelined SRC only needs warm nodes that are already loaded by the execution
prefetcher. It does not need to synchronously drain every queued speculative
prefetch task before spawning SRC; missing warm nodes are safe performance
misses because SRC falls through to pathdb.
Add a snapshot-fast prefetcher stop mode for StopAndCollectWarmSnapshot.
The new mode rejects new work, drops queued/unstarted tasks, avoids starting
new trie/pathdb reads after stop is requested, and still waits for in-flight
subfetcher goroutines to exit before reading trie witness maps.
Keep the existing full-drain behavior for normal StopPrefetcher callers.
Add tests covering queued-task discard, full-drain preservation, and the
synchronous production terminateForSnapshot path.
* core/state: bound snapshot-fast prefetch drain latency
Pipelined SRC only needs warm nodes that the execution prefetcher has already
loaded. Missing warm nodes are safe cache misses because SRC falls through to
pathdb, so the warm-snapshot stop path should not wait for one large
speculative prefetch batch to finish.
Split subfetcher account and storage prefetches into bounded chunks. In
snapshot-fast mode, the subfetcher now checks for stop between chunks and exits
before starting more trie/pathdb reads. Normal full-drain termination still
processes every chunk, preserving existing StopPrefetcher semantics.
Add tests covering snapshot-fast account/storage chunk cancellation and the
full-drain invariant for chunked account prefetches.
* core, core/state, miner: detach import prefetcher for SRC
Move pipelined import prefetcher shutdown out of the import thread.
After CommitSnapshot, the execution StateDB now detaches its trie
prefetcher and hands it to the SRC goroutine. SRC then synchronously
drains/reports the detached prefetcher before computing the root.
WarmSnapshot remains optional: when enabled, SRC builds the snapshot from
the fully drained detached prefetcher; when disabled, SRC just waits for
the prefetcher to finish and discards the warm nodes. This lets us A/B the
prefetch lifecycle independently from the snapshot reader.
Add explicit import-thread vs SRC-thread metrics:
- chain/imports/pipelined/prefetch_detach
- chain/imports/pipelined/src/prefetch_wait
- chain/imports/pipelined/src/prefetch_report
- chain/imports/pipelined/src/prefetch_subfetchers
Remove the superseded snapshot-fast stop path and its tests. The
remaining detached-prefetcher lifecycle uses full-drain semantics only,
with focused tests for detach, nil/empty handles, and one-shot stop
consumption.
* consensus, miner: restore Bor early block announcement timing
Restore the Bor timing split removed in 5d45f02: normal producers wait in
Prepare until the parent slot boundary, while Giugliano+ primary producers
return immediately from Seal so blocks can be announced before their own
timestamp.
Thread the required waitOnPrepare flag through the consensus Engine interface.
Normal mining passes true, while speculative/prefetch pipeline paths pass false
and perform their own parent-boundary wait before sealing. This preserves early
announcement semantics without blocking speculative header preparation.
The import-side future-block checks already match develop: post-Giugliano
headers are accepted once local time has reached the parent timestamp, subject
to the existing upper bound.
Also add/update tests for Prepare wait behavior, Seal early return behavior,
and explicit Prepare call sites.
* miner, cli: disable production pipelined SRC
Remove the production-side pipelined SRC config and CLI surface while
keeping import-side pipelined SRC configurable. Miner pipeline eligibility
now stays hard-disabled, the worker pipeline gauge reports disabled, and
block-production interrupt timers use the normal non-pipelined boundary.
Keep the previous production eligibility logic as a commented re-enable
reference so the constraints are easy to recover if this path is revisited.
Update CLI defaults and docs for import SRC: default import pipelining and
verbose logs off, default warm-snapshot on, with an explicit note that
warm-snapshot has no effect when import SRC is disabled.
* miner: preserve disabled pipeline behavior
* miner: snapshot pipeline eligibility for cleanup
* core/state: skip invalid FlatDiff storage prefetch roots
When pipelined SRC exposes the previous block through a FlatDiff overlay, accounts loaded from the overlay carry block N's post-state storage root. The miner/import prefetch readers, however, are opened at the committed parent root. For accounts that already existed in the committed parent, the overlay path resolves the committed storage root and uses that as the prefetch root. For accounts created only by the FlatDiff, there is no committed-parent storage trie to prefetch.
The merge with BlockSTM v2 left those new FlatDiff accounts falling back to their post-state storage root. That lets storage prefetch scheduling hand a block-N root to a committed-parent reader, which pathdb reports as Unexpected trie node at the storage-trie root.
Set the prefetch root to the empty storage root when the account is absent from the committed parent, and make all storage prefetch/get-prefetched/used paths skip empty roots. This keeps the execution overlay correct while preventing best-effort prefetch from opening an impossible trie.
Update the FlatDiff-new-account regression test to pin the empty-root behavior.
* core/state: fix V2 FlatDiff storage prefetch roots
The combined pipelined-SRC + BlockSTM branch was still logging bursts of pathdb hash mismatches during startup:
Unexpected trie node location=diff ... path=[]
The previous FlatDiff prefetch-root fix covered the normal stateObject paths, but V2's FinaliseFastWithPrefetch still snapshotted dirty storage slots using obj.data.Root. For accounts loaded from a FlatDiff overlay, obj.data.Root is the previous block's post-state storage root. The execution prefetcher, however, is opened at the committed parent root, so scheduling a storage trie with the FlatDiff post-state root can ask pathdb for a root that does not belong to that reader's state.
Carry the resolved prefetch root through snapshotDirtyStorageSlots and have FinaliseFastWithPrefetch schedule storage prefetches with that root. This preserves normal accounts by falling back to data.Root, uses the committed-parent root for FlatDiff overlay accounts, and skips new FlatDiff accounts whose storage trie did not exist at the committed parent root.
Add regressions for both cases: existing FlatDiff accounts must prefetch with the committed storage root, and new FlatDiff accounts must not schedule a storage prefetch against the FlatDiff post-state root.
* core: expand pipelined import post-exec metrics
Add finer-grained timers around persistPipelinedImport so mainnet experiments can explain post-execution overhead instead of relying on the aggregate post_exec timer alone.
The new metrics split out witness capture, collect bookkeeping, error-path prefetch cleanup, SRC block construction, pending-state publication, and a residual bucket for any post_exec time that is not covered by the known phases. Slow pipelined import logs now include the accounted and residual totals too, making it easier to tell whether a spike is a known phase or missing instrumentation.
* core: log pipelined import mode on V2 failures
* eth: mask pipelined head state sync checks
* core/state: prioritize FlatDiff in SafeBase reads
Pipelined import can execute block N+1 against committed root N-1 plus the pending FlatDiff for N. V2 SafeBase reads were checking the shared storage cache before the FlatDiff, so a slot warmed from the committed root could mask the newer parent-block overlay value. That can make parallel execution charge gas/refunds from stale SSTORE state and fail ValidateStateCheap with a gas-used mismatch. A retry can appear clean once the pending SRC layer has been committed and the same block executes against parent root N directly.
Route SafeBase and lazy StateDB committed-storage reads through a shared FlatDiff storage overlay helper. Explicit FlatDiff storage entries now win over shared trie caches, and FlatDiff destruct entries cover all slots by returning zero for old pre-destruction storage that was not rewritten by a resurrection.
Add regressions for FlatDiff entries winning over SafeBase shared cache, destruct masks hiding shared-cache values, and lazy destruct+resurrect overlays not exposing old storage.
* core/state: apply FlatDiff to SafeBase account reads
Pipelined V2 flatdiff import executes block N+1 over committed root N-1 plus the collected FlatDiff for block N. The previous FlatDiff SafeBase fix made storage slots consult that overlay before shared trie caches, but account scalar reads could still fall through to pooled StateDB copies.
Those pooled copies may already hold stateObjects loaded from root N-1. In that case GetBalance, GetNonce, GetCode, GetCodeHash, Exist, or GetStorageRoot can observe stale pre-FlatDiff account data while storage reads observe the FlatDiff view. Mainnet devnode logs showed this as flatdiff-only gas mismatches that immediately disappeared on direct retry, with the divergent transactions being EIP-7702 type-4 calls that are sensitive to account/code/existence state.
Add a FlatDiff account overlay helper and have SafeBase scalar account getters consult it before acquiring pooled StateDB readers. Account updates in the FlatDiff now provide balance, nonce, code hash, storage root, existence, and changed code bytes; destruct entries mask stale stateObjects as non-existent account data. Uncovered accounts still use the existing pooled read path.
Add regression coverage for stale stateObjects loaded before the FlatDiff reference is attached, covering both updated-account and destructed-account cases.
Tests: go test ./core/state
* core/state: fail V2 on SafeBase read errors
StateDB getters such as GetState, GetBalance, GetNonce, GetCode, GetCodeHash, Exist, and GetStorageRoot record database read failures internally and return zero-ish values to their caller. SafeBase previously cached those returned values unconditionally. If a pooled StateDB copy hit a transient missing-node, stale PBSS layer, or similar read failure during V2 execution, the zero-ish result could become a stable SafeBase cache entry for the rest of the block.
That is unsafe for pipelined SRC and PBSS because SRC can advance or flatten pathdb layers while the next block is executing or prefetching. A stale-layer read should make the current speculative V2 result unusable, not silently convert missing account/storage/code data into consensus state.
Track the first read error observed by SafeBase, avoid caching pooled StateDB read results unless the read completed cleanly, and replace any pooled StateDB copy that has recorded an error instead of returning it to the worker pool. ExecuteV2BlockSTM now carries the SafeBase/base read error in V2ExecutionResult, and V2StateProcessor aborts the block with v2: base read so the importer can retry through the normal fallback path.
The FlatDiff overlay paths still cache their explicit overlay values directly because they do not perform a database read. The guarded cache writes only apply to fallback reads through StateDB.
Add SafeBase regression coverage for storage, account scalar, and code read failures to prove failed zero-ish results do not poison caches. Update the V2 gas determinism fixture selection to skip incomplete embedded witnesses now that base read failures are surfaced instead of ignored.
Tests: go test ./core/state ./core
* core/state: keep SafeBase FlatDiff-agnostic
SafeBase should be a concurrent read-through cache over the block's logical base, not the owner of FlatDiff semantics. A StateDB constructed with a FlatDiff reference is the ground truth for block N+1 execution on top of committed root N-1, so every SafeBase miss must go through StateDB getters.
Remove the SafeBase storage cache and account-scalar FlatDiff bypasses. Those paths made SafeBase reason about pending system-contract writes, shared trie storage caches, and FlatDiff coverage directly, which duplicated StateDB rules and could let raw root-N-1 cached values win before StateDB applied the overlay.
Move the remaining overlay ordering into StateDB and stateObject. FlatDiff account coverage now masks stale stateObjects loaded before the reference unless the current execution has already dirtied that account. FlatDiff-backed objects are marked so repeated reads reuse the overlay-backed object, and FlatDiff storage coverage now beats stale originStorage populated from committedParentRoot before the reference was attached.
Keep SafeBase's read-error handling intact: failed StateDB reads are not cached and poisoned pooled copies are replaced, so a missing-node or stale-layer read cannot become a cached zero-ish base value.
Tests: go test -count=1 ./core/state -run 'Test(StateDB_FlatDiff|SafeBase_)'; go test -count=1 ./core; git diff --check
* core: instrument pipelined SRC overlap
Add direct observability for the pipelined import question: how much of SRC for block N actually overlaps execution of block N+1, and whether that overlap correlates with slower execution or SRC work.
The implementation timestamps each pending SRC goroutine, carries the SRC handle into the next block's PipelineImportOpts, and records the intersection between the winning execution branch's Process window and the previous SRC window. This gives a per-block overlap signal instead of relying only on temporal dashboard correlation.
New overlap metrics:
- chain/imports/pipelined/overlap/execution: duration for which the previous block's SRC was running during the current block's execution.
- chain/imports/pipelined/overlap/execution_percent: overlap/execution ratio for the current block, emitted as 0..100.
- chain/imports/pipelined/overlap/blocks: count of blocks whose execution had positive overlap with previous SRC.
- chain/imports/pipelined/overlap/no_overlap: count of pipeline-hit blocks where previous SRC had no execution overlap.
New execution-path metrics:
- chain/imports/pipelined/execution: winning execution branch duration, measured around the processor Process call only.
- chain/imports/pipelined/execution/with_overlap: execution duration for blocks with positive previous-SRC overlap.
- chain/imports/pipelined/execution/no_overlap: execution duration for blocks without previous-SRC overlap.
- chain/imports/pipelined/execution/overlap_0_percent: execution duration for blocks with 0% overlap.
- chain/imports/pipelined/execution/overlap_1_25_percent: execution duration for blocks with >0% and <25% overlap.
- chain/imports/pipelined/execution/overlap_25_50_percent: execution duration for blocks with >=25% and <50% overlap.
- chain/imports/pipelined/execution/overlap_50_75_percent: execution duration for blocks with >=50% and <75% overlap.
- chain/imports/pipelined/execution/overlap_75_100_percent: execution duration for blocks with >=75% overlap.
New SRC-path metrics:
- chain/imports/pipelined/src/open_statedb: time to open the temporary StateDB used by SRC.
- chain/imports/pipelined/src/apply_flatdiff: time to replay the FlatDiff into the SRC StateDB.
- chain/imports/pipelined/src/commit: time spent in SRC CommitWithUpdate/root computation, also preserving the existing chain/state/commit parity sample.
- chain/imports/pipelined/src/with_next_exec_overlap: total SRC wall-clock for SRCs that overlapped the next block's execution.
- chain/imports/pipelined/src/no_next_exec_overlap: total SRC wall-clock for SRCs that did not overlap the next block's execution.
The SRC with/no-next-exec split is intentionally one block delayed: SRC_N is classified when block N+1 records its execution overlap and then N is collected. The trailing pending SRC at the end of a short run may remain unclassified, which is acceptable for long catch-up windows.
Update TestPipelinedImportMetrics to assert the new metric streams fire, the execution split matches overlap/no-overlap counters, the overlap percent buckets classify every pipeline hit, and the SRC next-exec split mirrors the overlap classification.
Validated with:
- go test ./core -run 'TestPipelinedImportMetrics|TestPipelinedImportSRC_MakeWitnessFalse|TestPipelinedImportSRC_MultipleBlocks'
- go test ./core
* Optimize witness-off pipelined SRC import
Reduce overhead in the witness-off pipelined SRC path used during catch-up import.
This change splits the pipelined block write path so persistPipelinedImport can complete the durable block batch and post-write status resolution, start the SRC goroutine, and then overlap SRC work with the synchronous head/event publication tail. The public SpawnSRCGoroutine behavior is preserved for existing call sites; import-owned SRC state now uses a private spawn helper and auto-collection waits on that exact pending SRC instance instead of the global pending SRC slot.
For witness-off SRC, add ApplyFlatDiffForCommitFast. The fast replay bypasses journal revert-entry construction because FlatDiff replay in the SRC goroutine is irrevocable, while still marking addresses dirty so Finalise and CommitWithUpdate perform the normal trie root calculation. Witness-producing SRC keeps the existing ApplyFlatDiffForCommit path so proof-node collection remains conservative.
The fast path preserves correctness for pure destructs, destruct-and-resurrect accounts, new accounts, code deployment, and storage writes. It intentionally keeps the parent-root account storage root for overlay-produced FlatDiff account metadata and applies only the current FlatDiff storage slots as dirty writes; this avoids regressing chained pipeline blocks where block N changed storage and block N+1 only changes account scalars.
Add parity tests comparing trie-only journaled replay, multi-reader journaled replay, fast witness-off replay, and direct execution roots. Add a targeted regression test for preserving the parent storage root after execution from a FlatDiff overlay. Classify the new StateDB method in the V2 method parity allowlist as a pipelined SRC block-level operation.
* core, core/state: carry committed SRC nodes to the next pipelined SRC
The witness-off pipelined SRC opened a cold StateDB every block:
CommitWithUpdate re-resolved from pathdb every trie path the previous
block's SRC had just committed, roughly doubling trie-node traffic and
competing with BlockSTM execution for CPU. Index each SRC's commit node
set (shared blobs, hashes already computed by the committer) as a
hash-verified WarmSnapshot and hand it to the next block's SRC, whose
commit-time trie openings consult it before falling through to pathdb.
Value reads keep the multi-reader.
Consecutive blocks rewrite largely overlapping trie paths, so the carry
covers exactly the nodes the prefetcher-based warm snapshot could never
serve: the exec-side prefetcher runs at the grandparent root, making
every path the previous block rewrote a guaranteed hash miss.
Gated on pipeline.warm-snapshot, which previously had no effect when
witnesses were off; witness-on behavior is unchanged. Lookups are
hash-keyed with pathdb fallthrough, so root determinism does not depend
on carry contents.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, core/state: retain recent SRC commit nodes in a bounded warm ring
The single-block carry answered only ~20% of the witness-off SRC's
trie-node fetches: a block's write set overlaps its immediate
predecessor far less than it overlaps the preceding few dozen blocks
(hot contracts and the trie's upper levels repeat constantly, deeper
paths recur on longer horizons).
Replace the per-block handoff with WarmNodeRing, a BlockChain-owned
cache of the recent SRCs' committed node sets. Generations are evicted
oldest-first past 128 blocks or 256MB of retained blob payload; blobs
stay shared with the committer's output, so the budget measures
extended lifetime rather than copies. Lookups go through one merged
map — cost independent of depth — with per-entry generation tags so
evicting an old generation never drops a key a newer one refreshed.
Hash-keyed lookups keep correctness independent of ring contents;
abandoned-fork entries are structural misses, so reorgs need no
invalidation.
The reader wrapper now accepts a WarmNodeSource (WarmSnapshot or
WarmNodeRing), leaving the witness-on prefetcher-snapshot path
untouched. The carry plumbing through collect/spawn is removed: each
SRC adds its commit nodes to the ring and every SRC opens against it.
New warm_ring/{nodes,bytes,generations} gauges report fill level.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, eth, internal/cli: allow disabling exec prefetchers on pipelined import
CPU profiling of mainnet catch-up with pipelined witness-off import
attributed ~22% of process CPU to the two execution-side prefetch lanes:
the speculative block prefetcher (a full second execution of every block
on a throwaway StateDB, 16%) and the executing StateDB's trie prefetcher
(6%), whose output is detached and discarded on this path — the SRC
takes its warmth from the WarmNodeRing instead. Both lanes also drive a
large share of pathdb diff-layer walks and pebble read contention.
Add pipeline.exec-prefetch (default true, current behavior). When
disabled, pipelined witness-off imports skip both prefetchers; the
shared per-block caches remain and are populated by the BlockSTM
workers. Non-pipelined imports, witness-on imports, and the miner are
unaffected. All downstream prefetcher interactions (detach, stop, V2
settlement origin-read assist) are nil-safe, so the absence degrades
cleanly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: regenerate CLI docs
make docs picks up the previously undocumented gRPC token flag on the
peers/debug/chain/removedb subcommands.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, eth, internal/cli: scope exec-prefetch flag to the trie prefetcher
Benchmarking pipeline.exec-prefetch=false as shipped regressed catch-up
throughput ~25%: the speculative block prefetcher is load-bearing for
the execution lane. It runs ahead of the BlockSTM workers, pulling the
block's state reads into pebble/OS caches and pre-filling the shared
jumpdest caches; without it those reads go to cold storage inline and
execution p50 rose 118ms to 152ms even though process CPU dropped from
57% to 34% and the SRC lane got faster. CPU profiles cannot see this
benefit — cache warming manifests as avoided latency elsewhere.
Narrow the flag to the executing StateDB's trie prefetcher only (~6% of
process CPU), whose output really is detached and discarded on the
pipelined witness-off path. The speculative block prefetcher now always
runs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* triedb/pathdb: index diff-layer trie nodes for direct lookup
Trie-node reads resolved by walking the diff-layer chain, probing every
layer's node maps until a hit or the disk layer — up to maxDiffLayers
map probes per read. Mainnet catch-up profiling attributed 44% of all
process CPU to this walk (26% pure map probing), paid by every node
consumer: the SRC goroutine, trie prefetcher subfetchers, the
speculative block prefetcher, and RPC.
States already avoid this via the lookup structure; give nodes the
equivalent. A reference-counted index maps (owner, path, hash) to the
node blob for every live diff-layer node, maintained alongside lookup
under the layerTree lock on layer add/cap/flatten/removal. Node
requests always carry the expected hash (parent nodes embed child
hashes), so the triple is content-addressed: a hit is correct no matter
which layer or fork supplied it, and a miss guarantees the node is in
no live diff layer, letting the reader go straight to the disk layer.
The chain walk survives as the authority for noHashCheck readers and
irregular cases (stale layers, hash mismatches) so correctness never
depends on the index. Blobs are shared with the diff layers' node sets;
the index adds only map overhead. Reference counting keeps entries
alive while any sibling fork still owns them; deletion markers are not
indexed (post-deletion parents never reference them).
New metrics: pathdb/nodeindex/{hit,miss,count,bytes}.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, state: remove the witness-off SRC warm node ring
The pathdb node index serves every trie node the ring held (committed
nodes live in diff layers, which the index covers in one probe), and
feeding the ring ran in the SRC goroutine after each commit, costing
lane time and ~255MB of long-lived map churn. Mainnet catch-up A/B on
adjacent block ranges: ring off 334.5 mgasps / src p50 59.8ms vs ring
on 316.8 / 68.7ms.
The witness-on WarmSnapshot handoff is unchanged; pipeline.warm-snapshot
now gates only that path and is inert for witness-off import.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core: skip exec trie prefetcher when nothing consumes it in witness mode
With pipeline.exec-prefetch disabled, the executing StateDB's trie
prefetcher now runs only when the witness-on warm-snapshot handoff
consumes its output. Witness-on catch-up profiling showed it at ~15% of
process CPU with the handoff disabled, all discarded; the SRC captures
the witness on its own trie-only walk, so completeness never depended
on it.
The execution witness previously reached the StateDB only through
StartPrefetcher, so skipping the prefetcher dropped it and SRC failed
hard. Attach it via SetWitness on both processor branches instead. The
witness parity suites now import a third no-prefetcher chain and assert
root parity, proof-node-set parity, and stateless replay against the
baseline.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, state: let V2 honor the caller's trie-prefetch decision
BlockSTM V2's settle phase unconditionally swapped in its own
"v2-settle" trie prefetcher, resurrecting the prefetcher on every block
and making the exec-prefetch gate a no-op — profiling showed identical
prefetcher CPU with the flag on and off. V2 now performs the swap only
when the caller installed a prefetcher (new StateDB.HasPrefetcher
accessor), so execTriePrefetchEnabled's no-consumer decision finally
takes effect. Witness collection is unaffected: the witness rides on
the StateDB via SetWitness, and non-pipelined imports always run the
prefetcher for their own root computation.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, eth, internal/cli: drop the pipeline.exec-prefetch flag
Catch-up benchmarking showed the exec-side trie prefetcher is
load-bearing in both witness modes: its trie reads warm the shared
pathdb/pebble caches that the SRC goroutine hits moments later, and
disabling it starved witness-mode SRC so badly that collect waits
reached seconds per block. The disable position had no measured benefit
in any mode, so the knob is removed and the prefetcher always runs.
This also reverts the V2 settle-phase HasPrefetcher gate, which existed
only to make the flag effective.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, eth: lint cleanup after exec-prefetch removal and develop merge
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, internal/cli, tests: address review findings and deflake peering
Bound GetWitness's pipelined-witness wait to headers within 64 blocks of
the head so absent or pruned witness hashes return immediately instead of
letting peers tie the handler up for the full poll timeout. Switch the
WarmSnapshot key to a stack-built fixed-size array — the string-keyed
variant allocated on every lookup since Go's string(bytes) map-index
optimization does not cover struct-literal keys. Pre-initialize the
Pipeline block in readConfigFile so HCL config files omitting [pipeline]
no longer panic at startup (TOML files were unaffected). Replace the
fire-and-forget AddPeer pairs in the pipelined import integration tests
with a self-healing connect helper: Self() publishes its TCP port
asynchronously, so a single AddPeer could capture a port-0 enode and dial
a dead address forever — the CI failure mode of BasicImport.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, eth, internal/ethapi, tests: RPC correctness under pipelined import
Verify the read-only RPC surface on a node importing with pipelined SRC,
and close the one gap found: handlers that open tries directly by root
(eth_getProof, debug_storageRangeAt) failed transiently when queried at
the head whose SRC hadn't committed yet, since proofs need trie nodes the
FlatDiff overlay cannot provide.
- WaitForPipelinedStateCommit: bounded wait on the pending import's
collectedCh; no-op for any root other than the pending head's. Wired
into eth_getProof (via a new ethapi.Backend method) and
debug_storageRangeAt, converting the transient error into a few ms of
latency. Overlay-served value reads are untouched.
- core/pipelined_window_test.go: holds the SRC goroutine open via a
test-only hook to pin window semantics deterministically — overlay
reads correct, raw trie opens fail, wait gate blocks then releases,
post-settle values identical, proofs verify. Both schemes, race-clean.
- tests/bor/pipelined_rpc_test.go: BP -> pipelined importer sync with a
live tx stream; strict 12-method battery at pinned heights and latest
(getProof now strict), sync-time vs settled response identity, full
per-height parity vs a non-pipelined BP incl. cryptographic account
proof verification.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* consensus/bor, core, core/state: extract helpers from new pipeline code
Bring the new-code diffguard complexity violations under threshold and
remove production-code duplication:
- statedb: resolveFlatMutationObject extracted from applyFlatMutationFast;
the two witness-collection loops in IntermediateRoot deduped into
addObjectWitness/addWitnessNodes.
- blockchain: preload read-surface histograms + preloadFlatDiffReads
extracted from runSRCCompute into recordAndPreloadSRCWitnessReads.
- bor: FinalizeForPipeline now reuses commitSprintWork (extracted on
develop by #2314) instead of duplicating the sprint-start span and
state-sync commits; this also sets BorConsensusTime on the pipelined
path, matching FinalizeAndAssemble.
Behavior-preserving; legacy-grown giants (insertChainWithWitnesses,
IntermediateRoot, getStateObject, updateTrie) are declared as explicit
deviations in the PR description.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, core/state, core/txpool, eth, tests: address review findings
- PostExecState: match the cached FlatDiff by state root in addition to
block number (zero stored root — the miner's speculative path — keeps
the number-only match), so a same-height reorg can't serve a stale
overlay.
- waitForPendingSRCWitness: cap the collectedCh wait so a stalled SRC
goroutine can't pin peer/RPC handlers indefinitely.
- resolveWitnessSizes: fetch via new GetWitnessUncachedWait (bounded SRC
wait without inserting peer-driven reads into the witness cache) and
cap retained prefetch bytes at MaximumResponseSize.
- updateTrie: surface witness re-read GetStorage failures via setError
instead of silently producing an incomplete witness.
- minedBroadcastLoop: track the announceMinedBlock goroutine on h.wg so
handler shutdown doesn't race peer close.
- SetSpeculativeState: publish NewTxsEvent after releasing pool.mu,
matching the regular promotion flow.
- pipelined_rpc_test: assert fdlimit.Raise succeeds.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* consensus/bor, tests, triedb/pathdb: fix CI peering failure, review nits
The three pipelined integration tests failed on CI with both nodes at
zero peers for the full 60s window. Root cause: the dial scheduler
dedupes static nodes by ID (addStaticCh handling ignores re-adds), so
when the first AddPeer captures a port-0 enode — Self() publishes its
TCP port asynchronously after the listener starts — the "self-healing"
re-add loop never updates the record and the dialer keeps dialing a dead
address forever. On top of that, a transiently failed dial parks the
target in dial history for 35s, so a 60s deadline covers barely one
retry. connectAndWaitForPeers now waits for both enodes to publish a
real port before the first AddPeer and allows several history windows.
Also addresses review findings:
- bor_test: the waitOnPrepare subtests assert against the parent slot
boundary instead of fixed elapsed-time bounds, so loaded runners
can't flake them.
- pathdb: nodeFallback goes straight to tree.bottom() — nodeWalk's
primary attempt already is the entry-layer walk, so retrying it
deterministically hits the same stale disk layer (unlike
account/storage fallbacks, whose primary attempt is the lookup-index
path).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* miner, tests: guard GetTd underflow, tighten RPC test setup
- speculative_chain_reader: guard the parent-number subtraction in GetTd
against a genesis parent (can't happen on the speculative path, but an
underflow to MaxUint64 is worth a two-line guard).
- pipelined_rpc_test: handle crypto.GenerateKey errors; drop the
unbounded listener-wait loops — connectAndWaitForPeers now waits for
published listener ports under a deadline.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core: skip witness cache poll on witness-off nodes, merge warmKey doc
waitForPipelinedWitness gated its 2s cache poll on header existence and
head recency, but not on whether the node produces witnesses at all. On
a pipelined node importing with makeWitness=false, no witness ever lands
in the cache, so a single GetWitness request naming recent existing
hashes could park the WIT handler for the full poll timeout per hash.
Latch the most recent pipelined import's makeWitness and skip the poll
when it is false; the pending-SRC fast path already handles the
witness-off pending block. Extends the MakeWitnessFalse test to assert
the fast miss.
Also merges warmKey's duplicated doc comment into one block.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* tests: restore global logger on cleanup, use NoErrorf
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, miner: cover pipelined SRC branches
* core, triedb/pathdb, eth: fix pipelined-import failure paths
Addresses the review findings on the async collector's failure handling,
plus one issue independent of the feature flag.
triedb/pathdb: a lookup-index rejection is authoritative again — it means
the reader's state is neither the disk layer nor a descendant of it, so
the disk layer's flat data belongs to a different state. Falling back
there could serve a flattened-past root's newer values, or a reorged-away
branch its canonical sibling's writes, with no hash to catch it (unlike
trie nodes). The remaining fallback covers the genuine race where the
located layer goes stale mid-read, and now only reads the base disk layer
when bottomIfAncestorOf confirms it is an ancestor of the reader's state.
core: the SRC collector no longer touches chainmu. syncx.ClosableMutex's
TryLock is a blocking receive that only reports failure once the mutex is
closed, and the thread that collects holds chainmu for the whole batch
while waiting on collectedCh — so the mismatch path's TryLock deadlocked
both sides permanently, wedging every later insert and Stop(). Recovery
moved to the collecting thread as recoverFailedPipelinedImport, which now
runs for SRC errors too (previously only mismatches attempted anything,
leaving an unverified block as permanent head) and completes the
rollback: clears the pending entry, drops the FlatDiff overlay, deletes
the rejected block's canonical hash and tx lookups, moves the head back,
dereferences the divergent root in hash mode, and emits a corrective
ChainHeadEvent so the txpool re-resets off committed state.
core: clamp the insert index at 0 when attributing a collect failure back
to the previous block — first-block-of-batch is the normal cross-batch
shape, and a negative index reached the downloader's upper-bound-only
guard and indexed out of range. Guard hardened there as well.
core/state: track this block's own destructs separately from the parent
destructs the FlatDiff replay paths seed, and consult that set before the
storage overlay. The overlay holds the parent's post-state, so a
same-block destruct-and-recreate was reading pre-destruct storage where
every other node reads zero. Reordering the existing check would have
zeroed legitimate parent destruct-and-resurrect slots instead.
core: a supplied witness now disables pipelining for that block and
flushes any in-flight SRC first — the witness memdb replaces the shared
trie read backend, which starves both the running SRC and this block's
own parent-root reads.
core: a locally unavailable parent state no longer records a valid block
as bad; the removed state.New pre-check used to return that class early.
Tests: root-mismatch rollback over a two-block batch (hangs without the
collector fix), in-block destruct over an overlay (returns pre-destruct
storage without the guard), a positive control asserting the pipeline
actually ran, and TestV2GasDeterminism now fails rather than skips when
every fixture misses base reads.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, core/state: judge worker base-read failures per settled incarnation
The develop merge brings #2333's witness completeness fix into the
pipelined branch, and this commit resolves the collision between that
work and the branch's base-read error tracking. SafeBase read methods
now return errors instead of promoting every pooled-copy failure into
one block-global error: a speculative incarnation chasing stale values
may legitimately read outside the canonical state (on witness-backed
replay such state simply doesn't exist) and is then invalidated, so a
block-global error over-triggers. Each ParallelStateDB records its own
BaseReadErr, cleared on Reset.
The fatal check runs inside the settle callback, the only point where a
pdb is provably the settled incarnation. Scanning the executor's states
array after the fact reads recycled pool objects: settlement returns
each pdb to the pool, a later tx's Reset reuses it, and stale slots
alias the mutated object — a speculative error from a later tx then
surfaces under earlier indices (observed as ten aliased slots reporting
one invalidated incarnation's miss), while a genuinely settled error can
vanish before the scan.
The witness regeneration oracles are anchored against each block's real
mainnet roots; previously the round trip compared only against its own
stage-1 replay, so a fixture whose replay silently diverged from mainnet
would pass against itself. The settle-time gate then exposed three
fixture blocks whose witness/code archive lacks an EIP-7702 authority's
pre-state code blob (validateAuthorization reads the code to parse the
delegation; the fixtures were captured by a client whose V2 base reads
silently nil-served the miss). Those are pinned by exact hash as
known-incomplete fixtures. Serial replay of them still anchors only
because a missing blob reads as empty code, which accepts the
authorization just like the real delegation designator does.
Also constructs the detached-prefetcher warm-snapshot test through the
real constructor: a hand-rolled triePrefetcher literal has nil meters,
and report() panics when another test's global metrics.Enable() wins
the race.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core/stateless: validate peer-supplied witness header numbers explicitly
ValidateWitnessPreState trusted the transport decoder to guarantee a
non-nil, non-genesis block number on the witness context header. RLP
decoding does guarantee non-nil today, but this function is the boundary
check for peer-supplied witnesses — it shouldn't inherit its input
invariants from whichever decoder happened to run. A nil number would
panic on Uint64(), and a genesis number would underflow the parent
lookup to MaxUint64 (harmless miss, but an unrelated probe). Both now
return validation errors, and the parent number is computed once.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core: roll back failed pipelined imports on every chainmu flush path
recoverFailedPipelinedImport was reachable only from the collect path;
the reorg/gap, ProcessBlock-error and witness-fed flushes cleared the
pending entry and returned the error without undoing the published head,
leaving the rejected block canonical while the import continued around
it. flushPendingImportSRC now takes a rollback flag: the three callers
that hold chainmu pass true; shutdown passes false (it cannot take
chainmu, and the startup rewind covers it).
The rollback also emits RemovedLogsEvent for the rejected block's logs
(collected before its indexes are deleted), matching the retraction a
reorg sends, so filter and subscription clients drop them.
insertChainWithWitnesses returns -1 instead of clamping to 0 when a
cross-batch SRC failure has no previous in-batch index: both consumers
treat out-of-band indices as "failure outside this batch" and only log,
whereas the clamp blamed (and bad-block reported) the batch's first
block for a previous batch's failure.
TestFlatDiffOverlay_ParentDestructKeepsOverlayStorage pins why same-
block and parent-block destructs live in separate sets: a parent-block
destruct seeded by the FlatDiff replay paths must not shadow overlay
storage written after the resurrection.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* core, miner: seal only on committed state, never the FlatDiff overlay
Found by running the kurtosis e2e topology with the pipeline enabled on
validators: the network deterministically split at the first sprint
boundary. commitWork opened sealing state through StateAtWithReaders,
which serves the FlatDiff overlay whenever the parent is the latest
pipelined import — an overlay statedb is rooted at the grandparent with
the parent's writes installed as unjournaled read-only objects, so the
root sealed into the header omits every parent write the new block does
not rewrite (on the devn…
Backmerge v2.10.2 (master) into develop. One conflict in core/state/statedb_test.go: both sides added tests at the same spot; kept TestLogCount (develop) and the two CopyWithoutLogHistory tests (master).
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.
Tip: disable this comment in your organization's Code Review settings.
Codecov Report❌ Patch coverage is ❌ Your patch check has failed because the patch coverage (88.54%) is below the target coverage (90.00%). You can increase the patch coverage or adjust the target coverage. Additional details and impacted files@@ Coverage Diff @@
## develop #2473 +/- ##
===========================================
+ Coverage 55.62% 57.60% +1.97%
===========================================
Files 918 959 +41
Lines 167160 176094 +8934
===========================================
+ Hits 92978 101433 +8455
- Misses 68714 68984 +270
- Partials 5468 5677 +209
... and 23 files with indirect coverage changes
🚀 New features to boost your workflow:
|
Backmerge
master(tagv2.10.2, merged fromv2.10.2-candidatein #2466) intodevelop.This is a merge commit, not a squash. Please merge it with a merge commit so develop keeps master's history.
Conflict
Only one:
core/state/statedb_test.go. Both sides added new tests at the same place. I kept all three tests:TestLogCount(develop)TestCopyWithoutLogHistoryPreservesLogIndexandTestCopyWithoutLogHistoryKeepsLogsForDirtyJournal(master)Checks
go build ./...okgo test ./core/state/ -run 'TestLogCount|TestCopy'okNot included
7646fe031from the v2.10.1-hotfix2 line (core/blockchain_statesync_test.go) is not on master, so it is not in this PR. The fix itself is already on develop through #2371. Only its regression test is missing.🤖 Generated with Claude Code