Skip to content

fix(rpc-client): retry rate-limit responses returned as an HTTP 200 string result - #582

Closed
SQD-Trevor-Agent wants to merge 1 commit into
masterfrom
alert-fix/Tr2KEE-ratelimit-200-retry
Closed

SQD-Trevor-Agent wants to merge 1 commit into
masterfrom
alert-fix/Tr2KEE-ratelimit-200-retry

Conversation

@SQD-Trevor-Agent

Copy link
Copy Markdown

Cause (proven)

dump-tac-mainnet-0 was in CrashLoopBackOff (721 restarts). The archive's highest block was frozen at 25,979,220 while the chain head is 26,125,782 (~146k behind), so the hotblocks rolling window advanced past the frozen archive block and the portal_hotblocks_gap alert fired for tac-mainnet.

The fatal error, from the dump pod's previous-container logs:

DataValidationError: server returned unexpected result: "You've been rate limited, please upgrade your plan.\n" is not an object
    at /squid/evm/evm-rpc/lib/rpc.js:1014:19
    at EvmRpcClient.receiveResult (/squid/util/rpc-client/lib/client.js:384:24)
  rpcUrl:    http://erpc.erpc:8082/main/evm/239   (erpc → uniblock)
  rpcMethod: eth_getTransactionReceipt
  rpcResponse: {"jsonrpc":"2.0","id":549,"result":"You've been rate limited, please upgrade your plan.\n"}

When rate-limited, the upstream (uniblock, via the erpc proxy) returns HTTP 200 with the JSON-RPC result set to a plain rate-limit string instead of a 429 status or a JSON-RPC error object. Because there is no res.error and the status is 200, the response bypasses every retryable path in RpcClient (isConnectionError handles 429s, RetryError, rate-limit RpcErrors, etc.) and goes straight to the result validator. getResultValidator(Receipt) then throws a non-retryable DataValidationError (… is not an object). The dumper runs with retryAttempts: MAX_SAFE_INTEGER, so a genuine 429 retries forever with backoff — but this validation error propagates and crash-loops the process.

Fix

In RpcClient.receiveResult, detect a rate-limit signal delivered as a 200 string result and throw a RetryError before validation. RetryError is already treated as a connection error, so the existing retry/backoff (and erpc provider failover to ankr) handle it exactly like a 429. Detection is precise (typeof result === 'string' && /rate limit/i, matching the existing isRateLimitError regex), so it only catches rate-limit strings and does not convert genuine data-corruption validation failures into retries.

This lives in rpc-client (not evm-rpc) because the signal is a transport/provider concern that affects every method (blocks, receipts, traces, logs) and every chain, not just the evm receipts path.

Related to #551 (retry transient malformed block results) — same class of "transient upstream response crash-loops the dumper", but a different cause and path: #551 wraps the eth_getBlockByNumber block validator for an object that is missing a field; this handles a rate-limit string result on any method, centrally in the client.

Test (red→green)

util/rpc-client/src/client.rate-limit.test.ts spins up a local HTTP server returning the exact {result: "You've been rate limited…"} 200 response and asserts client.call('eth_getTransactionReceipt', …) rejects with a RetryError, plus unit coverage for the isRateLimitResult predicate.

  • pre-fix: 1 failed — expected Error: server returned unexpected result: … to be an instance of RetryError (reproduces the production crash)
  • post-fix: 10 passed
  • tsc --noEmit clean; rush build --to @subsquid/rpc-client green.

Falsification

If dump-tac-mainnet-0 still exits with a DataValidationError after this change, the malformed result is not a rate-limit string (the regex missed it) — capture the new rpcResponse and widen the signature. If the dump merely lags without crashing, the remaining issue is provider capacity (uniblock rate-limiting), which is a provider/ops matter (swap/upgrade), not this code path.

Provider note (operator action, not part of this PR)

The trigger is uniblock persistently rate-limiting tac-mainnet (chainId 239) via erpc. This code fix stops the crash-loop; if rate-limiting is sustained the archive will still lag. An operator should consider swapping/upgrading the tac-mainnet erpc upstream — but per convention that mitigation is handled out-of-band, not in this PR.

…tring result

Some providers (notably via an erpc proxy) signal rate limiting with an
HTTP 200 whose JSON-RPC `result` is a plain string like
"You've been rate limited, please upgrade your plan." instead of a 429
or a JSON-RPC error object. This bypassed the retry machinery and hit the
result validator, throwing a non-retryable DataValidationError that
crash-looped the dumper. Treat such a result as a RetryError so the
existing retry/backoff (and provider failover) handle it like a 429.
@mo4islona

Copy link
Copy Markdown
Contributor

Superseded by #584, which retries error text delivered as result in general (not only rate-limit messages) together with the other transient error shapes listed in #583.

@mo4islona mo4islona closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants