Skip to content

fix(evm-rpc): retry transient provider and proxy errors instead of failing - #584

Merged
mo4islona merged 2 commits into
masterfrom
fix/evm-rpc-retry-transient-errors
Sep 29, 2026
Merged

mo4islona merged 2 commits into
masterfrom
fix/evm-rpc-retry-transient-errors

Conversation

@mo4islona

@mo4islona mo4islona commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Several transient failures from RPC providers and load-balancing proxies were classified as fatal. evm-dump exited on them and the data service restarted ingestion, although the same call succeeds when repeated. The error shapes, with handling before and after this change and the ones left fatal on purpose, are listed in #583.

Changes

  • rpc-client: HTTP 520–524 (Cloudflare origin errors) are retried like 502/504.
  • evm-rpc, EvmRpcClient.isConnectionError:
    • rate limits: code -32429, and the messages throughput limit and exceeded … capacity;
    • transient errors: upstream and aggregator failures, messages that say the failure is temporary, and "block not there yet" errors from a backend lagging the head. "Temporary", "please retry" and "at this time" count only when neither the JSON-RPC code (-32700, -32600, -32601, -32602) nor the message blames the request, so advice like "please retry with a smaller block range" stays fatal;
    • a proxy error is judged by the causes nested under data, not by its outer code and message, which the proxy normalizes. An unsupported-method, ignored-method, unauthorized, client-side or execution cause keeps the error fatal;
    • an HTTP 500 whose JSON-RPC body is such an error; for a batch body, only when every error in it is transient, whatever their order. A bare 500 is still governed by retryInternalServerErrors.
  • evm-rpc, result validation: null where a value is due, or a non-hex string in result, throws RetryError instead of DataValidationError. Malformed objects still fail. See the trade-off below.

Metrics

Errors that used to stop evm-dump or restart ingestion are now retried, so an endpoint that keeps failing no longer shows up as restarts. To keep that visible:

  • sqd_chain_rpc_retried_errors_total{url, kind}, in evm-dump and the EVM data service.
    • kind is one of rate_limit, transient, no_result, internal, timeout, http, connection. Outside EVM it can also be retry or other.
    • url is redacted like in the other RPC metrics.
    • Only errors the client actually retries are counted. The error that fails a request after its attempts are spent is not, and neither is a failed batch that reduceBatchOnRetry splits in halves instead of retrying.
    • In the data service the RPC clients run in worker threads. Their counts are read on each scrape, with a 1 s timeout per worker. A worker that does not answer keeps its last counts, and the final counts of closed backfill workers are kept, so the counter does not go down.
  • sqd_hotblocks_ingestion_restarts_total{reason}: reason is the name of the error that stopped ingestion, or fork, or ended.
  • API: RpcClient.getRetryKind() and getMetrics().retriedErrors in rpc-client. EvmRpcClient refines the kinds.

Trade-off

null or error text in result means the node does not have the data. Usually that is a backend lagging the head, which a retry fixes, but it can also be data the node will never have, for example pruned transaction history. The response does not tell these apart.

These retries draw on the client's shared retry budget, which evm-dump sets unbounded. So on permanently missing data the dump now retries with backoff, logging each attempt as a connection failure and counting it as no_result, instead of exiting. In the data service the budget is 5 attempts, after which the error surfaces as before. A per-request cap on no-data retries is left as a TODO in getResultValidator.

Related

Tests

  • evm/evm-rpc/test/rpc.transient-errors.test.ts:
    • every observed error shape, transient and permanent;
    • proxy errors with nested causes;
    • HTTP 500 bodies;
    • through a local server: error text or null in place of a block throws RetryError, a malformed block throws DataValidationError, and the call recovers on retry;
    • retry kinds, and the counter after a recovered call.
  • util/rpc-client/src/client.connection-error.test.ts: HTTP status classification, retry kinds, and the counts against a local server: every retry is counted, while the error that fails a request, or any error when retries are off, is not.
  • util/util-internal-data-service/src/rpc-metrics.test.ts: summing across workers, keeping the counts of a closed worker, a hung or dead worker, a worker that stops answering keeping its last counts, a worker removed during a scrape counted once, the Prometheus output, and the last values surviving a failed read.
  • util/util-internal-data-service/src/data-service.test.ts: a failed ingestion session increments the restart counter by error name.
  • util-internal-dump-cli and evm-data-service have no test setup. Their wiring is checked by compilation and by the tests of the pieces they call.
  • Results:
    • rpc-client: 34 passed;
    • evm-rpc: 325 passed, 29 skipped;
    • util-internal-data-service: 14 passed;
    • tsc: clean for rpc-client, util-internal-data-service, evm-rpc (with tsconfig.test.json), evm-data-service and util-internal-dump-cli.

🤖 Generated with Claude Code

mo4islona and others added 2 commits September 30, 2026 00:00
Providers and load-balancing proxies report transient failures in shapes the
classifier did not recognize: a rate limit under code -32429, error text or
null as the `result` of a 200 response, upstream or head-lag failures under
generic codes, and Cloudflare 520-524. evm-dump exited on them and the data
service restarted ingestion, although the same call succeeds when repeated.

Plain errors are classified by message, and "please retry"-like phrases count
only when neither the code nor the message blames the request. A proxy
normalizes its outer code and message, so its errors are judged by the causes
nested under `data`; an unsupported or ignored method, or rejected
credentials, stays fatal.

Closes #583

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Errors that used to stop evm-dump or restart real-time ingestion are now
retried, so a dump stuck on a persistently failing endpoint no longer shows
up as restarts. Counting retried errors by kind, per endpoint, keeps that
condition alertable. In the data service the RPC clients live in worker
threads, so their counts are read on each scrape and the final counts of
closed workers are kept.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@mo4islona
mo4islona force-pushed the fix/evm-rpc-retry-transient-errors branch from f4b33c1 to 9515ca6 Compare September 29, 2026 21:00
@mo4islona
mo4islona merged commit 25ad162 into master Sep 29, 2026
2 checks passed
@mo4islona
mo4islona deleted the fix/evm-rpc-retry-transient-errors branch September 29, 2026 21:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

evm-rpc: in-body rate-limit error with code -32429 is not retried and crashes the dump

1 participant