Skip to content

Restore eth_getLogs range reduction after alloy migration - #6729

Draft
datanexus-vincent wants to merge 1 commit into
graphprotocol:masterfrom
datanexus-vincent:fix/eth-getlogs-range-reduction
Draft

datanexus-vincent wants to merge 1 commit into
graphprotocol:masterfrom
datanexus-vincent:fix/eth-getlogs-range-reduction

Conversation

@datanexus-vincent

Copy link
Copy Markdown

Problem

log_stream shrinks the eth_getLogs block range when a provider rejects a request as too heavy. It detected that by substring-matching the rendered error against a list that included ServerError(-32005) and ServerError(-32000). Those are rust-web3's Debug rendering of the error code.

Since the alloy migration in v0.42.0 the error is an alloy RpcError, which renders as error code -32005: ... or ErrorResp(ErrorPayload { code: -32005, .. }). The code fingerprints never match. A provider with a range or result cap below GRAPH_ETHEREUM_MAX_BLOCK_RANGE_SIZE now gets ten identical retries, Unexpected RPC error, and a block stream restart, and the subgraph never advances. Besu with --rpc-max-logs-range is the simplest reproduction.

Change

The substring list is replaced by a structural classification of the alloy error with three outcomes:

Outcome Trigger Behaviour
Too heavy A known cap message Shrink the range without retrying
Maybe too heavy Code -32000, -32002, -32003 or -32005 with no known message, or a bare HTTP 503 Retry, then shrink if it still fails
Unrecognized Everything else Retry, then surface the error, as before

Caps are matched by message, not code. geth and reth report a cap under -32602 and Alchemy under -32600, which they also use for genuinely malformed requests, so the code alone cannot be trusted.

Cap messages are also read from HTTP 400, 403, 413 and 503 bodies, and from 200 bodies alloy failed to deserialize. graph-node's HTTP transport turns any non-2xx response into an HttpError carrying the raw body, so a provider that pairs its cap with an error status was otherwise invisible. Ankr does this with a 413.

Covered, each with its verbatim message in the tests: geth, reth, Besu, Alchemy, Infura, QuickNode, Ankr, zkSync Era, Monad.

Behaviour changes to be aware of

  • -32000, -32005 and HTTP 503 are no longer shrunk on immediately. The pre-alloy list skipped retries for these. They also cover transient failures, for example geth's header not found from a lagging node, Infura's rate limit, and a gateway with no backend. They are now retried first and shrunk only if the failure persists.
  • -32002 and -32003 are new. These are geth's request timeout and response-too-large codes. They get the same retry-then-shrink treatment when the message is not recognised.
  • A cap under an unlisted message and an ambiguous code converges more slowly. It pays the full retry schedule, about three minutes at the defaults, before each shrink. Adding the message to the list removes the delay.
  • A cap under an unlisted message and any other code does not shrink. That matches the behaviour before this change.
  • A client-side timeout still does not shrink the range. That is unchanged.

The reduction policy itself, dividing the step by ten down to a single block, is untouched.

Testing

Unit tests cover every message fragment in isolation, the verbatim provider messages under their real codes, the ambiguous codes, unrelated errors, cap messages inside HTTP and undeserializable bodies, and statuses whose bodies must not be read.

cargo test -p graph-chain-ethereum, cargo clippy -p graph-chain-ethereum --all-targets and cargo check -p graph-chain-ethereum --release pass. The provider messages come from client sources and provider documentation, not from live requests.

log_stream shrinks the block range when a provider rejects eth_getLogs
as too heavy. It recognised that by substring-matching the rendered
error, including "ServerError(-32005)" and "ServerError(-32000)", which
are rust-web3's Debug rendering. Since the alloy migration in v0.42.0
those strings never appear, so a provider enforcing a range or result
cap got ten identical retries, "Unexpected RPC error", and a block
stream restart loop instead of a smaller range.

Classify the alloy error structurally instead:

- A known cap message shrinks the range at once, without retrying.
  Caps are matched by message because most providers report them under
  a generic code such as -32602.
- An ambiguous code (-32000, -32002, -32003, -32005) or a bare HTTP 503
  is retried first and shrinks the range only if it keeps failing.
  These also cover transient failures such as a lagging node or a rate
  limit, which the old list shrank on immediately.
- Anything else is retried and surfaced, as before.

Cap messages are also read from HTTP 400/403/413/503 bodies and from
200 bodies alloy could not deserialize, since the transport hides the
JSON-RPC error in both cases.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants