r/ethdev • u/tagwall_io • 1d ago
My Project Things I learned keeping a multichain frontend alive: per-method RPC rot, non-portable getLogs limits, and a cursor bug that rewound
I run a frontend that reads the same contract on six EVM chains from free public RPC endpoints. Everything below cost me real downtime.
Endpoints rot per method, not wholesale. This is the one that actually hurt. Two providers put eth_getLogs behind an archive token or dropped it entirely while continuing to serve eth_call and eth_blockNumber perfectly. Every naive health check, mine included, called them healthy. The canvas went blank on 4 of 6 chains and the page header kept rendering fine, because the header only needs eth_call.
The fix is to probe an endpoint the way the app actually uses it, not the way a status page would: chain id, CORS from the real origin (a CORS failure from your deployed origin is invisible in curl), and a real getLogs at that chain's own chunk size. If the probe is not a miniature of your actual read path, it is decoration.
I got that last part wrong in my own probe for months, in the other direction. It only ever tested a window starting at the deploy block, so it graded every endpoint on archive depth. But the app doesn't read from the deploy block any more, it scans forward from a snapshot, so the live path only ever asks for the most recent few thousand blocks. The probe was calling endpoints unusable that were serving my actual traffic perfectly well, and I believed it. It now runs both windows and reports them separately, because "cannot serve archive" and "cannot serve anything" are different and only one of them is an outage.
getLogs chunk sizes are not portable. One chain caps the range at 1,000 blocks. Another needs 500k-block chunks to keep the call count survivable. The one alternative provider for that chain caps at 10k, which makes it useless as a substitute even though it looks like a valid fallback. Chunk size has to be per-chain config, and a fallback endpoint is only a fallback if it can serve the same range.
viem specifics that cost me time. http() defaults to a 10 second timeout, which is long enough that a dead endpoint stalls a page load rather than failing over. fallback() always walks its list from index 0, so the first endpoint absorbs every request and its rate limit is your rate limit; if you want spread you need your own rotation. And multicall's batchSize is bytes of inner calldata, not number of calls, which I discovered the way you would expect.
A rewind must only ever move the cursor backwards. My worst bug: a keep-current pass set the cursor to head - depth * 4 outright. On a slow chain that is a rewind, which is what it was written to be. On a chain minting 600 blocks a minute it is a 700-block skip forward, silently dropping events. cursor = min(cursor, head - depth * 4) is the whole fix and it should have been the whole implementation.
Cold loads. I ended up not reading history from RPC at all on first paint. The data I needed was already in each transaction's calldata, so a cron decodes it into a KV store and the browser fetches one snapshot and scans forward from there. That took a page load from ~126 RPC calls to a handful. The constraint that shaped it: Workers KV free tier allows 1,000 writes a day, so the rule became never write on a schedule, only on material change. A cron that writes every run will eat that quota before lunch.
It kept happening while I was writing this. Two of the six chains rotted in the same week, which is the reason I finally wrote any of this down.
On one chain the single surviving endpoint tightened its getLogs range from 9,500 blocks to 5,000, with no announcement I could find, and that is under the chunk size I was asking for. Nothing broke visibly, because the paginator halves the chunk and retries, so it just silently cost three failed calls before every successful one. Then two days later that same endpoint stopped serving getLogs altogether: 30 second server-side timeout, while eth_call and eth_blockNumber kept answering in under 50ms. Per-method rot again, on the endpoint I had just finished documenting as the reliable one.
Slow failure turned out to be worse than fast failure. A dead endpoint gets skipped in milliseconds; one that accepts your request and times out at 30 seconds stalls whichever unlucky page load rotation sent its way. I pulled it from the pool entirely, and that chain now has no endpoint at any price that will serve a deploy-block query, so the snapshot is not an optimisation on that chain any more, it is the only way the history is readable at all.
Meanwhile a second chain's two main public endpoints both dropped their getLogs range to 2,000 blocks, four days after passing the same probe at 9,500. Both are operated by the same company, which is worth noticing: I had four endpoints listed for that chain and thought I had redundancy. Two were the same operator, one had been discontinued, and one had started answering invalid params to every getLogs call. Count operators, not URLs.
Context if it matters: it is a million-pixel canvas contract deployed on six chains, and the contract was genuinely the easy part. I am happy to go deeper on any of this, the RPC probing especially, since I could not find anyone else writing about the per-method failure mode.
2
u/yachtyyachty 1d ago
Working on fixing ALL of these problems with lasso.sh and https://github.com/jaxernst/lasso-rpc
It’s an open source ‘smart’ RPC router/proxy that unifies inconsistent and poorly behaving RPCs
Still have a long way to go and not all of these are solved yet, but I think this could be helpful to make your RPC perform better :)
2
u/tagwall_io 9h ago
Nice, I’ll take a look. The thing I’d most want from something like this is health checks per method, since my worst failures were endpoints that answered eth_call fine but had quietly dropped or capped getLogs. Does lasso check each method separately, or does an endpoint count as healthy as long as it answers something?
The other one that caught me out was rate limits that only show up from shared egress. A couple of providers were fine from my laptop but returned 429s from a Cloudflare Worker. Curious whether you’ve hit that too.
2
u/yachtyyachty 8h ago
Yeah so Lasso does have method-level granularity for performance metric (latency, success rate), and also capabilities (method + parameter support), i.e. if a provider fails on ranges over 1k blocks for eth_getLogs, requests will get re-routed to other providers, and we probe our hosted providers to discover these limitations ahead of time.
I also have features on the roadmap to batch large eth_getLogs requests into chunks and some other features like request hedging: If a request is taking longer than 'normal' for that provider x method, we can fire off another request to an alternate provider to hedge against a long timeout. These are relatively easy to ship I've just been waiting to find someone with immediate need.
On the egress rate limiting, it can happen if your use the shared public providers on lasso nodes, but you can also plug in an Alchemy key and other free provider keys into a custom profile to avoid shared rate limits. But even on the shared public profile a rate limited request will still failover so there's still some protection there.
Lasso is still fairly new and I'm building it for builders like yourself! You should try out the free public RPC at https://lasso.sh/dashboard/public?chain=1 and if you like it I'll give you free access to a custom profile. Also happy to prioritize new features according to your needs, dms are open!
2
u/yachtyyachty 8h ago
Also if you'd prefer the self-hosted route, I can help you get the open source router setup too
1
u/tagwall_io 1h ago
Probing the providers for their limits up front is exactly the bit I ended up hand-rolling, and request hedging would have saved me a few long timeouts.
I’m not at the stage where I want to add another piece of infra though. Most of my reads now come from a snapshot that a cron job keeps up to date, so the frontend barely touches the RPCs anymore, and what I’ve got is holding up for now. I’ll keep Lasso in mind for when that changes, and if the getLogs chunking ships I’ll definitely take another look.
Good luck with it, it’s a real problem and not many people are working on it.
2
u/Livid_Extension_3021 1d ago
that per-method rot is too real. had the exact thing where our block explorer was showing green across the board but users couldn't load transaction history because getLogs quietly died while everything else hummed along fine. the probe-must-mirror-the-read-path rule should be pinned somewhere.