<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: tagwall</title>
    <description>The latest articles on DEV Community by tagwall (@tagwall).</description>
    <link>https://dev.to/tagwall</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174445%2Ffa9d2d47-3a8e-46f2-b556-10c93f41177e.jpg</url>
      <title>DEV Community: tagwall</title>
      <link>https://dev.to/tagwall</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tagwall"/>
    <language>en</language>
    <item>
      <title>Free RPC endpoints die one method at a time. Here's everything that broke keeping a six-chain frontend alive</title>
      <dc:creator>tagwall</dc:creator>
      <pubDate>Sat, 10 Oct 2026 02:55:44 +0000</pubDate>
      <link>https://dev.to/tagwall/free-rpc-endpoints-die-one-method-at-a-time-heres-everything-that-broke-keeping-a-six-chain-131p</link>
      <guid>https://dev.to/tagwall/free-rpc-endpoints-die-one-method-at-a-time-heres-everything-that-broke-keeping-a-six-chain-131p</guid>
      <description>&lt;p&gt;I run a frontend that reads the same smart contract on six EVM chains, using free public RPC endpoints. Here are a few of the things which cost me downtime, an expensive lesson I hope you can learn from.&lt;/p&gt;

&lt;h2&gt;
  
  
  Endpoints rot per method, not all at once
&lt;/h2&gt;

&lt;p&gt;This is the one that actually hurt the most. Two providers put &lt;code&gt;eth_getLogs&lt;/code&gt; behind an archive token, or dropped it entirely, while still answering &lt;code&gt;eth_call&lt;/code&gt; and &lt;code&gt;eth_blockNumber&lt;/code&gt; perfectly. Every simple health check including mine, called them healthy. The canvas went blank on 4 of 6 chains while the page header kept rendering fine, because the header only needs &lt;code&gt;eth_call&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The fix is to probe an endpoint the way your app actually uses it, not the way a status page would. Check the chain id, check CORS from your real origin (a CORS failure from the deployed site is invisible in curl), and run a real &lt;code&gt;getLogs&lt;/code&gt; at that chain's own chunk size. If the probe isn't a miniature of your actual read path, it's useless.&lt;/p&gt;

&lt;p&gt;I got that wrong in my own probe for months, in the other direction. It only ever tested a window starting at the deploy block, so it graded every endpoint on archive depth. But the app doesn't read from the deploy block any more. It scans forward from a snapshot, so the live path only asks for the last few thousand blocks. The probe was failing endpoints that were serving my real traffic fine. It now tests both windows and reports them separately, because "can't serve archive" and "can't serve anything" are different problems and only one of them can cause an outage.&lt;/p&gt;

&lt;h2&gt;
  
  
  getLogs chunk sizes aren't portable
&lt;/h2&gt;

&lt;p&gt;One chain caps the range at 1,000 blocks. Another needs 500k-block chunks to keep the number of calls survivable. The only alternative provider for that second chain caps at 10k, which makes it useless as a backup even though it looks like a valid fallback. Chunk size has to be configured per chain, and a fallback endpoint is only a fallback if it can serve the same range.&lt;/p&gt;

&lt;h2&gt;
  
  
  viem details that cost me time
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;http()&lt;/code&gt; defaults to a 10 second timeout. That's long enough for a dead endpoint to stall a page load instead of failing over.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;fallback()&lt;/code&gt; always starts at the first endpoint in its list, so that endpoint absorbs every request and its rate limit becomes yours. If you want load spread, you need your own rotation.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;multicall&lt;/code&gt;'s &lt;code&gt;batchSize&lt;/code&gt; is bytes of calldata, not number of calls. I found that out the way you'd expect.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A rewind must only ever move the cursor backwards
&lt;/h2&gt;

&lt;p&gt;My worst bug. A keep-current pass set the cursor to &lt;code&gt;head - depth * 4&lt;/code&gt; outright. On a slow chain that's a rewind, which is what it was written to do. On a chain producing 600 blocks a minute, it's a 700-block jump forward that silently drops events.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;cursor = min(cursor, head - depth * 4)&lt;/code&gt; is the whole fix, and it should have been the whole implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stop reading history from RPC on page load
&lt;/h2&gt;

&lt;p&gt;In the end I stopped reading history from RPC on first load at all. Everything I needed was already in each paint transaction's calldata, so a cron job decodes it into Cloudflare KV, and the browser fetches one snapshot and only scans forward from there. A cold page load went from about 126 RPC calls to a handful.&lt;/p&gt;

&lt;p&gt;The constraint that shaped it is the KV free tier's 1,000 writes a day. The rule became never write on a schedule, only when something actually changed. A cron that writes every run will use up that quota before lunch.&lt;/p&gt;

&lt;h2&gt;
  
  
  It kept happening while I wrote this
&lt;/h2&gt;

&lt;p&gt;Two of the six chains rotted in the same week, which is why I finally wrote any of this down.&lt;/p&gt;

&lt;p&gt;On one chain, the last working endpoint quietly tightened its &lt;code&gt;getLogs&lt;/code&gt; range from 9,500 blocks to 5,000, below the chunk size I was asking for. Nothing visibly broke, because my paginator halves the chunk and retries, so it just wasted three failed calls before every successful one. Two days later the same endpoint stopped serving &lt;code&gt;getLogs&lt;/code&gt; altogether, with a 30 second server-side timeout, while &lt;code&gt;eth_call&lt;/code&gt; and &lt;code&gt;eth_blockNumber&lt;/code&gt; kept answering in under 50ms.&lt;/p&gt;

&lt;p&gt;Slow failure turned out to be worse than fast failure. A dead endpoint gets skipped in milliseconds. One that accepts your request and times out after 30 seconds stalls whichever page load was unlucky enough to get it. I pulled it from the pool, and that chain now has no free endpoint that will serve a query from the deploy block, so the snapshot isn't an optimisation there any more. It's the only way the history can be read at all.&lt;/p&gt;

&lt;p&gt;Meanwhile, on a second chain, both main public endpoints dropped their &lt;code&gt;getLogs&lt;/code&gt; range to 2,000 blocks, four days after passing the same probe at 9,500. Both are run by the same company. I had four endpoints listed for that chain and thought I had redundancy. Two were the same operator, one had been discontinued, and one had started answering &lt;code&gt;invalid params&lt;/code&gt; to every &lt;code&gt;getLogs&lt;/code&gt; call. Count operators, not URLs.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it's for
&lt;/h2&gt;

&lt;p&gt;The contract is a million-pixel canvas called tagwall, deployed on six chains, and the contract was genuinely the easy part. Every pixel is stored on chain, there's no owner or admin, and the frontend is open source. If you're curious, it's at &lt;a href="https://tagwall.io" rel="noopener noreferrer"&gt;https://tagwall.io&lt;/a&gt;, and I'm happy to go deeper on any of the above in the comments, the RPC probing especially.&lt;/p&gt;

</description>
      <category>ethereum</category>
      <category>web3</category>
      <category>webdev</category>
      <category>cloudflare</category>
    </item>
  </channel>
</rss>
