<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Darshan Turakhia</title>
    <description>The latest articles on DEV Community by Darshan Turakhia (@darshan_turakhia).</description>
    <link>https://dev.to/darshan_turakhia</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3848023%2F326f61e3-2775-436f-9506-0b35ed05d998.png</url>
      <title>DEV Community: Darshan Turakhia</title>
      <link>https://dev.to/darshan_turakhia</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/darshan_turakhia"/>
    <language>en</language>
    <item>
      <title>How a Zero-Downtime Table Swap Silently Dropped Orders From Postgres Logical Replication for 91 Hours</title>
      <dc:creator>Darshan Turakhia</dc:creator>
      <pubDate>Tue, 06 Oct 2026 10:59:16 +0000</pubDate>
      <link>https://dev.to/darshan_turakhia/how-a-zero-downtime-table-swap-silently-dropped-orders-from-postgres-logical-replication-for-91-4767</link>
      <guid>https://dev.to/darshan_turakhia/how-a-zero-downtime-table-swap-silently-dropped-orders-from-postgres-logical-replication-for-91-4767</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;09:14 UTC, Thursday. #finance-eng: "Why has daily revenue been flat at $118,402 for four days straight? We ran a 20%-off promo Tuesday, this should have spiked, not flatlined." No PagerDuty alert fired. No panel on the production dashboard is red. Checkout is taking orders normally. Whatever broke didn't trip a single automated check, because nothing in production is actually broken.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  the scramble
&lt;/h2&gt;

&lt;p&gt;First theory: the promo's attribution tracking is broken, not revenue itself. Marketing's UTM tagging has flaked before. Someone pulls the Stripe payouts dashboard directly, bypassing the internal reporting pipeline entirely. Payouts for the last four days are up, in line with the promo. Money is moving. The number everyone is staring at is wrong, not the business.&lt;/p&gt;

&lt;p&gt;Second theory: the BI tool is serving a stale cached query. The finance dashboard sits on top of an internal analytics Postgres instance, fed by logical replication from the primary rather than queried against production directly, a deliberate choice made a year ago so heavy reporting queries can never compete with checkout for connections or I/O. Someone forces a hard refresh on the dashboard, bypassing the BI tool's own cache layer. Same flat number.&lt;/p&gt;

&lt;p&gt;Third theory, and the one that finally moves: query the analytics replica directly instead of through the dashboard.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

09:41 UTC, analytics replica

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;SELECT count(*), max(created_at) FROM orders;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;count&lt;/th&gt;
&lt;th&gt;max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;9,214,880&lt;/td&gt;
&lt;td&gt;2026-10-02 14:02:41+00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;-- same query against production, same moment&lt;br&gt;
SELECT count(*), max(created_at) FROM orders;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;count&lt;/th&gt;
&lt;th&gt;max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;9,222,563&lt;/td&gt;
&lt;td&gt;2026-10-06 09:40:12+00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Production has 7,683 more rows than the replica, and the replica's newest row is timestamped almost exactly four days in the past. That timestamp isn't a random stale point. It's 12 minutes after the rename-swap migration that shipped Tuesday afternoon. The dashboard was never broken. The data feeding it stopped arriving.&lt;/p&gt;


&lt;h2&gt;
  
  
  the hunt
&lt;/h2&gt;

&lt;p&gt;Tuesday's deploy converted the orders table's primary key from &lt;code&gt;int4&lt;/code&gt; to &lt;code&gt;int8&lt;/code&gt; using the standard zero-downtime shadow-table pattern: build &lt;code&gt;orders_new&lt;/code&gt; with the wider key, backfill it in batches, dual-write via trigger while the backfill catches up, then an atomic rename swap inside one transaction:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

Tuesday 14:02:17 UTC, cutover

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;BEGIN;&lt;br&gt;
ALTER TABLE orders RENAME TO orders_old;&lt;br&gt;
ALTER TABLE orders_new RENAME TO orders;&lt;br&gt;
COMMIT;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Clean, fast, no lock held longer than the two renames. The on-call engineer who ran it watched checkout error rates for twenty minutes afterward and saw nothing. By every metric anyone was watching, the migration worked.&lt;/p&gt;

&lt;p&gt;First check on the replication path: is the slot even still alive?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

09:52 UTC, primary

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;SELECT slot_name, active, confirmed_flush_lsn&lt;br&gt;
FROM pg_replication_slots&lt;br&gt;
WHERE slot_name = 'analytics_sub';&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;slot_name&lt;/th&gt;
&lt;th&gt;active&lt;/th&gt;
&lt;th&gt;confirmed_flush_lsn&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;analytics_sub&lt;/td&gt;
&lt;td&gt;t&lt;/td&gt;
&lt;td&gt;4A2/8F1C3D00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;-- run again two minutes later&lt;br&gt;
 analytics_sub | t      | 4A2/8F22A188&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Active, and the LSN is climbing between queries. That reads as healthy replication by every normal heuristic, and it's the dead end that costs the most time: the publication backing this subscription also covers &lt;code&gt;customers&lt;/code&gt;, &lt;code&gt;invoices&lt;/code&gt;, and three other tables that are still being written to constantly. Those writes keep the logical decoding worker busy and the confirmed LSN advancing regardless of what's happening with orders specifically. A slot-level health check can look perfectly fine while one table inside it has gone completely silent.&lt;/p&gt;

&lt;p&gt;The next check is narrower, not "is the slot alive" but "what tables is this publication actually sending":&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

10:08 UTC, primary

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;SELECT pubname, tablename FROM pg_publication_tables&lt;br&gt;
WHERE pubname = 'orders_pub';&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;pubname&lt;/th&gt;
&lt;th&gt;tablename&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;orders_pub&lt;/td&gt;
&lt;td&gt;orders_old&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;orders_pub&lt;/td&gt;
&lt;td&gt;customers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;orders_pub&lt;/td&gt;
&lt;td&gt;invoices&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;code&gt;orders_old&lt;/code&gt;. Not &lt;code&gt;orders&lt;/code&gt;. The publication has been faithfully, correctly, silently replicating a table that nothing has written to since 14:02:17 UTC on Tuesday.&lt;/p&gt;


&lt;h2&gt;
  
  
  the find
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;ALTER PUBLICATION ... ADD TABLE&lt;/code&gt; resolves the table name to its object ID at the moment the command runs and stores that OID in &lt;code&gt;pg_publication_rel&lt;/code&gt;. It does not store the name. From that point forward, the publication tracks whatever object owns that OID, not whatever object currently answers to the name &lt;code&gt;orders&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The rename-swap pattern is specifically designed to be atomic and cheap, which it achieves by never actually copying data at cutover time. It just reassigns two names to two existing relations. The relation that was &lt;code&gt;orders&lt;/code&gt; for the three years before Tuesday did not stop existing. It got renamed to &lt;code&gt;orders_old&lt;/code&gt; and kept its OID, and the publication kept following that OID exactly as configured. The relation that's been named &lt;code&gt;orders&lt;/code&gt; since 14:02:17 UTC is a different object with a different OID, one the publication was never told about, because the migration added a new table under a temporary name and the publication definition was never revisited after the swap.&lt;/p&gt;

&lt;p&gt;Nothing in this chain is a bug. The publication did precisely what &lt;code&gt;ALTER PUBLICATION ADD TABLE orders&lt;/code&gt; asked it to do, a year ago, for the table that was &lt;code&gt;orders&lt;/code&gt; at the time. A rename is not a schema change the replication system has any reason to treat as suspicious, since tables get renamed for all kinds of reasons that have nothing to do with swapping in a different relation underneath a stable name. There was no constraint violation, no decoding error, no WAL gap, nothing for any alert to fire on. The failure mode is a correctly configured system doing exactly what it was told, applied to an assumption, "the table named orders is the table I meant," that a swap-based migration quietly invalidated.&lt;/p&gt;


&lt;h2&gt;
  
  
  the fix
&lt;/h2&gt;

&lt;p&gt;Point the publication at the table that's actually live, and drop the dead one once its replicated history is no longer needed on the subscriber:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

10:31 UTC, primary

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;ALTER PUBLICATION orders_pub DROP TABLE orders_old;&lt;br&gt;
ALTER PUBLICATION orders_pub ADD TABLE orders;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Adding a table to an existing publication doesn't push any of its current rows to subscribers by itself. It only starts streaming changes from that point forward. The subscriber still needs the 91 hours of orders it missed, plus confirmation that its local &lt;code&gt;orders&lt;/code&gt; table matches the live one row for row, not just going forward from here:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

10:33 UTC, analytics replica

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;ALTER SUBSCRIPTION analytics_sub REFRESH PUBLICATION WITH (copy_data = true);&lt;br&gt;
-- 11 minutes to copy and index 9.2M rows&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;REFRESH PUBLICATION&lt;/code&gt; diffs the subscription's known relations against the publication's current membership, and for anything newly added it runs a full initial copy for that relation only, while &lt;code&gt;customers&lt;/code&gt; and &lt;code&gt;invoices&lt;/code&gt; keep streaming without interruption the entire time. Eleven minutes later the replica's &lt;code&gt;orders&lt;/code&gt; table matches production exactly, and the finance dashboard's next scheduled refresh picks up four days of orders in one jump.&lt;/p&gt;

&lt;p&gt;The longer-term fix is a check that doesn't depend on anyone remembering to update the publication by hand during a migration that was never framed as a replication change in the first place:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

nightly parity check, new cron

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;-- runs against every table/subscriber pair, alerts on the table name, not the slot&lt;br&gt;
SELECT p.tablename,&lt;br&gt;
       now() - max(o.created_at) AS staleness&lt;br&gt;
FROM pg_publication_tables p&lt;br&gt;
JOIN orders o ON true&lt;br&gt;
WHERE p.pubname = 'orders_pub' AND p.tablename = 'orders'&lt;br&gt;
GROUP BY p.tablename&lt;br&gt;
HAVING now() - max(o.created_at) &amp;gt; interval '2 hours';&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It also became a standing step in the shadow-table migration runbook: any rename-swap that touches a table covered by logical replication requires a publication membership check in the same change window as the cutover, not as a follow-up ticket filed whenever someone happens to notice.&lt;/p&gt;




&lt;h2&gt;
  
  
  the aftermath
&lt;/h2&gt;

&lt;p&gt;91 hrs Between cutover and the orders table rejoining replication&lt;/p&gt;

&lt;p&gt;7,683 Orders missing from the analytics replica at discovery&lt;/p&gt;

&lt;p&gt;$0 Actual revenue lost — Stripe and production had every order&lt;/p&gt;

&lt;p&gt;11 min To resync the table via a targeted copy_data refresh&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Postgres publications bind to a table's object ID at &lt;code&gt;ADD TABLE&lt;/code&gt; time, not its name. A rename-swap migration that reassigns an existing name to a new relation is invisible to that binding — the publication keeps following the OID, which is now the wrong table.&lt;/li&gt;
&lt;li&gt;  A healthy replication slot (active, LSN advancing) says the slot is working. It says nothing about whether a specific table inside a multi-table publication is still part of that stream. Check &lt;code&gt;pg_publication_tables&lt;/code&gt; directly when one table's data looks stale; don't infer table-level health from slot-level metrics.&lt;/li&gt;
&lt;li&gt;  &lt;code&gt;ALTER SUBSCRIPTION ... REFRESH PUBLICATION WITH (copy_data = true)&lt;/code&gt; resyncs only the relations that are new to the subscription. It's a safe, targeted recovery step, not a full subscription rebuild, and it doesn't interrupt tables that were never affected.&lt;/li&gt;
&lt;li&gt;  Any migration pattern built around renaming tables, not just this one, should carry an explicit step to re-check every publication and every foreign key, trigger, or grant that was bound to the old name by object reference rather than by name. A clean rename at the schema level can still leave dependent systems pointed at the wrong object.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The migration itself worked exactly as designed, start to finish, with zero downtime and zero checkout errors. The thing it broke wasn't checkout. It was a dependency three systems away that nobody thought to list as a stakeholder for a primary-key type change, and the only signal it ever produced was a dashboard that quietly stopped changing.&lt;/p&gt;

</description>
      <category>database</category>
      <category>postgres</category>
      <category>backend</category>
      <category>programming</category>
    </item>
    <item>
      <title>Reimplementing a Trie Reminded Me How Autocomplete Actually Works</title>
      <dc:creator>Darshan Turakhia</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:06:26 +0000</pubDate>
      <link>https://dev.to/darshan_turakhia/reimplementing-a-trie-reminded-me-how-autocomplete-actually-works-5ghd</link>
      <guid>https://dev.to/darshan_turakhia/reimplementing-a-trie-reminded-me-how-autocomplete-actually-works-5ghd</guid>
      <description>&lt;p&gt;I spent an embarrassing amount of time last week reimplementing something I'd "known" since a data structures class a decade ago: the trie. Typing it out from scratch again reminded me why it's one of the few data structures that actually earns the "elegant" label people throw around too easily.&lt;/p&gt;

&lt;p&gt;Here's the pitch in one sentence: a trie stores strings by character, along shared paths from a root, so two words with the same prefix walk the exact same nodes until they stop agreeing. Insert "car," "card," and "care," and you get one shared path for "c-a-r," which then splits into three different endings. Nothing about adding "care" touches the nodes "car" and "card" already built.&lt;/p&gt;

&lt;p&gt;That sharing is the whole point, and it's easy to undersell how much it buys you. A hash set can tell you "yes, this exact string is in the collection." That's it. It cannot tell you "give me everything that starts with these three letters" without scanning the entire set and checking each entry one by one. A trie answers that second question by walking to the end of the prefix and then just... looking at what's underneath. Lookup cost depends on how long your prefix is, not on how many million other strings happen to be sitting next to it in memory. That's the entire mechanism behind every autocomplete dropdown you've ever typed into: walk the prefix, then collect every real word hanging off wherever you land.&lt;/p&gt;

&lt;p&gt;The part that actually bit me while rebuilding this: you need an explicit marker for "a real word ends here." It sounds obvious written down, but it's the single easiest thing to forget, and the bug it causes is sneaky. Say you only ever insert "card." The path c → a → r → d exists in your trie. All four nodes are real. But "car" was never inserted as its own entry, it's just a prefix that happens to lead somewhere. If your node doesn't carry a boolean flag for "an insertion actually terminated here," you have no way to distinguish a real stored word from a string that merely happens to be a prefix of a longer one. I've seen (and once written) autocomplete bugs where a partial prefix gets suggested as a complete match because of exactly this missing flag.&lt;/p&gt;

&lt;p&gt;The tradeoff nobody mentions in the five-minute version of this concept: a naive implementation wastes a lot of memory. The textbook approach gives every node a fixed-size array, one slot per possible next character. Twenty-six slots if you're doing lowercase English, more if you need digits or punctuation or, god forbid, full Unicode. Most of those slots sit empty on any node that doesn't branch in every direction, and that waste adds up fast across a trie with real-world vocabulary in it. The fix is one of two things: back each node with a hash map instead of a fixed array, so you only pay for children that actually exist, or compress the whole structure with something like a radix trie, which collapses long runs of single-child nodes into one node holding a whole substring instead of one character. "Cardboard" doesn't need five separate one-character hops after "card," it needs one node labeled "board."&lt;/p&gt;

&lt;p&gt;Autocomplete is the use case everyone reaches for first, and it's a good one, but it's not the only place this shows up. Spell-checkers use the same walk to confirm a word exists in a dictionary and to find near-miss suggestions by exploring nearby paths. IP routers use a binary version of the same idea, walking a trie over address bits instead of letters, to do longest-prefix-match routing: find the most specific rule that matches a destination address. Different alphabet, same shared-prefix-sharing-a-path trick underneath.&lt;/p&gt;

&lt;p&gt;I ended up writing both a longer explanation and a small interactive visualizer, mostly because I wanted to actually watch the branching happen rather than just trust my own mental model of it. If you want to poke at it yourself, &lt;a href="https://nodique.com/guides/what-is-a-trie" rel="noopener noreferrer"&gt;the guide&lt;/a&gt; walks through the insert/search mechanics and the end-of-word gotcha in more detail, and &lt;a href="https://nodique.com/tools/trie-visualizer" rel="noopener noreferrer"&gt;the visualizer&lt;/a&gt; lets you insert your own words and watch the prefix lookup highlight in real time, suggestions and all.&lt;/p&gt;

&lt;p&gt;If you've never implemented one, it's a genuinely good weekend exercise. Small enough to build in an hour. Annoying enough, in the best way, that you'll start second-guessing every search box you touch afterward.&lt;/p&gt;

</description>
      <category>programming</category>
      <category>computerscience</category>
      <category>webdev</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How a Rolling Deploy</title>
      <dc:creator>Darshan Turakhia</dc:creator>
      <pubDate>Tue, 29 Sep 2026 10:58:00 +0000</pubDate>
      <link>https://dev.to/darshan_turakhia/how-a-rolling-deploy-cm6</link>
      <guid>https://dev.to/darshan_turakhia/how-a-rolling-deploy-cm6</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;13:52 UTC. p99 latency on &lt;code&gt;pricing-svc&lt;/code&gt; alerts at 1.8 seconds, four times its usual ceiling. The dashboard shows six healthy pods, six pods passing their readiness probe, six pods with room on their CPU requests. One of them is pegged at 97%. The other five are idling under 12%.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  the setup
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;checkout-api&lt;/code&gt; calls &lt;code&gt;pricing-svc&lt;/code&gt; over gRPC for every cart total, roughly 4,000 RPCs a minute at midday. Both run in the same namespace, and &lt;code&gt;pricing-svc&lt;/code&gt; is exposed through a standard ClusterIP Service, resolved once by DNS to a single virtual IP with kube-proxy handling the fan-out to whichever pods are in the endpoint list.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;

&lt;span class="nx"&gt;checkout&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;api&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="nx"&gt;pricing&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ts&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;const client = new PricingServiceClient(&lt;br&gt;
  'pricing-svc.default.svc.cluster.local:50051',&lt;br&gt;
  grpc.credentials.createInsecure()&lt;br&gt;
);&lt;br&gt;
// one channel per pod, reused for every call it makes&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line is the part nobody thought hard about. gRPC runs on HTTP/2, and HTTP/2 multiplexes many concurrent RPCs over a single TCP connection. A gRPC client opens one connection to a target and keeps it open, reusing it for every subsequent call instead of opening a new one per request. For a REST client that reconnects constantly, ClusterIP round robin evens out over thousands of short-lived connections. For a gRPC client, the load balancing decision happens exactly once, at connection time, and then holds for as long as that connection stays up. &lt;code&gt;checkout-api&lt;/code&gt; runs eight pods. Each one makes that decision independently, once, on startup.&lt;/p&gt;




&lt;h2&gt;
  
  
  the scramble
&lt;/h2&gt;

&lt;p&gt;First theory: bad node. The hot pod, &lt;code&gt;pricing-svc-7d4f-x2k9&lt;/code&gt;, is cordoned and drained, forcing it onto a different node entirely. Eleven minutes later the alert fires again, same symptom, different pod, &lt;code&gt;pricing-svc-7d4f-m8p1&lt;/code&gt;, now pegged while the rest idle. Different node, different hardware, same shape of failure. That rules out anything node-specific and burns fifteen minutes doing it.&lt;/p&gt;

&lt;p&gt;Second theory: the HPA is fighting itself, scaling on a metric that doesn't reflect real load. &lt;code&gt;kubectl describe hpa pricing-svc&lt;/code&gt; shows six of six replicas, target CPU utilization at 60%, current average across the deployment sitting at 22%. The average looks fine because five pods are nearly idle. Averages hide a skew this sharp, and the HPA has no reason to add capacity when its own number says the fleet is underloaded.&lt;/p&gt;




&lt;h2&gt;
  
  
  the hunt
&lt;/h2&gt;

&lt;p&gt;Instead of watching pods cycle, the next step is watching the connections themselves. Shelling into a &lt;code&gt;checkout-api&lt;/code&gt; pod and checking its open sockets to the pricing service:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;

kubectl &lt;span class="nb"&gt;exec &lt;/span&gt;checkout-api-9f3a-vv2q &lt;span class="nt"&gt;--&lt;/span&gt; ss &lt;span class="nt"&gt;-tnp&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; :50051

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;ESTAB  0  0  10.2.4.19:44822  10.2.6.31:50051&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One line. One connection. Every one of the eight &lt;code&gt;checkout-api&lt;/code&gt; pods shows the exact same thing: a single established socket to a single backend IP, held open since the pod started. Cross-referencing those destination IPs against the pod list settles it: all eight client pods are connected to the same two pricing-svc pods, and nothing is connected to the other four at all.&lt;/p&gt;

&lt;p&gt;A Prometheus query against &lt;code&gt;grpc_server_handled_total&lt;/code&gt;, split by pod, shows just how lopsided it got: those two pods have handled 83% of all RPCs since the timestamp on a &lt;code&gt;checkout-api&lt;/code&gt; deploy from six hours earlier, a routine dependency bump with nothing in the diff touching networking. The deploy itself didn't cause an error. It just decided, once, where six hours of subsequent traffic would go.&lt;/p&gt;




&lt;h2&gt;
  
  
  the find
&lt;/h2&gt;

&lt;p&gt;The rollout replaced all eight &lt;code&gt;checkout-api&lt;/code&gt; pods inside a forty-second window, standard behavior for a deployment with no &lt;code&gt;maxUnavailable&lt;/code&gt; tuning. At the moment those eight new pods came up and each opened its one gRPC connection, &lt;code&gt;pricing-svc&lt;/code&gt; was mid-rollout too, from an unrelated change earlier that morning. Only two of its six pods had passed the readiness probe and made it into the Service's endpoint list. kube-proxy's iptables rules can only route to what's registered, so every new connection from every &lt;code&gt;checkout-api&lt;/code&gt; pod landed on one of those two.&lt;/p&gt;

&lt;p&gt;By the time the other four &lt;code&gt;pricing-svc&lt;/code&gt; pods turned ready, thirty seconds later, it didn't matter. The gRPC client's default load balancing policy is &lt;code&gt;pick_first&lt;/code&gt;: resolve the target once, connect to the first address, and stay connected until that connection breaks. It doesn't re-resolve, doesn't rebalance, doesn't notice four more pods just joined the pool. The connection from that forty-second window was healthy, so it just kept being used, for six hours, until midday traffic pushed those two pods past the point where 1.8-second p99 was survivable.&lt;/p&gt;




&lt;h2&gt;
  
  
  the fix
&lt;/h2&gt;

&lt;p&gt;Immediate: restart the two hot &lt;code&gt;pricing-svc&lt;/code&gt; pods. Their connections drop, the eight &lt;code&gt;checkout-api&lt;/code&gt; pods reconnect, and this time all six backends are already ready, so the new connections spread out. Latency back to baseline in under two minutes. That's a mitigation, not a fix, since the exact same readiness race can happen on the next deploy.&lt;/p&gt;

&lt;p&gt;The real fix is making the client aware there's more than one backend to pick from. That means a headless Service, so DNS returns every pod IP instead of one virtual IP, paired with a client-side load balancing policy that actually uses all of them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;

&lt;span class="s"&gt;k8s/pricing-svc-headless.yaml&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;apiVersion: v1&lt;br&gt;
kind: Service&lt;br&gt;
metadata:&lt;br&gt;
  name: pricing-svc-headless&lt;br&gt;
spec:&lt;br&gt;
  clusterIP: None&lt;br&gt;
  selector:&lt;br&gt;
    app: pricing-svc&lt;br&gt;
  ports:&lt;br&gt;
    - port: 50051&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;

&lt;span class="nx"&gt;checkout&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;api&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="nx"&gt;pricing&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;after&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;const client = new PricingServiceClient(&lt;br&gt;
  'dns:///pricing-svc-headless.default.svc.cluster.local:50051',&lt;br&gt;
  grpc.credentials.createInsecure(),&lt;br&gt;
  { 'grpc.lb_policy': 'round_robin' }&lt;br&gt;
);&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With &lt;code&gt;round_robin&lt;/code&gt;, the client resolves every backend IP and opens a connection to each one, spreading calls across all of them instead of pinning to whichever pod answered first. As a second layer, &lt;code&gt;pricing-svc&lt;/code&gt; now caps how long any single connection can live, so even a client stuck on an old load balancing policy is forced to reconnect and re-resolve periodically instead of holding one connection indefinitely:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;

&lt;span class="nx"&gt;pricing&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;svc&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="nx"&gt;server&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ts&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;const server = new grpc.Server({&lt;br&gt;
  'grpc.max_connection_age_ms': 5 * 60 * 1000,&lt;br&gt;
  'grpc.max_connection_age_grace_ms': 30 * 1000,&lt;br&gt;
});&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  the aftermath
&lt;/h2&gt;

&lt;p&gt;83% Of all pricing-svc RPCs handled by 2 of 6 pods for six hours&lt;/p&gt;

&lt;p&gt;1.8s p99 latency at the point the imbalance became visible&lt;/p&gt;

&lt;p&gt;5 min max_connection_age_ms added as a floor on any future skew&lt;/p&gt;

&lt;p&gt;3 Other internal gRPC clients found using ClusterIP + pick_first&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  A ClusterIP Service load balances TCP connections, not requests. That distinction only matters for protocols that hold a connection open and reuse it, which is exactly what HTTP/2 and gRPC are built to do.&lt;/li&gt;
&lt;li&gt;  &lt;code&gt;pick_first&lt;/code&gt; is a reasonable default for a client with one backend. Against a Service backed by multiple pods, it turns whatever pod answered first during a narrow readiness window into the only pod that matters, indefinitely.&lt;/li&gt;
&lt;li&gt;  Average CPU across a deployment can look completely healthy while two pods carry the entire fleet. Per-pod metrics on anything that does client-side connection reuse are not optional, they're the only view that shows this failure mode at all.&lt;/li&gt;
&lt;li&gt;  A connection age limit on the server is a cheap insurance policy even after fixing the client. It bounds how long any future imbalance, caused by a bug nobody's found yet, can last before the system corrects itself on its own.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nothing about that rollout failed. Both deployments finished, both health checks passed, and kube-proxy did exactly what a ClusterIP Service is supposed to do. The problem was a forty-second window where "correct" and "evenly distributed" briefly meant two different things, and gRPC's connection reuse turned that window into six hours.&lt;/p&gt;

</description>
      <category>programming</category>
      <category>backend</category>
      <category>devops</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How Kafka</title>
      <dc:creator>Darshan Turakhia</dc:creator>
      <pubDate>Sat, 26 Sep 2026 10:56:53 +0000</pubDate>
      <link>https://dev.to/darshan_turakhia/how-kafka-2obk</link>
      <guid>https://dev.to/darshan_turakhia/how-kafka-2obk</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;03:14 UTC. Security's on-call channel gets a message from the fraud-scanning job: a customer API key flagged and revoked four days earlier is still authorizing requests, all of them from ap-south-1. The revoke ticket has been closed as resolved since Monday.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  the setup
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;authz-gateway&lt;/code&gt; sits in front of every API route and checks incoming keys against a local cache before anything touches the database. It used to call Postgres on every request; that was a flat 40ms added to p99 regardless of what the request actually needed. Six months earlier, someone replaced it with a Kafka Streams state store hydrated from a compacted topic, &lt;code&gt;api-key-status&lt;/code&gt;, keyed by &lt;code&gt;api_key_id&lt;/code&gt;. p99 for the auth check dropped to 4ms. An active key is a normal record. A revoked key is a tombstone, a record with the key present and the value set to &lt;code&gt;null&lt;/code&gt;, which is how compacted topics represent deletion.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;

&lt;span class="nx"&gt;gateway&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="nx"&gt;authz&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;js&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;function isKeyValid(apiKeyId) {&lt;br&gt;
  const record = keyStatusStore.get(apiKeyId); // local RocksDB state store&lt;br&gt;
  return record !== undefined &amp;amp;&amp;amp; record.status === 'active';&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The consumer group backing that store runs with &lt;code&gt;group.instance.id&lt;/code&gt; set, static membership, turned on months earlier after a rolling deploy triggered a full rebalance and briefly spiked auth latency across every pod, not just the ones restarting. Static membership means a pod that drops out during a deploy keeps its partition assignment reserved instead of forcing an immediate reassignment. It was a good fix for that problem. It also meant nobody was watching for what happens when a pod goes quiet for reasons that have nothing to do with a deploy.&lt;/p&gt;




&lt;h2&gt;
  
  
  the scramble
&lt;/h2&gt;

&lt;p&gt;First move: check the source of truth. Postgres's &lt;code&gt;api_keys&lt;/code&gt; table shows the key correctly marked &lt;code&gt;revoked_at&lt;/code&gt; four days ago. Whatever's wrong isn't the revoke itself, it happened, it's recorded, it's just not everywhere it needs to be.&lt;/p&gt;

&lt;p&gt;Second theory: the tombstone never made it onto the topic. A console consumer against &lt;code&gt;api-key-status&lt;/code&gt; from the revoke's approximate timestamp finds it immediately, a null-value record for that exact key, committed and replicated. The event exists. The dead end didn't cost much, five minutes to rule out, but it mattered: it meant the problem was downstream of Kafka, not upstream of it.&lt;/p&gt;

&lt;p&gt;Third theory, on-call's: a stale edge cache. This service doesn't have one on the authz path, ruled out by reading its own deployment topology rather than by testing anything.&lt;/p&gt;




&lt;h2&gt;
  
  
  the hunt
&lt;/h2&gt;

&lt;p&gt;The fraud job's flagged requests all traced to three pods, all in ap-south-1, all part of the same consumer group. &lt;code&gt;kafka-consumer-groups.sh --describe&lt;/code&gt; against that group showed those three partitions current, no lag, caught up to the log's high-water mark. Not stuck. Not behind. Just wrong.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;

kafka-consumer-groups.sh &lt;span class="nt"&gt;--describe&lt;/span&gt; &lt;span class="nt"&gt;--group&lt;/span&gt; authz-gateway

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;PARTITION  CURRENT-OFFSET  LOG-END-OFFSET  LAG  CONSUMER-ID&lt;br&gt;
4          88213           88213           0    authz-gateway-7f2a...&lt;br&gt;
7          61940           61940           0    authz-gateway-7f2a...&lt;br&gt;
9          73301           73301           0    authz-gateway-7f2a...&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Caught up with a wrong answer is a different bug than stuck behind. The next thing to check wasn't the topic, it was those three pods' history. Kubernetes events showed the answer: an ap-south-1-scoped NetworkPolicy change went out five days earlier as part of a node pool migration, and a rule ordering mistake in it blocked egress from those three pods to the Kafka broker subnet for roughly 30 hours before someone caught it in an unrelated ticket and reverted it. Nothing restarted those pods. Their liveness probe was an HTTP ping against the gateway's own port, which had nothing to do with whether it could reach Kafka, so it stayed green the entire time.&lt;/p&gt;

&lt;p&gt;With no broker access, the consumer's background heartbeat thread stopped reaching the group coordinator. Kafka correctly detected that after &lt;code&gt;session.timeout.ms&lt;/code&gt; and reassigned those three partitions to healthy members elsewhere, who kept the cache correct for everyone else. The revoke tombstone was produced and consumed by those healthy members within seconds, exactly as designed. It's what happened to the three isolated pods, once the network policy reverted and they reconnected, that broke.&lt;/p&gt;




&lt;h2&gt;
  
  
  the find
&lt;/h2&gt;

&lt;p&gt;Kafka Streams persists its state store to local disk and checkpoints the offset it last processed. On reconnect, if that checkpoint is still within the log, it doesn't replay from scratch, it does a delta restore: read forward from the checkpointed offset, apply whatever's new. That's the whole point of a local store, cheap recovery. The three pods' checkpoints were roughly 6 hours stale when the network came back, right around when the revoke had been written.&lt;/p&gt;

&lt;p&gt;The topic's &lt;code&gt;delete.retention.ms&lt;/code&gt; was left at Kafka's default, 86400000ms, 24 hours. That setting exists specifically so a lagging consumer gets a grace window to see a tombstone before the log cleaner physically removes it. The three pods were unreachable for roughly 30 hours, past that window. By the time they reconnected and asked to resume from their checkpoint, the tombstone that checkpoint needed to see had already been compacted away. The delta restore read forward past where the tombstone used to be and found nothing for that key, so it never touched the in-memory record still sitting there from before the isolation: active.&lt;/p&gt;

&lt;p&gt;Nothing in that path throws an error. The consumer isn't lagging, the topic isn't corrupt, the restore completes successfully. It just completes with a state store that is now permanently wrong for one key, with no future event on the topic ever going to fix it, because as far as the topic is concerned that key's history now starts after the revoke.&lt;/p&gt;




&lt;h2&gt;
  
  
  the fix
&lt;/h2&gt;

&lt;p&gt;Immediate: force the three pods to drop their local state directory and rebuild from the earliest offset rather than delta-restoring from a checkpoint that could no longer be trusted.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

recovery, run once per affected pod

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;kubectl exec authz-gateway-7f2a-xyz -- rm -rf /data/kafka-streams/authz-gateway&lt;br&gt;
kubectl delete pod authz-gateway-7f2a-xyz&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the two changes meant to keep this from being invisible again. First, a health check that actually reflects Kafka connectivity instead of just the local HTTP server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;

&lt;span class="nx"&gt;gateway&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="nx"&gt;health&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;js&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;after&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;let lastPollAt = Date.now();&lt;/p&gt;

&lt;p&gt;consumer.on('poll', () =&amp;gt; {&lt;br&gt;
  lastPollAt = Date.now();&lt;br&gt;
});&lt;/p&gt;

&lt;p&gt;app.get('/healthz', (req, res) =&amp;gt; {&lt;br&gt;
  const staleMs = Date.now() - lastPollAt;&lt;br&gt;
  if (staleMs &amp;gt; 60_000) {&lt;br&gt;
    return res.status(503).json({ error: &lt;code&gt;kafka poll stale for \${staleMs}ms\&lt;/code&gt; });&lt;br&gt;
  }&lt;br&gt;
  res.status(200).json({ ok: true });&lt;br&gt;
});&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Second, &lt;code&gt;delete.retention.ms&lt;/code&gt; on &lt;code&gt;api-key-status&lt;/code&gt; went from 24 hours to 7 days, trading disk for a grace window wide enough to cover a multi-day network isolation, not just a routine restart. It doesn't fix the underlying risk, a consumer down longer than the retention window can still miss a tombstone, but it moves the failure mode from "a botched network policy" to "something has been broken for a week and nobody noticed," which is a bar this team was willing to accept.&lt;/p&gt;




&lt;h2&gt;
  
  
  the aftermath
&lt;/h2&gt;

&lt;p&gt;96 hrs Time between the revoke and the fraud job catching it in ap-south-1&lt;/p&gt;

&lt;p&gt;3 Requests authorized with the revoked key before detection&lt;/p&gt;

&lt;p&gt;24h → 7d delete.retention.ms on api-key-status&lt;/p&gt;

&lt;p&gt;2 Other Kafka-backed services found with the same HTTP-only healthcheck&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Static group membership solves rebalance storms during deploys, but it also means a consumer that's gone quiet for an unrelated reason keeps its assignment and its local state without raising anything from outside. It needs its own signal, not the assumption that a healthy pod implies a healthy consumer.&lt;/li&gt;
&lt;li&gt;  &lt;code&gt;delete.retention.ms&lt;/code&gt; is a grace window, not a guarantee. It's sized against how long a consumer might lag, not against how long one might be completely unreachable, and those are different failure modes with very different durations.&lt;/li&gt;
&lt;li&gt;  A local state store that restores by delta from a checkpoint is fast exactly because it trusts the checkpoint. That trust is only as good as the log's willingness to still contain everything between the checkpoint and now, and compaction doesn't make that promise indefinitely.&lt;/li&gt;
&lt;li&gt;  A healthcheck that only proves a process is running, not that it's connected to the one dependency that determines whether its answers are correct, will report green through the exact failure it exists to catch.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Kafka never logged an error either. It deleted exactly what its retention setting told it to delete, on schedule, while the only three pods that needed to see it first were on the wrong side of a firewall rule.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>programming</category>
      <category>backend</category>
      <category>devops</category>
    </item>
    <item>
      <title>How Skipping a Cache Reset for Speed Let Apollo Client Show the Wrong Tenant</title>
      <dc:creator>Darshan Turakhia</dc:creator>
      <pubDate>Wed, 23 Sep 2026 10:55:43 +0000</pubDate>
      <link>https://dev.to/darshan_turakhia/how-skipping-a-cache-reset-for-speed-let-apollo-client-show-the-wrong-tenant-2p18</link>
      <guid>https://dev.to/darshan_turakhia/how-skipping-a-cache-reset-for-speed-let-apollo-client-show-the-wrong-tenant-2p18</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;09:41 PST. A support lead pastes a screenshot into the incident channel: her own screen, mid quick-switch into Acme Corp's account, showing three ticket subjects that mention Globex Inc, a different customer entirely. The message underneath is two words: "is this normal."&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  the setup
&lt;/h2&gt;

&lt;p&gt;The internal tool support agents used to manage tickets across customer accounts was a multi-tenant React app on top of a GraphQL API, Apollo Client for data fetching. Tenant scoping happened entirely server-side: every request carried a JWT with a &lt;code&gt;tenantId&lt;/code&gt; claim, and every resolver filtered its database queries by that claim. The GraphQL queries themselves never mentioned a tenant. A ticket list looked like this, the same query text for every account an agent ever loaded:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight graphql"&gt;&lt;code&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;TicketsOverview.graphql&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;query TicketsOverview {&lt;br&gt;
  tickets {&lt;br&gt;
    id&lt;br&gt;
    subject&lt;br&gt;
    status&lt;br&gt;
    customer&lt;br&gt;
  }&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That was a deliberate simplification. Plumbing a tenant argument through every query in the app would have meant threading it through dozens of components for no functional gain, since the server already refused to return data outside the caller's token. Nobody on the frontend team had reason to think of tenant identity as something the client needed to track for itself.&lt;/p&gt;

&lt;p&gt;Six weeks earlier, product shipped Quick Switch: a dropdown that let an agent jump from one customer account to another without a full sign-out and page reload. The old flow called &lt;code&gt;client.resetStore()&lt;/code&gt; on every account change, which cleared the Apollo cache and refetched every active query, a UX that took 300 to 500ms and showed a loading spinner over the whole page. Quick Switch skipped that on purpose, to make account switching feel instant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;

&lt;span class="nx"&gt;useQuickSwitch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;before&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;async function switchTenant(tenantId: string) {&lt;br&gt;
  const { accessToken } = await api.post('/auth/switch-tenant', { tenantId });&lt;br&gt;
  setAccessToken(accessToken);&lt;br&gt;
  // resetStore() intentionally skipped here — it added a visible&lt;br&gt;
  // loading flash and product wanted the switch to feel instant.&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every query in the app used Apollo's default &lt;code&gt;cache-first&lt;/code&gt; fetch policy. That was fine for a single-tenant session: read from cache if present, otherwise hit the network. Nobody had reasoned through what &lt;code&gt;cache-first&lt;/code&gt; does the instant the same query, with the same variables, needs to mean something different because the identity behind it changed.&lt;/p&gt;




&lt;h2&gt;
  
  
  the scramble
&lt;/h2&gt;

&lt;p&gt;First theory: a session bug on the auth service, a Redis-backed token cache serving a stale &lt;code&gt;tenantId&lt;/code&gt; claim to the wrong pod. The on-call engineer pulled the auth service's logs for the support lead's account switch. The issued JWT was correct, Acme Corp's tenant ID, timestamped to the millisecond of the click. Ruled out within ten minutes.&lt;/p&gt;

&lt;p&gt;Second theory: the GraphQL gateway was routing to a stale replica that hadn't caught up on a recent write, some kind of read-after-write lag. Someone checked the actual network response in the support lead's browser recording, the one piece of evidence that made this theory look promising at first. The GraphQL response body for &lt;code&gt;TicketsOverview&lt;/code&gt;, once it arrived, contained exactly Acme Corp's tickets. Correct tenant, correct data, right there in the payload.&lt;/p&gt;

&lt;p&gt;That was the detail that stalled the investigation for the better part of an hour: every request and every response, read individually, was correct. The bug wasn't in anything that crossed the network.&lt;/p&gt;




&lt;h2&gt;
  
  
  the hunt
&lt;/h2&gt;

&lt;p&gt;The break came from reproducing it deliberately instead of waiting for another report. Someone throttled their local network to Slow 3G in dev tools, specifically to stretch out whatever window the bug lived in, then opened Apollo Client Devtools and watched the &lt;code&gt;TicketsOverview&lt;/code&gt; cache entry while triggering Quick Switch.&lt;/p&gt;

&lt;p&gt;The component re-rendered twice. The first render, milliseconds after the switch, showed the previous tenant's tickets, read straight from cache. The second render, 800ms to 1.5 seconds later depending on network conditions, showed the correct tenant's tickets, once the network response landed and overwrote the cache entry.&lt;/p&gt;

&lt;p&gt;Apollo Client's default cache key is the query name plus its serialized variables. &lt;code&gt;TicketsOverview&lt;/code&gt; had no variables. Every tenant, every session, every agent, all produced the identical cache key: &lt;code&gt;TicketsOverview:{}&lt;/code&gt;. The cache had no idea two different customers' data had ever passed through that key, because from the cache's perspective, nothing about the query had changed. Tenant identity lived exclusively in an HTTP header the cache never looked at.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

what cache-first actually did on switch

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;Agent clicks "Switch to Acme Corp"&lt;/li&gt;
&lt;li&gt;New JWT (tenantId=acme) stored, old JWT (tenantId=globex) discarded&lt;/li&gt;
&lt;li&gt;TicketsOverview component re-renders&lt;/li&gt;
&lt;li&gt;Apollo checks cache for key TicketsOverview:{} -&amp;gt; HIT (Globex's data, still cached)&lt;/li&gt;
&lt;li&gt;cache-first policy: return the hit immediately, refetch in background&lt;/li&gt;
&lt;li&gt;UI paints Globex's tickets under Acme Corp's header, for 0.8-1.5s&lt;/li&gt;
&lt;li&gt;Network response for Acme's data arrives, overwrites cache, UI re-renders correctly
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  the find
&lt;/h2&gt;

&lt;p&gt;Root cause: the Apollo Client cache had no concept of tenant identity, because tenant scoping had only ever been designed as a server-side concern. That was a correct design for the network layer, the server never returned the wrong tenant's data. It was an incomplete design for the client, because &lt;code&gt;cache-first&lt;/code&gt; answers queries from a cache that has no way to know the caller's identity changed. Quick Switch removed the one step, &lt;code&gt;resetStore()&lt;/code&gt;, that had been silently compensating for that gap by wiping the cache clean on every identity change. Once that step was gone, the gap was directly exposed to whoever happened to switch accounts fastest.&lt;/p&gt;




&lt;h2&gt;
  
  
  the fix
&lt;/h2&gt;

&lt;p&gt;The immediate fix restored a cache clear on switch, but a narrower one than the old full &lt;code&gt;resetStore()&lt;/code&gt;, which also triggers a refetch of every active query on the page, most of which weren't tenant-scoped UI at all:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;

&lt;span class="nx"&gt;useQuickSwitch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;after&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;async function switchTenant(tenantId: string) {&lt;br&gt;
  const { accessToken } = await api.post('/auth/switch-tenant', { tenantId });&lt;br&gt;
  await client.clearStore(); // empties the cache, no refetch of inactive queries&lt;br&gt;
  setAccessToken(accessToken);&lt;br&gt;
  // active queries now re-render against an empty cache and fetch fresh,&lt;br&gt;
  // instead of painting stale data first&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;clearStore()&lt;/code&gt; empties the cache without immediately refetching everything the way &lt;code&gt;resetStore()&lt;/code&gt; does, so active components see a brief loading state rather than a full-page spinner, then fetch clean. That closed the immediate hole: an agent now sees a loading skeleton for a beat, never another customer's data.&lt;/p&gt;

&lt;p&gt;The structural fix addressed the actual gap, that the cache had no tenant awareness at all, so the next feature that skips a manual clear doesn't reopen the same hole. Every tenant-scoped query now carries an explicit &lt;code&gt;tenantId&lt;/code&gt; variable sourced from a reactive variable tied to the active session, and a type policy scopes the cache key by it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;

&lt;span class="nx"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ts&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;export const activeTenantVar = makeVar(null);&lt;/p&gt;

&lt;p&gt;export const cache = new InMemoryCache({&lt;br&gt;
  typePolicies: {&lt;br&gt;
    Query: {&lt;br&gt;
      fields: {&lt;br&gt;
        tickets: {&lt;br&gt;
          keyArgs: ['tenantId'],&lt;br&gt;
        },&lt;br&gt;
      },&lt;br&gt;
    },&lt;br&gt;
  },&lt;br&gt;
});&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight graphql"&gt;&lt;code&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;TicketsOverview.graphql&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;after&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;query TicketsOverview($tenantId: ID!) {&lt;br&gt;
  tickets(tenantId: $tenantId) {&lt;br&gt;
    id&lt;br&gt;
    subject&lt;br&gt;
    status&lt;br&gt;
    customer&lt;br&gt;
  }&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With &lt;code&gt;tenantId&lt;/code&gt; in both the query variables and the type policy's &lt;code&gt;keyArgs&lt;/code&gt;, Acme Corp's tickets and Globex's tickets now live under distinct cache keys, the same way two different pages of a paginated list already did. Even if a future feature forgets to clear the store on an identity change, the cache itself can no longer confuse one tenant's entry for another's; the identity is part of the key, not an invisible side channel.&lt;/p&gt;




&lt;h2&gt;
  
  
  the aftermath
&lt;/h2&gt;

&lt;p&gt;1,900 Quick Switch actions recorded in the 30 days before the fix&lt;/p&gt;

&lt;p&gt;46 Switches, across 12 agents, where session telemetry showed the stale render held for over 500ms&lt;/p&gt;

&lt;p&gt;6 weeks Time between Quick Switch shipping and the first reported sighting&lt;/p&gt;

&lt;p&gt;0 Stale cross-tenant renders recorded since the tenantId-scoped cache shipped&lt;/p&gt;

&lt;p&gt;Because this was an internal tool, not customer-facing, no external data was exposed. It still landed as a corrective action under the company's SOC 2 tenant-isolation control, since that commitment covers who can see whose data regardless of which side of the product a screen sits on. A runtime assertion now runs on every tenant-scoped cache read in production: an Apollo Link compares the tenant embedded in a cached response against the tenant on the current session's JWT, and logs an alert if they ever disagree again. A synthetic test performs a scripted Quick Switch on a throttled connection in CI and fails the build if any tenant-scoped field renders data tagged with a different tenant ID than the active session, even for a single frame.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  A cache-first policy is only as safe as the assumption that the same query and variables always mean the same data. Tenant identity, or any other value carried outside the query itself, breaks that assumption the moment it changes without the cache knowing.&lt;/li&gt;
&lt;li&gt;  Removing a slow safety step to make something feel instant is a real product tradeoff, not automatically a mistake, but it needs to be made with full knowledge of what the slow step was actually doing. Here it was quietly acting as the only thing scoping the cache by identity.&lt;/li&gt;
&lt;li&gt;  A bug built entirely out of individually correct network requests and responses won't show up by inspecting the network tab alone. It has to be watched as a render sequence over time, which is why throttling the connection to stretch the window mattered more than any single log line.&lt;/li&gt;
&lt;li&gt;  If a value determines which data a query is allowed to return, it belongs in the cache key, not only in a header. Otherwise the client is trusting the server's authorization to also double as the client's own cache boundary, and those are two different systems that happen to usually agree.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The server had been doing its job the entire six weeks, never once returning a ticket to a tenant that didn't own it. The leak was one layer up, in a cache that had no way of knowing two different customers had ever asked it the exact same question.&lt;/p&gt;

</description>
      <category>react</category>
      <category>javascript</category>
      <category>webdev</category>
      <category>frontend</category>
    </item>
    <item>
      <title>How a Retry Loop</title>
      <dc:creator>Darshan Turakhia</dc:creator>
      <pubDate>Sat, 19 Sep 2026 10:55:58 +0000</pubDate>
      <link>https://dev.to/darshan_turakhia/how-a-retry-loop-3p0b</link>
      <guid>https://dev.to/darshan_turakhia/how-a-retry-loop-3p0b</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;03:14 UTC. Site search starts returning 503s. Not slow, not degraded — down. The on-call engineer's first Kibana query for search-service error rates times out too, which is its own bad sign, because Kibana is querying the same cluster that's supposedly the problem.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  the setup
&lt;/h2&gt;

&lt;p&gt;The cluster in question is a six-node self-managed Elasticsearch 8.x deployment handling two very different jobs. During the day it serves product and content search for the site, low latency, high query volume, nothing exotic. Overnight it also feeds a batch export: a Node job called &lt;code&gt;order-events-exporter&lt;/code&gt; that scrolls the &lt;code&gt;order_events&lt;/code&gt; index, about 600 million documents, and streams every doc into Snowflake for the analytics team. It's run at 03:00 UTC for two years without incident.&lt;/p&gt;

&lt;p&gt;The export uses the Scroll API, which works by taking a point-in-time snapshot of the index's segments the moment the first request runs, then letting you page through that snapshot with a &lt;code&gt;scroll_id&lt;/code&gt; until you're done. The snapshot is what makes it consistent, but keeping it alive costs the cluster something: every open scroll context pins the segments it's reading in place, which blocks merges and keeps otherwise-deletable segments resident in memory until the context is closed or its keep-alive expires.&lt;/p&gt;




&lt;h2&gt;
  
  
  the scramble
&lt;/h2&gt;

&lt;p&gt;The paged engineer's first theory was a hot shard from a relevance-ranking deploy that had shipped the previous afternoon. It touched the search query builder, so it was the obvious suspect. Rolling it back didn't help; the 503s kept coming at the same rate.&lt;/p&gt;

&lt;p&gt;Second theory: traffic spike. The dashboard showed request volume roughly flat against the same time yesterday, which ruled that out inside a few minutes, but the CPU graph in the same panel was doing something odd, climbing in a slow, steady ramp rather than the spiky pattern a traffic surge produces. Steady ramps usually mean something is accumulating, not spiking.&lt;/p&gt;

&lt;p&gt;Third theory, and the one that actually got investigated properly: a JVM heap problem. Old-gen usage on the two data nodes handling the &lt;code&gt;order_events&lt;/code&gt; shards was sitting above 90%, high enough to explain GC pauses long enough to look like request timeouts. But "the heap is full" isn't a root cause, it's a symptom, and nobody could yet say full of what.&lt;/p&gt;




&lt;h2&gt;
  
  
  the hunt
&lt;/h2&gt;

&lt;p&gt;The cluster logs had already said what was wrong, twenty minutes before the page fired. It just hadn't been paired with the right alert:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

elasticsearch.log, 02:54:11 UTC

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;[order-events-node-2] TooManyScrollContextsException: Trying to create too many scroll&lt;br&gt;
contexts. Must be less than or equal to: [500]. This limit can be set by changing the&lt;br&gt;
[search.max_open_scroll_context] setting.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That exception is real and specific: Elasticsearch caps concurrent open scroll contexts at 500 by default, cluster-wide, precisely because each one is a standing memory liability. Once the cap is hit, every &lt;em&gt;new&lt;/em&gt; scroll request gets rejected, which explained why the export job had been failing since shortly after 02:54, but not yet why site search itself was down; a rejected scroll creation shouldn't touch ordinary, non-scroll search queries at all.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;GET _nodes/stats/indices/search&lt;/code&gt; gave the number that connected the two:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;

GET \_nodes/stats/indices/search

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;"order-events-node-2": {&lt;br&gt;
  "search": {&lt;br&gt;
    "open_contexts": 1847,&lt;br&gt;
    "scroll_current": 1847,&lt;br&gt;
    "scroll_total": 2214&lt;br&gt;
  }&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;1,847 open contexts, well past the 500 the cluster is supposed to allow. That number only makes sense if contexts were being created faster than they were ever being closed or expired, and closed faster than they were also being &lt;em&gt;renewed&lt;/em&gt; past their keep-alive, which pointed straight at the export job's retry logic:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;

&lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;exporter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;js&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;before&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;async function exportBatch() {&lt;br&gt;
  for (let attempt = 0; attempt &amp;lt; 5; attempt++) {&lt;br&gt;
    try {&lt;br&gt;
      let res = await client.search({&lt;br&gt;
        index: 'order_events',&lt;br&gt;
        scroll: '10m',&lt;br&gt;
        size: 5000,&lt;br&gt;
        body: { query: { match_all: {} } },&lt;br&gt;
      });&lt;br&gt;
      return await drainScroll(res);&lt;br&gt;
    } catch (err) {&lt;br&gt;
      log.warn(&lt;code&gt;export attempt \${attempt} failed, retrying\&lt;/code&gt;, err);&lt;br&gt;
      // falls through to the next loop iteration and calls&lt;br&gt;
      // client.search() again from scratch&lt;br&gt;
    }&lt;br&gt;
  }&lt;br&gt;
  throw new Error('export failed after 5 attempts');&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a client-side timeout, the &lt;code&gt;catch&lt;/code&gt; block didn't resume the existing &lt;code&gt;scroll_id&lt;/code&gt;, because it never had one to resume, the request had already timed out client-side before the response carrying it arrived. It just looped and called &lt;code&gt;client.search()&lt;/code&gt; again, which creates a brand-new scroll context server-side, every single time. The old context the timed-out request had actually created was never told to close.&lt;/p&gt;

&lt;p&gt;Normally this would be forgiving. A scroll context's default keep-alive is short, and an abandoned one expires and gets cleaned up on its own. But eight months earlier, a different slowness ticket had bumped this job's keep-alive from &lt;code&gt;1m&lt;/code&gt; to &lt;code&gt;10m&lt;/code&gt; "to give the export more breathing room," and three weeks before this incident, a mapping change had added a large nested &lt;code&gt;line_items&lt;/code&gt; field to &lt;code&gt;order_events&lt;/code&gt; to support a new analytics report. Scroll pages that used to return in roughly 80ms were now taking 600-700ms under load, comfortably past the export client's 5-second per-page timeout during the cluster's overnight compaction window. The job was retrying constantly, each retry left a fresh 10-minute-lived context behind, and contexts were piling up roughly ten times faster than they were expiring.&lt;/p&gt;




&lt;h2&gt;
  
  
  the find
&lt;/h2&gt;

&lt;p&gt;Root cause: the export job's retry loop restarted the scroll from scratch on every timeout instead of resuming or explicitly closing the abandoned context, and a keep-alive that had been widened months earlier for an unrelated fix let each orphaned context live ten times longer than the default. Combined with a recent mapping change that slowed scroll pages enough to trigger the retry loop constantly, contexts accumulated past the 500-context cap within about twelve minutes of the job starting. Past that point, the pinned segments those contexts held open pushed old-gen heap usage over the cluster's parent circuit breaker threshold, which trips for &lt;em&gt;all&lt;/em&gt; request types, not just scroll, once triggered:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

elasticsearch.log, 03:13 UTC — one minute before the page

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;[order-events-node-2] [parent] Data too large, data for [] would be&lt;br&gt;
[16.2gb/95%], which is larger than the limit of [15.9gb/94%], real usage: [16.1gb],&lt;br&gt;
new bytes reserved: [112kb]: circuit_breaking_exception&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the line that took down site search. The parent circuit breaker doesn't distinguish between a batch export's scroll request and a shopper's product query, once it trips it rejects both.&lt;/p&gt;




&lt;h2&gt;
  
  
  the fix
&lt;/h2&gt;

&lt;p&gt;Immediate mitigation was a manual context flush, which closes every open scroll on the cluster and gives the circuit breaker room to reset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

manual recovery during the incident

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;curl -X DELETE "localhost:9200/_search/scroll/_all"&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Heap usage dropped from 95% to 61% within about ninety seconds and search traffic recovered on its own. The durable fix had two parts. First, the retry logic now resumes by &lt;code&gt;scroll_id&lt;/code&gt; and always cleans up on final failure instead of leaking:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;

&lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;exporter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;js&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;after&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;async function exportBatch() {&lt;br&gt;
  let scrollId;&lt;br&gt;
  try {&lt;br&gt;
    let res = await client.search({&lt;br&gt;
      index: 'order_events',&lt;br&gt;
      scroll: '1m',&lt;br&gt;
      size: 5000,&lt;br&gt;
      body: { query: { match_all: {} } },&lt;br&gt;
    });&lt;br&gt;
    scrollId = res._scroll_id;&lt;br&gt;
    return await drainScroll(res, scrollId);&lt;br&gt;
  } finally {&lt;br&gt;
    if (scrollId) {&lt;br&gt;
      await client.clearScroll({ scroll_id: scrollId }).catch(() =&amp;gt; {});&lt;br&gt;
    }&lt;br&gt;
  }&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Second, the export was migrated off the Scroll API entirely, onto &lt;code&gt;search_after&lt;/code&gt; with a point-in-time (PIT) ID, which is Elasticsearch's own recommended replacement for scroll-based deep pagination and defaults to a much shorter, non-negotiable keep-alive per request rather than one long-lived session an app can forget to close. The keep-alive that had been widened months earlier no longer exists as a setting to accidentally leave too generous.&lt;/p&gt;




&lt;h2&gt;
  
  
  the aftermath
&lt;/h2&gt;

&lt;p&gt;38 min Site search returning 503s cluster-wide&lt;/p&gt;

&lt;p&gt;1,847 Peak open scroll contexts against a default cap of 500&lt;/p&gt;

&lt;p&gt;12 min Time from job start to the context cap being exceeded&lt;/p&gt;

&lt;p&gt;0 Incidents since migrating off Scroll to search_after + PIT&lt;/p&gt;

&lt;p&gt;A new Datadog metric now tracks &lt;code&gt;open_contexts&lt;/code&gt; per node with an alert at 300, 60% of the default cap, so the next accumulation gets caught while it's still just a batch job failing, not the whole cluster.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  A retry loop that doesn't know how to resume isn't retrying, it's restarting, and restarting a stateful server-side operation without cleaning up the state it already created is how you turn a transient timeout into a resource leak.&lt;/li&gt;
&lt;li&gt;  A keep-alive bump made for one incident months ago can quietly change the blast radius of an unrelated bug later. Nobody revisiting the retry logic knew the keep-alive had been widened, or that it mattered.&lt;/li&gt;
&lt;li&gt;  Elasticsearch's circuit breaker protects the whole cluster by design, which is exactly why an isolated batch job's misbehavior became a site-wide outage instead of staying contained to itself.&lt;/li&gt;
&lt;li&gt;  &lt;code&gt;TooManyScrollContextsException&lt;/code&gt; was in the logs twenty minutes before the page fired. The gap wasn't missing information, it was a missing alert on a log line nobody had thought to wire up, because nothing had ever hit that limit before.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Scroll API had been the right tool for this export for two years. It only stopped being the right tool the moment something downstream, a slower mapping, a wider keep-alive, a retry loop that didn't know its own state, started leaning on the one property Scroll assumes you'll respect: that whoever opens a context is the one responsible for closing it.&lt;/p&gt;

</description>
      <category>elasticsearch</category>
      <category>backend</category>
      <category>database</category>
      <category>devops</category>
    </item>
    <item>
      <title>B-Tree vs LSM-Tree: the storage-engine tradeoff behind every database you use</title>
      <dc:creator>Darshan Turakhia</dc:creator>
      <pubDate>Tue, 15 Sep 2026 17:32:22 +0000</pubDate>
      <link>https://dev.to/darshan_turakhia/b-tree-vs-lsm-tree-the-storage-engine-tradeoff-behind-every-database-you-use-3l3l</link>
      <guid>https://dev.to/darshan_turakhia/b-tree-vs-lsm-tree-the-storage-engine-tradeoff-behind-every-database-you-use-3l3l</guid>
      <description>&lt;p&gt;Every relational database you've ever used made a bet at the storage-engine level, before you wrote a single query: reads matter more than writes, or writes matter more than reads. That bet shows up as a choice between two data structures, a B-tree or an LSM-tree. Most engineers never think about it. It's still deciding how their database behaves under load.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual tradeoff
&lt;/h2&gt;

&lt;p&gt;A B-tree updates data in place. It walks down to the exact leaf page a key belongs on and rewrites it right there, so the tree stays fully sorted at all times. A read never does more than follow one path from root to leaf. There's no ambiguity about where a key lives, because there's only ever one copy of it.&lt;/p&gt;

&lt;p&gt;An LSM-tree (log-structured merge-tree) refuses to do that. It never touches old data on a write. Instead it buffers writes in memory, flushes them to new immutable files on disk, and merges those files together later in the background. A write is a sequential append, not a random rewrite, which is why LSM-trees dominate anywhere ingest volume is the bottleneck: Cassandra, RocksDB, LevelDB, HBase. B-trees stay the default for read-heavy, point-lookup workloads: Postgres, MySQL's InnoDB, basically every traditional RDBMS.&lt;/p&gt;

&lt;p&gt;What's worth understanding is what each side gives up, not which one "wins."&lt;/p&gt;

&lt;p&gt;B-trees pay at write time. Once your dataset outgrows memory, a page update that used to be a cheap in-memory write becomes a real disk seek. Under heavy write load the tree also spends real time rebalancing, splitting pages as they fill so every leaf stays at the same depth.&lt;/p&gt;

&lt;p&gt;LSM-trees pay at read time instead. A single key can exist simultaneously in the in-memory memtable, the most recently flushed file, and several older files. The newest write wins, but a lookup has to check each location until it finds a match. Bloom filters make this bearable in practice (a probabilistic "this file definitely doesn't have your key" check that skips most files without opening them), but a worst-case LSM-tree lookup is still checking more places than a B-tree ever has to.&lt;/p&gt;

&lt;p&gt;There's also a cost LSM-trees defer rather than eliminate: compaction, the background process merging small files into larger ones and dropping superseded key versions. This is the part people underestimate. It's not free maintenance sitting quietly in the background, it's competing for the same disk I/O and CPU your live traffic needs. A compaction backlog is the classic LSM-tree failure mode, and it's an ugly one: writes outpace merging, read latency climbs as lookups check more uncompacted files, and the whole thing spirals if nothing throttles incoming writes to let compaction catch up. If you've operated Cassandra under sustained write pressure, you've probably met this one at 3am.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this actually bites you
&lt;/h2&gt;

&lt;p&gt;This decision almost never gets made consciously by application code. It gets made once, implicitly, by whichever database a team picked years ago. So the first time most engineers actually think about it is after a write-heavy workload starts struggling on a B-tree-backed system that was never built for it. Time-series data, event logging, anything ingesting far more than it reads back in the same window: that's LSM-tree territory. Building it on Postgres because "that's what we use" is how you end up fighting the storage engine instead of your actual problem.&lt;/p&gt;

&lt;p&gt;Some engines hedge instead of picking a side outright. MySQL's MyRocks storage engine swaps InnoDB's B-tree for an LSM-tree underneath the same relational interface, specifically for workloads write-bound enough to justify the read-side cost.&lt;/p&gt;

&lt;p&gt;I wrote up the fuller comparison, including read and write amplification and where each structure shows up in real systems, &lt;a href="https://nodique.com/comparisons/b-tree-vs-lsm-tree" rel="noopener noreferrer"&gt;here&lt;/a&gt;, if you want the longer version of this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  Seeing it instead of reading about it
&lt;/h2&gt;

&lt;p&gt;The part that's genuinely hard to build intuition for from a blog post is why a B-tree stays shallow even at billions of rows. So I built &lt;a href="https://nodique.com/tools/b-tree-visualizer" rel="noopener noreferrer"&gt;an interactive B-tree visualizer&lt;/a&gt;. Insert keys and watch a node split once it fills, the middle key pushing up into the parent every time. Search afterward and it highlights exactly which nodes the lookup touches. Entirely client-side, no backend. Took me about four minutes of playing with it to have the mechanism actually click in a way the text version never quite managed.&lt;/p&gt;

&lt;p&gt;Next time someone asks why adding an index didn't help their write-heavy table, or why their time-series database "just feels different" from Postgres under load, this is the mechanism underneath both answers.&lt;/p&gt;

</description>
      <category>database</category>
      <category>backend</category>
      <category>computerscience</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
