<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mubashir Taj</title>
    <description>The latest articles on DEV Community by Mubashir Taj (@mubashirtaj).</description>
    <link>https://dev.to/mubashirtaj</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4141856%2F104507fc-2bcf-482e-abf8-6205f16f43ff.png</url>
      <title>DEV Community: Mubashir Taj</title>
      <link>https://dev.to/mubashirtaj</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mubashirtaj"/>
    <language>en</language>
    <item>
      <title>Caching from Zero to Production</title>
      <dc:creator>Mubashir Taj</dc:creator>
      <pubDate>Tue, 29 Sep 2026 13:19:41 +0000</pubDate>
      <link>https://dev.to/mubashirtaj/caching-from-zero-to-production-hh4</link>
      <guid>https://dev.to/mubashirtaj/caching-from-zero-to-production-hh4</guid>
      <description>&lt;p&gt;Every backend eventually hits the same wall: the database cannot keep up, but the data barely changes. That is the moment caching stops being an optimization and becomes a requirement. This article walks through the full path, from the basic pattern to the failure modes that take down production systems, with Go examples throughout.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Why cache at all
&lt;/h2&gt;

&lt;p&gt;It comes down to orders of magnitude. Reading from memory takes nanoseconds. Reading from disk takes milliseconds. A network round trip to the database takes longer still. When the same data is requested thousands of times per second and changes rarely, serving it from the database every time is pure waste.&lt;/p&gt;

&lt;p&gt;The two numbers that justify a cache: request volume and data stability. High volume plus stable data means caching will pay for itself. Low volume or constantly changing data means you are adding operational complexity for little gain. Cache deliberately, not by default.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The cache-aside pattern
&lt;/h2&gt;

&lt;p&gt;Cache-aside, also called lazy loading, is the most common pattern and the right starting point. The application owns the logic:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Check the cache for the key.&lt;/li&gt;
&lt;li&gt;On a hit, return the value.&lt;/li&gt;
&lt;li&gt;On a miss, read from the database, write the result into the cache with a TTL, then return it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here is a compact Go example using Redis:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;GetUser&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rdb&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;sql&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DB&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;User&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="s"&gt;"user:"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;

    &lt;span class="c"&gt;// 1. Try the cache first&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;rdb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Bytes&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt; &lt;span class="n"&gt;User&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Unmarshal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cached&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c"&gt;// 2. Miss: fall through to the database&lt;/span&gt;
    &lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;fetchUserFromDB&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c"&gt;// 3. Populate the cache for next time (TTL: 5 minutes)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Marshal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;rdb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Minute&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The cache stays simple because it never talks to the database directly. The trade-off is the first request after every expiry pays the full database cost, and stale data is possible between the write to the database and the next cache refresh. For most read-heavy workloads, that trade-off is correct.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. TTLs, and why fixed TTLs are dangerous
&lt;/h2&gt;

&lt;p&gt;The TTL is the most consequential number in your caching layer, and teams often set it once and forget it. A fixed TTL on every key creates a subtle trap: keys written at the same time expire at the same time.&lt;/p&gt;

&lt;p&gt;Consider a hot key absorbing 5,000 requests per second. Its TTL expires. In the next instant, all 5,000 requests fall through to the database simultaneously. CPU climbs from 40% to 100% in seconds, the connection pool saturates, and latencies spike. This is the thundering herd, also called cache stampede. The cache did not fail. The expiry pattern synchronized thousands of requests into a single attack on your own database.&lt;/p&gt;

&lt;p&gt;The cheapest fix is jitter. Instead of a fixed 300-second TTL, use 300 seconds plus a random offset of up to 60 seconds. Expirations spread across time, and the synchronized cliff disappears:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="c"&gt;// Jittered TTL: spreads expirations so hot keys never die together&lt;/span&gt;
&lt;span class="n"&gt;ttl&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Minute&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Duration&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rand&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Intn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="m"&gt;60&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Second&lt;/span&gt;
&lt;span class="n"&gt;rdb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ttl&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One line of randomness removes an entire class of incident. This belongs in every caching layer from day one.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Stampede protections
&lt;/h2&gt;

&lt;p&gt;Jitter handles the synchronized case, but a genuinely cold key under heavy load still funnels every request to the database. Three protections cover the remaining scenarios.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Request coalescing.&lt;/strong&gt; When 5,000 requests need the same missing key, only one should query the database. In Go, &lt;code&gt;golang.org/x/sync/singleflight&lt;/code&gt; does exactly this: concurrent callers for the same key share a single in-flight execution.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="n"&gt;group&lt;/span&gt; &lt;span class="n"&gt;singleflight&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Group&lt;/span&gt;

&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;GetProduct&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;Product&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;group&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Do&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"product:"&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;func&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;interface&lt;/span&gt;&lt;span class="p"&gt;{},&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;GetProductCached&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c"&gt;// cache-aside logic from section 2&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;Product&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;5,000 concurrent requests become one database query. The rest wait microseconds for the shared result.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stale-while-revalidate.&lt;/strong&gt; Serve the slightly old value instantly and refresh it in the background. Implementation: on a miss, check for a stale copy kept under a separate key with a longer TTL. If present, return it immediately and trigger an async refresh. Users get a fast response, the database gets one query instead of thousands. This is the right choice when slightly stale data is acceptable, which covers far more endpoints than teams assume.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Probabilistic early refresh.&lt;/strong&gt; For the highest-traffic keys, refresh the cache before the TTL expires, with probability increasing as expiry approaches. This keeps hot keys warm permanently and removes the miss path entirely for your most critical data.&lt;/p&gt;

&lt;p&gt;Use all three together and stampedes become a non-issue rather than an incident category.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Eviction policies
&lt;/h2&gt;

&lt;p&gt;Memory is finite, so Redis must decide what to remove when it fills up. The &lt;code&gt;maxmemory-policy&lt;/code&gt; setting controls this, and the default (no eviction, writes start failing) surprises people.&lt;/p&gt;

&lt;p&gt;The policies worth knowing:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;allkeys-lru&lt;/strong&gt;: evict the least recently used keys across everything. The best default for a general-purpose cache, because it automatically keeps hot data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;volatile-lru&lt;/strong&gt;: evict least recently used, but only among keys that have a TTL. Use this when some keys must never be evicted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;allkeys-lfu&lt;/strong&gt;: evict the least frequently used keys. Better than LRU when access patterns have long-term hot keys mixed with one-off scans that would pollute an LRU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;volatile-ttl&lt;/strong&gt;: evict keys with the shortest remaining TTL first. A reasonable choice when TTLs genuinely reflect data priority.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Set it explicitly in your Redis configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;maxmemory&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;&lt;span class="n"&gt;gb&lt;/span&gt;
&lt;span class="n"&gt;maxmemory&lt;/span&gt;-&lt;span class="n"&gt;policy&lt;/span&gt; &lt;span class="n"&gt;allkeys&lt;/span&gt;-&lt;span class="n"&gt;lru&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An unset eviction policy is a latent outage. Decide this before production, not during it.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Write strategies
&lt;/h2&gt;

&lt;p&gt;Reads are only half the story. When data changes, the cache and the database must agree, or at least disagree briefly and predictably.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cache-aside with invalidation&lt;/strong&gt; (the common path): on write, update the database, then delete the cache key. The next read repopulates it. Simple and correct for most cases. The small risk is a read repopulating stale data between the database write and the delete, which the delete-then-verify patterns address if you need them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write-through&lt;/strong&gt;: on write, update the cache and the database together. Reads never miss, but every write pays both costs, and you cache data nobody may read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write-behind&lt;/strong&gt;: on write, update the cache immediately and the database asynchronously. Lowest write latency, but a crash before the async write means data loss. Only for data you can afford to lose or reconstruct.&lt;/p&gt;

&lt;p&gt;For most systems: cache-aside with invalidation on writes. Reach for the others when a measured requirement demands it.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Observability: the metrics that matter
&lt;/h2&gt;

&lt;p&gt;A cache you cannot observe is a liability. Four metrics cover nearly everything:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Hit ratio&lt;/strong&gt; (hits / total requests). Below 80% on a read-heavy endpoint means the cache is not earning its keep. Investigate key design, TTLs, or whether the data is even cacheable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;p99 latency with and without cache&lt;/strong&gt;. The gap between cache hits and database fallbacks tells you the real value of the layer, and a shrinking gap means the database path needs attention.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Eviction rate&lt;/strong&gt;. A sudden rise means working set exceeded memory. Either grow the instance or tighten what you cache.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stampede signals&lt;/strong&gt;: watch for correlated spikes in database connections and cache misses on the same key pattern. That correlation is the fingerprint of a herd forming.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Alert on eviction rate and on miss-rate spikes for your top keys. Those two catch most caching incidents before users notice.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Production checklist
&lt;/h2&gt;

&lt;p&gt;Before calling your caching layer production-ready:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;TTLs are jittered, not fixed.&lt;/li&gt;
&lt;li&gt;Request coalescing (singleflight or equivalent) guards every hot path.&lt;/li&gt;
&lt;li&gt;An eviction policy is set explicitly and matches the workload.&lt;/li&gt;
&lt;li&gt;Write invalidation is tested, including the failure case where the delete fails.&lt;/li&gt;
&lt;li&gt;Hit ratio, eviction rate, and p99 latency are dashboarded and alerted.&lt;/li&gt;
&lt;li&gt;The cache failing does not take down the application. Every cache call has a timeout and a fallback to the database. A cache outage should degrade latency, never cause an outage.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That last point deserves emphasis. The cache is an optimization layer. If its absence breaks you, the architecture has the dependency backwards.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;Caching is simple to start and easy to get wrong at scale. The pattern fits in thirty lines of code, but production demands jittered TTLs, coalesced requests, explicit eviction, and real observability. Get those right and the cache becomes the quietest, most reliable part of your stack.&lt;/p&gt;

&lt;p&gt;What is your go-to stampede protection in production?&lt;/p&gt;

</description>
      <category>caching</category>
      <category>redis</category>
      <category>backend</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>hello developerws</title>
      <dc:creator>Mubashir Taj</dc:creator>
      <pubDate>Thu, 24 Sep 2026 19:33:20 +0000</pubDate>
      <link>https://dev.to/mubashirtaj/hello-developerws-4bok</link>
      <guid>https://dev.to/mubashirtaj/hello-developerws-4bok</guid>
      <description>&lt;p&gt;hello developers&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
