<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Yatin Annam</title>
    <description>The latest articles on DEV Community by Yatin Annam (@yatinannam).</description>
    <link>https://dev.to/yatinannam</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4162362%2Fbad905dd-bbdd-4c89-9117-7d3ab1485665.jpg</url>
      <title>DEV Community: Yatin Annam</title>
      <link>https://dev.to/yatinannam</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yatinannam"/>
    <language>en</language>
    <item>
      <title>I Built a Semantic Cache for RAG. The Hard Part Was Knowing When NOT to Cache.</title>
      <dc:creator>Yatin Annam</dc:creator>
      <pubDate>Fri, 09 Oct 2026 15:46:52 +0000</pubDate>
      <link>https://dev.to/yatinannam/i-built-a-semantic-cache-for-rag-the-hard-part-was-knowing-when-not-to-cache-30fa</link>
      <guid>https://dev.to/yatinannam/i-built-a-semantic-cache-for-rag-the-hard-part-was-knowing-when-not-to-cache-30fa</guid>
      <description>&lt;p&gt;Every time a RAG application answers a question, it may need to retrieve documents and make an LLM call.&lt;/p&gt;

&lt;p&gt;But what happens when someone asks essentially the same question again?&lt;/p&gt;

&lt;p&gt;We could reuse the previous answer. The challenge is knowing when two questions are similar enough to share an answer—and when they only &lt;em&gt;look&lt;/em&gt; similar.&lt;/p&gt;

&lt;p&gt;That's the problem I explored while building &lt;strong&gt;Weir&lt;/strong&gt;, an open-source, cost-aware gateway for Retrieval-Augmented Generation (RAG).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/yatinannam/weir" rel="noopener noreferrer"&gt;https://github.com/yatinannam/weir&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Weir?
&lt;/h2&gt;

&lt;p&gt;Weir sits in front of an existing RAG service. It decides whether a request can safely use a cached answer or needs retrieval and generation, and chooses a model tier when generation is necessary.&lt;/p&gt;

&lt;p&gt;It combines three main ideas:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Semantic caching:&lt;/strong&gt; Reuse answers for sufficiently similar questions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model routing:&lt;/strong&gt; Use a cheaper model for suitable questions and a larger model when needed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost accounting:&lt;/strong&gt; Track actual request cost, counterfactual cost, and latency so savings can be measured.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It also includes request logging, monitoring, evaluation tooling, and a Docker-based demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  The surprising problem with semantic caching
&lt;/h2&gt;

&lt;p&gt;Suppose a user asks:&lt;/p&gt;

&lt;p&gt;"What are the ICU visiting hours?"&lt;/p&gt;

&lt;p&gt;Later, another user asks:&lt;/p&gt;

&lt;p&gt;"What are the general ward visiting hours?"&lt;/p&gt;

&lt;p&gt;The questions are structurally similar. An embedding-based cache might assign them a high similarity score.&lt;/p&gt;

&lt;p&gt;But serving the ICU answer to the second user would be incorrect.&lt;/p&gt;

&lt;p&gt;This is why Weir uses an entity guard alongside embedding similarity. It checks distinguishing information such as numbers, negation, and important domain terms before accepting a potential cache hit.&lt;/p&gt;

&lt;p&gt;The goal isn't to cache as aggressively as possible. It's to save unnecessary work without silently returning the wrong answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I evaluated it
&lt;/h2&gt;

&lt;p&gt;I compared four configurations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A baseline that always uses the larger model without caching&lt;/li&gt;
&lt;li&gt;Semantic caching alone&lt;/li&gt;
&lt;li&gt;Model routing alone&lt;/li&gt;
&lt;li&gt;Full Weir, combining caching and routing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The evaluation used a fictional hospital FAQ with 40 documents and 158 evaluation questions, including paraphrases, look-alike questions, and unanswerable questions.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Cold-pass evaluation
&lt;/h3&gt;

&lt;p&gt;Each of the 158 questions was asked once.&lt;/p&gt;

&lt;p&gt;Full Weir reduced measured cost per 1,000 requests from $0.0988 to $0.0786—a 20% reduction. The evaluation judge score and key-fact score remained unchanged against the baseline.&lt;/p&gt;

&lt;p&gt;The trade-off matters: median latency increased from 767 ms to 876 ms in this workload. Cost reduction does not automatically mean every workload becomes faster.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Repeat-heavy replay
&lt;/h3&gt;

&lt;p&gt;I also tested a 300-request replay in which 73.7% of requests were repeats.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Baseline&lt;/th&gt;
&lt;th&gt;Full Weir&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cache hit rate&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;93.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost per 1,000 requests&lt;/td&gt;
&lt;td&gt;$0.0970&lt;/td&gt;
&lt;td&gt;$0.0055&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median latency&lt;/td&gt;
&lt;td&gt;747 ms&lt;/td&gt;
&lt;td&gt;14 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Judge score&lt;/td&gt;
&lt;td&gt;4.79&lt;/td&gt;
&lt;td&gt;4.79&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Key-fact score&lt;/td&gt;
&lt;td&gt;0.947&lt;/td&gt;
&lt;td&gt;0.948&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On this workload, the measured cost reduction was approximately 94%, and median latency was about 50 times lower.&lt;/p&gt;

&lt;p&gt;The important qualification is that these are synthetic, repeat-heavy benchmark results—not proof of equivalent savings on a production workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  What about incorrect cache hits?
&lt;/h2&gt;

&lt;p&gt;At the chosen similarity threshold, embedding similarity alone accepted some dangerous look-alike matches in the evaluation set. Adding the entity guard reduced wrong accepted cache hits to zero on that test set.&lt;/p&gt;

&lt;p&gt;That is encouraging, but it is not a guarantee of safety. A small synthetic dataset cannot cover every possible ambiguity or domain-specific distinction.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened under load?
&lt;/h2&gt;

&lt;p&gt;I also tested the system under load, including provider failures and database delays.&lt;/p&gt;

&lt;p&gt;The tests surfaced a genuine overload mode: when cache lookups began timing out under CPU pressure, more requests bypassed the cache and reached retrieval and generation, adding further load.&lt;/p&gt;

&lt;p&gt;I documented that limitation rather than presenting the load test as an unconditional capacity claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;Weir includes a one-command demo launcher:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv run demo.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You'll need Docker Desktop and &lt;code&gt;uv&lt;/code&gt;. A Groq API key is optional; without one, the demo uses a stub model, so generation is simulated while the cache, retrieval, guard, and routing components remain real.&lt;/p&gt;

&lt;p&gt;The repository includes the implementation, evaluation reports, architecture, limitations, and instructions for running the tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd love feedback on
&lt;/h2&gt;

&lt;p&gt;I'm particularly interested in feedback from people building RAG applications or LLM infrastructure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What failure cases would you add to the semantic-cache evaluation?&lt;/li&gt;
&lt;li&gt;How would you make the entity guard more robust across domains?&lt;/li&gt;
&lt;li&gt;What real-world workload would you use to test whether these savings translate to production?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'd appreciate technical criticism, benchmark suggestions, and contributions.&lt;/p&gt;

</description>
      <category>rag</category>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
