Every time a RAG application answers a question, it may need to retrieve documents and make an LLM call.
But what happens when someone asks essentially the same question again?
We could reuse the previous answer. The challenge is knowing when two questions are similar enough to share an answer—and when they only look similar.
That's the problem I explored while building Weir, an open-source, cost-aware gateway for Retrieval-Augmented Generation (RAG).
GitHub: https://github.com/yatinannam/weir
What is Weir?
Weir sits in front of an existing RAG service. It decides whether a request can safely use a cached answer or needs retrieval and generation, and chooses a model tier when generation is necessary.
It combines three main ideas:
- Semantic caching: Reuse answers for sufficiently similar questions.
- Model routing: Use a cheaper model for suitable questions and a larger model when needed.
- Cost accounting: Track actual request cost, counterfactual cost, and latency so savings can be measured.
It also includes request logging, monitoring, evaluation tooling, and a Docker-based demo.
The surprising problem with semantic caching
Suppose a user asks:
"What are the ICU visiting hours?"
Later, another user asks:
"What are the general ward visiting hours?"
The questions are structurally similar. An embedding-based cache might assign them a high similarity score.
But serving the ICU answer to the second user would be incorrect.
This is why Weir uses an entity guard alongside embedding similarity. It checks distinguishing information such as numbers, negation, and important domain terms before accepting a potential cache hit.
The goal isn't to cache as aggressively as possible. It's to save unnecessary work without silently returning the wrong answer.
How I evaluated it
I compared four configurations:
- A baseline that always uses the larger model without caching
- Semantic caching alone
- Model routing alone
- Full Weir, combining caching and routing
The evaluation used a fictional hospital FAQ with 40 documents and 158 evaluation questions, including paraphrases, look-alike questions, and unanswerable questions.
1. Cold-pass evaluation
Each of the 158 questions was asked once.
Full Weir reduced measured cost per 1,000 requests from $0.0988 to $0.0786—a 20% reduction. The evaluation judge score and key-fact score remained unchanged against the baseline.
The trade-off matters: median latency increased from 767 ms to 876 ms in this workload. Cost reduction does not automatically mean every workload becomes faster.
2. Repeat-heavy replay
I also tested a 300-request replay in which 73.7% of requests were repeats.
| Metric | Baseline | Full Weir |
|---|---|---|
| Cache hit rate | 0% | 93.0% |
| Cost per 1,000 requests | $0.0970 | $0.0055 |
| Median latency | 747 ms | 14 ms |
| Judge score | 4.79 | 4.79 |
| Key-fact score | 0.947 | 0.948 |
On this workload, the measured cost reduction was approximately 94%, and median latency was about 50 times lower.
The important qualification is that these are synthetic, repeat-heavy benchmark results—not proof of equivalent savings on a production workload.
What about incorrect cache hits?
At the chosen similarity threshold, embedding similarity alone accepted some dangerous look-alike matches in the evaluation set. Adding the entity guard reduced wrong accepted cache hits to zero on that test set.
That is encouraging, but it is not a guarantee of safety. A small synthetic dataset cannot cover every possible ambiguity or domain-specific distinction.
What happened under load?
I also tested the system under load, including provider failures and database delays.
The tests surfaced a genuine overload mode: when cache lookups began timing out under CPU pressure, more requests bypassed the cache and reached retrieval and generation, adding further load.
I documented that limitation rather than presenting the load test as an unconditional capacity claim.
Try it yourself
Weir includes a one-command demo launcher:
uv run demo.py
You'll need Docker Desktop and uv. A Groq API key is optional; without one, the demo uses a stub model, so generation is simulated while the cache, retrieval, guard, and routing components remain real.
The repository includes the implementation, evaluation reports, architecture, limitations, and instructions for running the tests.
What I'd love feedback on
I'm particularly interested in feedback from people building RAG applications or LLM infrastructure:
- What failure cases would you add to the semantic-cache evaluation?
- How would you make the entity guard more robust across domains?
- What real-world workload would you use to test whether these savings translate to production?
I'd appreciate technical criticism, benchmark suggestions, and contributions.
Top comments (4)
@yatinannam , the "entity guard" is the exact missing piece in most naive semantic caching implementations. solving the "icu vs general ward" embedding trap is a massive win for production rag.
to answer your question on making the entity guard more robust across domains: instead of relying purely on heuristics, you could run a lightweight, local named entity recognition (ner) pipeline (like a small spacy model or a tiny quantized local llm) on the incoming query. extract the core domain entities and require an exact or fuzzy match on those entities before even calculating the vector similarity. if the core entities mismatch, it's an instant cache miss.
regarding the "overload mode" you mentioned—when cache lookups time out and everything cascades into the expensive rag pipeline—this is a fatal flaw for constrained budgets. on a $150 phone or a tight api budget, you cannot afford a thundering herd. a good failure case to test is a "circuit breaker" or "stale-while-revalidate" approach: if the cache lookup is slow, serve the previous answer with a "stale" flag, rather than hitting the heavy retrieval and generation pipeline.
cost-aware routing isn't just a nice-to-have; for indie devs and constrained environments, it's survival. weir is exactly the kind of pragmatic infrastructure the community needs. great work! 🐯
Wow, thanks for the detailed response! The ICU vs general ward example is exactly the type of failure case I was hoping the entity guard would catch.
The NER approach is interesting. My current guard relies on lexical checks for identifying information such as numbers, negation, and domain-specific terms. A lightweight NER layer could enhance adaptability to various domains, but I need to evaluate potential latency and new failure modes before integrating it into the cache path.
The point of overload is also important. I literally wrote up a failure mode where cache lookup timeouts were making requests skip the cache putting more pressure on retrieval and generation. A circuit breaker sounds to be explored here.
But I'd be more careful with stale-while-revalidate For static public FAQs, serving a clearly marked stale answer might be a reasonable fallback. Weir currently skips caching for questions that are clinical, personalized, or time-sensitive, because stale information can be worse than failing safely.
Thanks for the specific suggestions. Those are the very edge cases I want to dig into next!!!
The ICU vs general ward example is exactly the kind of thing that makes me distrust plain embedding similarity, so the entity guard idea hits home. Curious how it handles synonyms though, like
Good question! Currently, Weir handles synonyms through an explicit, configurable lexicon rather than trying to infer them dynamically.
For example, “ICU”, “intensive care unit”, and “critical care” are mapped to the same canonical term, while “general ward” maps to a separate one. This lets the guard recognize known aliases without treating two different wards as equivalent.
The limitation is that the current lexicon is curated around the fictional hospital knowledge base used for evaluation, so it won't automatically understand every unseen or domain-specific synonym. That's definitely an area I'd like to explore further, especially how to improve synonym coverage without introducing unsafe cache hits or adding unnecessary latency to every request.
Do you have a particular synonym pair in mind? I'd be interested in adding examples like that to the evaluation set!