DEV Community

Cover image for My agent paid twice for the same question. The fix was a Redis key
Rakesh Singh
Rakesh Singh

Posted on

My agent paid twice for the same question. The fix was a Redis key

Two people in the same workspace asked my agent the same question minutes apart. Both runs made the same tool calls and returned the same cited answer. I paid full price twice.

The obvious fix is cache-aside: hash the question, store the answer. It fails because an answer also depends on the workspace, the document version, the model, and the prompt.

The rule: the cache key is everything the answer depends on. This post covers the exact-match and tool-result caches in Redis. Semantic caching is next.

Where does the money go in an agent system?

The system answers questions over telecom specification documents. It cites the exact section, or it says it cannot answer.

It has the shape most production agentic systems have: a cheap deterministic path first, and an agent only when that path fails.

Retrieval path. Search the indexed documents, generate an answer with citations, abstain if the evidence is missing. One model call.

Agent path. A bounded loop that plans, calls tools, and composes an answer. Several model calls and several tool calls per question.

Per question, the agent path costs several times what the retrieval path costs. That is where caching pays first.

Why does cache-aside fail for LLM answers?

The wrong model is "the key is the question." Cache-aside assumes three things, and all three are false here.

The key is known. A product page has an ID. A question is a sentence, and two users write the same question in different words.

The value can be rebuilt for free. There is no database row behind an answer. Rebuilding it means paying the model again, and the second run may word it differently.

The value depends only on the key. An answer also depends on the document version, on what the user is allowed to read, and on the model and prompt that produced it.

This post fixes the third assumption. The first one, the same question in different words, is the semantic cache. That is the next post in this series.

What belongs in the cache key?

Everything the answer depends on. If changing a thing can change the answer, that thing is a segment of the key.

ans:{workspace}:{corpus_version}:{model}:{prompt_version}:{sha256(question)}
Enter fullscreen mode Exit fullscreen mode

![The cache key is split into five labelled parts: prefix, workspace, corpus version, model and prompt version, and the hash of the question, with the failure caused by omitting each one](
The cache key split into five labelled parts: prefix, workspace, corpus version, model and prompt version, and the hash of the question, with the failure caused by omitting each one

)
Each part of the key and what goes wrong when it is left out.

  • workspace: every cached answer was authorised for someone. The workspace is always part of the key. There is no shared cache across tenants.
  • corpus_version: a re-index bumps the version, and every answer built on the old documents becomes unreachable without a single delete. The old entries die by TTL.
  • model and prompt_version: the answer came from one model and one prompt. Change either, and the saved answer belongs to a system you no longer run.
  • sha256(question): the question after normalising: lowercase, trim, collapse whitespace.

The value is the answer, its citations, and a timestamp, as JSON, written with a TTL. Only complete answers with citations are stored.

The agent loop gets its own cache. An agent calls the same tool with the same arguments repeatedly, inside one run and across runs. Unique questions still share sub-steps.

tool:{workspace}:{corpus_version}:{tool_name}:{sha256(canonical_args)}
Enter fullscreen mode Exit fullscreen mode

Here is the order of checks on every request:

![Flow diagram. A question goes to the exact-match cache, then the semantic cache, then the retrieval path, then the agent path, which uses a tool-result cache. Cache hits go straight to the answer](
Flow diagram. A question goes to the exact-match cache, then the semantic cache, then the retrieval path, then the agent path, which uses a tool-result cache. Cache hits go straight to the answer

)
Red boxes are Redis. Blue boxes pay for model calls. The semantic cache in the middle is the next post.

Start caching where the cost is highest, and measure before adding more.

What does the key look like in Redis?

import hashlib
import json


def sha(s: str) -> str:
    return hashlib.sha256(s.encode()).hexdigest()


def answer_key(ws, corpus_version, model, prompt_version, question):
    q = " ".join(question.lower().split())  # lowercase, trim, collapse whitespace
    # corpus_version is the line that matters
    return f"ans:{ws}:{corpus_version}:{model}:{prompt_version}:{sha(q)}"


def tool_key(ws, corpus_version, tool_name, args: dict):
    canonical = json.dumps(args, sort_keys=True, separators=(",", ":"))
    return f"tool:{ws}:{corpus_version}:{tool_name}:{sha(canonical)}"
Enter fullscreen mode Exit fullscreen mode

The line that matters is corpus_version. Without it, a re-index leaves every old answer reachable, and you need a delete job to clean up. With it, old answers stop matching on their own.

The exact-match cache costs one GET per request. It only hits on identical text, so it never returns a wrong match.

For the tool cache, canonicalise the arguments first. Otherwise the same call with keys in a different order is a miss. This is the safest layer, because it matches exact inputs to exact outputs and involves no judgement about meaning.

Two numbers tell me whether these caches are working: the hit rate of each cache on its own, and the cost per question with the cache cold against the cost on a repeated question.

When is this key not enough?

Reworded questions miss. The hit rate is low for free-typed questions and high for anything the UI suggests, such as example questions. Rewording is a job for the semantic cache.

Not every tool can be cached. Cache only tools that are read-only and deterministic. A tool with side effects is never cached. A tool that reads live data gets a short TTL or none.

A careless version empties the cache for nothing. In my system the corpus version is derived from the content hashes that ingestion already computes. A re-index that changes nothing does not empty the cache.

Redis goes down. The cache is optional for correctness. On an error or timeout, treat it as a miss and run the uncached path. Keep the client timeout short, so a slow Redis cannot slow every request. The cache is not optional for cost: with no cache, every request pays full price.

A right key can still hold a wrong value. A run that timed out, hit its step budget, or returned a partial answer is never stored. Neither is an abstention: "not found" is correct only until the missing document is ingested. The other cases where a hit is wrong get their own post.

What can you check in the next 20 minutes?

  • List everything your answer depends on besides the question. Each item is a key segment.
  • Put the corpus version in the key and a TTL on every write.
  • For each tool, ask: read-only and deterministic? If yes, cache it, and sort the arguments before hashing.
  • Put a short timeout on the Redis call. On error, treat it as a miss.
  • Log the hit rate for each cache separately. One combined number hides which layer is doing the work.

What is in your LLM cache key besides the question, and which part did you add only after it served a wrong answer?

Top comments (0)