Everyone says their AI agent "has memory". Almost nobody shows what the memory actually changes. We wanted a number, so we built a test where the only thing that changes is memory, and replayed nine months of insurance claims through it.
The answer: the same model with the same prompt caught 0 of 12 fraud claims in August and September without memory, and 11 of 12 with it, without flagging a single honest customer. The more interesting part is the shape of the curve, including the months where memory caught nothing at all.
Watch the 3-minute demo: https://youtu.be/NUwMfwLhgdA · Try it in your browser: https://rishighosal.github.io/claimlens/ · Code: https://github.com/rishighosal/claimlens
The project in one paragraph
ClaimLens is a fraud-triage agent for an insurer's Special Investigation Unit (SIU). Organised fraud rings file claims that each look normal. The give-away is in the history: the same phone number under a different name, the same payee bank account, the same surveyor, the same story told almost word for word. ClaimLens uses Hindsight, an open-source agent memory system from Vectorize, to remember every claim and every investigator decision. For each new claim it asks memory targeted questions and scores the claim twice: once alone, once with what memory returned.

Why a demo isn't proof
A demo proves the happy path. You pick a claim, you know the answer, the agent gets it right. That tells you nothing about the two things that matter for an SIU:
Does it catch fraud it has never been told about?
Does it leave honest people alone?
And there's a subtle trap with memory: leakage. If the memory bank already contains the investigator's final verdict on a claim, of course the agent "detects" it. So the test has to replay time.
The replay: memory only ever knows the past
Our dataset is 215 synthetic motor and health claims from Hyderabad, January to September, with four hidden fraud rings and deliberate decoys. The replay starts with an empty Hindsight bank and walks through the claims in date order:
for c in claims: # sorted by intimation date
today = c\["intimation\_date"]
# outcomes decided up to today become memory, and not a day earlier
while verdict\_queue and verdict\_queue\[0]\[0] <= today:
\_, vc = verdict\_queue.pop(0)
pending.append(ClaimMemory.verdict\_item(vc, vc\["verdict"]))
if c\["claim\_id"] in scored:
await flush() # memory must hold everything before this claim
mem\_a, ev, base = await score\_pair(agent, c)
pending.append(ClaimMemory.claim\_item(c)) # only now does the claim itself enter memory
Three rules make it honest:
A claim is scored before it is retained, so it can never find itself.
An investigator outcome enters memory on the day it was decided (closed\_on), not on the day the claim arrived.
The ground-truth label is used only to count results. The agent never sees it.
We scored 70 claims: all 33 fraud claims plus 37 genuine ones. Everything else still flows into memory, exactly like a real claims desk.
Same model, both arms, always
The "without memory" arm is not a straw man. It is the same model (openai/gpt-oss-120b on Groq), the same system prompt and the same JSON schema. The only difference is the HISTORY section that Hindsight fills in.
We found one way this could silently break. On a free tier you hit rate limits, and our LLM client falls back to a smaller model. If one arm quietly ran on the backup, you'd be comparing two models, not memory versus no memory. So the pair scorer pins both arms to the same model:
async def score\_pair(agent, c):
(mem\_a, ev), base = await asyncio.gather(agent.assess\_with\_memory(c), agent.assess\_stateless(c))
if mem\_a\["model"] != base\["model"]:
# rate limits pushed one arm onto a backup model: re-score the other arm on the same one
pinned = Investigator(agent.memory, LLM(model=mem\_a\["model"], fallback\_model=""), strict=True)
base = await pinned.assess\_stateless(c)
return mem\_a, ev, base
In strict mode the evaluation also refuses to use the deterministic fallback scorer that the live app keeps as a safety net. If no model answers, it waits, then retries.
The results
A claim counts as "flagged" at a risk score of 61 or more ("refer to SIU").
Without memory With Hindsight
Fraud claims flagged, whole replay 0 of 33 14 of 33
Fraud claims flagged, Aug–Sep 0 of 12 11 of 12
Genuine claims wrongly flagged 0 0
ROC AUC 0.46 0.79
Fraud value flagged ₹0 ₹36,74,500
And month by month, with memory: 0% (March–June), 43% (July), 100% (August), 86% (September). Without memory it was 0% every month.

Read the curve, not the total
The headline "42% recall" sounds weak until you look at when the misses happen.
March to June, memory catches nothing, and that is correct. The rings' early claims were paid. Nobody had confirmed anything. "Same garage, same surveyor, previous claims approved" is not evidence, and we explicitly don't want an agent that accuses people on that.
July, investigators repudiate the first ring claims. Those outcomes are retained into Hindsight as their own memories, tagged with the same phone, account, vehicle and surveyor identifiers as the claims. Recall jumps to 43%.
August, 100%. September, 86%. Every new claim that shares a personal identifier with a confirmed-fraud claim now arrives with STRONG evidence attached.
That is what "an agent that improves over time" should look like: flat until there's something to learn from, then a step up after each confirmed case. No retraining happened anywhere in this curve. The model weights never changed. Only the memory did.
The per-ring view shows the same story: the first claims of every ring are missed, the later ones are caught.
Ring Claims Without memory With memory
Garage + surveyor collusion 13 0 6
Hospital admission ring 11 0 5
Early-claim intermediary 7 0 3
Recycled vehicle damage 2 0 0
The honest misses
The recycled vehicle ring scored 0 of 2. The same Hyundai Creta claimed the same front-left damage three times under different owners. In the replay both repeat claims scored 45, which means "standard review", not "refer to SIU". Nothing about that car had ever been confirmed as fraud, and our scoring rubric says one repeated identifier without a fraud finding is a review, not a referral. In the live app with the full history, the third claim scores 78 and is referred. We decided we'd rather under-refer than accuse.
Early ring claims are invisible by design. If you need to catch the first claim of a brand-new ring, memory alone won't do it. That's a job for document checks and field investigation. Memory is what stops the tenth claim.
What we'd tell anyone evaluating agent memory
Run an ablation, not a demo. Same model, same prompt, memory on and off. Anything else measures something else.
Replay time. If your memory can see the future, your numbers are fiction.
Pin the model. Fallbacks and retries quietly change what you're measuring.
Report false positives next to recall. For an SIU a wrongly accused customer costs more than a missed claim.
Show the curve. A learning agent should get better after each outcome it's told about. If your curve is flat, your memory isn't doing anything.
The replay script is in the repo (scripts/replay\_eval.py). It's also the pilot plan for a real insurer: run it on their closed claims and measure the same with/without difference on their data.
ClaimLens is built on Hindsight (docs). If you're new to the idea, Vectorize has a good explainer on what agent memory is. Code: https://github.com/rishighosal/claimlens
Top comments (0)