<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sudip Manna</title>
    <description>The latest articles on DEV Community by Sudip Manna (@sudip_manna_78303e6dd92dd).</description>
    <link>https://dev.to/sudip_manna_78303e6dd92dd</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147716%2F593b5c5d-deed-47a4-bd1c-f1aa8e8e5b3b.jpg</url>
      <title>DEV Community: Sudip Manna</title>
      <link>https://dev.to/sudip_manna_78303e6dd92dd</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sudip_manna_78303e6dd92dd"/>
    <language>en</language>
    <item>
      <title>I replayed nine months of claims to prove agent memory works</title>
      <dc:creator>Sudip Manna</dc:creator>
      <pubDate>Mon, 28 Sep 2026 17:32:01 +0000</pubDate>
      <link>https://dev.to/sudip_manna_78303e6dd92dd/i-replayed-nine-months-of-claims-to-prove-agent-memory-works-5731</link>
      <guid>https://dev.to/sudip_manna_78303e6dd92dd/i-replayed-nine-months-of-claims-to-prove-agent-memory-works-5731</guid>
      <description>&lt;p&gt;Everyone says their AI agent "has memory". Almost nobody shows what the memory actually changes. We wanted a number, so we built a test where the only thing that changes is memory, and replayed nine months of insurance claims through it.&lt;br&gt;
The answer: the same model with the same prompt caught 0 of 12 fraud claims in August and September without memory, and 11 of 12 with it, without flagging a single honest customer. The more interesting part is the shape of the curve, including the months where memory caught nothing at all.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Watch the 3-minute demo: &lt;a href="https://youtu.be/NUwMfwLhgdA" rel="noopener noreferrer"&gt;https://youtu.be/NUwMfwLhgdA&lt;/a&gt; · Try it in your browser: &lt;a href="https://rishighosal.github.io/claimlens/" rel="noopener noreferrer"&gt;https://rishighosal.github.io/claimlens/&lt;/a&gt; · Code: &lt;a href="https://github.com/rishighosal/claimlens" rel="noopener noreferrer"&gt;https://github.com/rishighosal/claimlens&lt;/a&gt;&lt;br&gt;
The project in one paragraph&lt;br&gt;
ClaimLens is a fraud-triage agent for an insurer's Special Investigation Unit (SIU). Organised fraud rings file claims that each look normal. The give-away is in the history: the same phone number under a different name, the same payee bank account, the same surveyor, the same story told almost word for word. ClaimLens uses Hindsight, an open-source agent memory system from Vectorize, to remember every claim and every investigator decision. For each new claim it asks memory targeted questions and scores the claim twice: once alone, once with what memory returned.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fled0h6ph470nzf3gaec2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fled0h6ph470nzf3gaec2.png" alt=" " width="800" height="381"&gt;&lt;/a&gt;&lt;br&gt;
Why a demo isn't proof&lt;br&gt;
A demo proves the happy path. You pick a claim, you know the answer, the agent gets it right. That tells you nothing about the two things that matter for an SIU:&lt;br&gt;
Does it catch fraud it has never been told about?&lt;br&gt;
Does it leave honest people alone?&lt;br&gt;
And there's a subtle trap with memory: leakage. If the memory bank already contains the investigator's final verdict on a claim, of course the agent "detects" it. So the test has to replay time.&lt;br&gt;
The replay: memory only ever knows the past&lt;br&gt;
Our dataset is 215 synthetic motor and health claims from Hyderabad, January to September, with four hidden fraud rings and deliberate decoys. The replay starts with an empty Hindsight bank and walks through the claims in date order:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;claims&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;                                   &lt;span class="c1"&gt;# sorted by intimation date
&lt;/span&gt;    &lt;span class="n"&gt;today&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;\&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;intimation\_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="c1"&gt;# outcomes decided up to today become memory, and not a day earlier
&lt;/span&gt;    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;verdict&lt;/span&gt;\&lt;span class="n"&gt;_queue&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;verdict&lt;/span&gt;\&lt;span class="n"&gt;_queue&lt;/span&gt;\&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;\&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;today&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        \&lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;vc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;verdict&lt;/span&gt;\&lt;span class="n"&gt;_queue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;pending&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ClaimMemory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;verdict&lt;/span&gt;\&lt;span class="nf"&gt;_item&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;vc&lt;/span&gt;\&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verdict&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;\&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claim\_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;scored&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;flush&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                              &lt;span class="c1"&gt;# memory must hold everything before this claim
&lt;/span&gt;        &lt;span class="n"&gt;mem&lt;/span&gt;\&lt;span class="n"&gt;_a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt;\&lt;span class="nf"&gt;_pair&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;pending&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ClaimMemory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;claim&lt;/span&gt;\&lt;span class="nf"&gt;_item&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;      &lt;span class="c1"&gt;# only now does the claim itself enter memory
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three rules make it honest:&lt;br&gt;
A claim is scored before it is retained, so it can never find itself.&lt;br&gt;
An investigator outcome enters memory on the day it was decided (&lt;code&gt;closed\_on&lt;/code&gt;), not on the day the claim arrived.&lt;br&gt;
The ground-truth label is used only to count results. The agent never sees it.&lt;br&gt;
We scored 70 claims: all 33 fraud claims plus 37 genuine ones. Everything else still flows into memory, exactly like a real claims desk.&lt;br&gt;
Same model, both arms, always&lt;br&gt;
The "without memory" arm is not a straw man. It is the same model (&lt;code&gt;openai/gpt-oss-120b&lt;/code&gt; on Groq), the same system prompt and the same JSON schema. The only difference is the HISTORY section that Hindsight fills in.&lt;br&gt;
We found one way this could silently break. On a free tier you hit rate limits, and our LLM client falls back to a smaller model. If one arm quietly ran on the backup, you'd be comparing two models, not memory versus no memory. So the pair scorer pins both arms to the same model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;score&lt;/span&gt;\&lt;span class="nf"&gt;_pair&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mem&lt;/span&gt;\&lt;span class="n"&gt;_a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ev&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;gather&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;assess&lt;/span&gt;\&lt;span class="n"&gt;_with&lt;/span&gt;\&lt;span class="nf"&gt;_memory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;assess&lt;/span&gt;\&lt;span class="nf"&gt;_stateless&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;mem&lt;/span&gt;\&lt;span class="n"&gt;_a&lt;/span&gt;\&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;base&lt;/span&gt;\&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="c1"&gt;# rate limits pushed one arm onto a backup model: re-score the other arm on the same one
&lt;/span&gt;        &lt;span class="n"&gt;pinned&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Investigator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;LLM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;mem&lt;/span&gt;\&lt;span class="n"&gt;_a&lt;/span&gt;\&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;fallback&lt;/span&gt;\&lt;span class="n"&gt;_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;strict&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;pinned&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;assess&lt;/span&gt;\&lt;span class="nf"&gt;_stateless&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;mem&lt;/span&gt;\&lt;span class="n"&gt;_a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;base&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In strict mode the evaluation also refuses to use the deterministic fallback scorer that the live app keeps as a safety net. If no model answers, it waits, then retries.&lt;br&gt;
The results&lt;br&gt;
A claim counts as "flagged" at a risk score of 61 or more ("refer to SIU").&lt;br&gt;
    Without memory  With Hindsight&lt;br&gt;
Fraud claims flagged, whole replay  0 of 33 14 of 33&lt;br&gt;
Fraud claims flagged, Aug–Sep 0 of 12 11 of 12&lt;br&gt;
Genuine claims wrongly flagged  0   0&lt;br&gt;
ROC AUC 0.46    0.79&lt;br&gt;
Fraud value flagged ₹0    ₹36,74,500&lt;br&gt;
And month by month, with memory: 0% (March–June), 43% (July), 100% (August), 86% (September). Without memory it was 0% every month.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu8yvwqm7ssoi5zbl29rj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu8yvwqm7ssoi5zbl29rj.png" alt=" " width="800" height="395"&gt;&lt;/a&gt;&lt;br&gt;
Read the curve, not the total&lt;br&gt;
The headline "42% recall" sounds weak until you look at when the misses happen.&lt;br&gt;
March to June, memory catches nothing, and that is correct. The rings' early claims were paid. Nobody had confirmed anything. "Same garage, same surveyor, previous claims approved" is not evidence, and we explicitly don't want an agent that accuses people on that.&lt;br&gt;
July, investigators repudiate the first ring claims. Those outcomes are retained into Hindsight as their own memories, tagged with the same phone, account, vehicle and surveyor identifiers as the claims. Recall jumps to 43%.&lt;br&gt;
August, 100%. September, 86%. Every new claim that shares a personal identifier with a confirmed-fraud claim now arrives with STRONG evidence attached.&lt;br&gt;
That is what "an agent that improves over time" should look like: flat until there's something to learn from, then a step up after each confirmed case. No retraining happened anywhere in this curve. The model weights never changed. Only the memory did.&lt;br&gt;
The per-ring view shows the same story: the first claims of every ring are missed, the later ones are caught.&lt;br&gt;
Ring    Claims  Without memory  With memory&lt;/p&gt;

&lt;p&gt;Garage + surveyor collusion 13  0   6&lt;br&gt;
Hospital admission ring 11  0   5&lt;br&gt;
Early-claim intermediary    7   0   3&lt;br&gt;
Recycled vehicle damage 2   0   0&lt;br&gt;
The honest misses&lt;br&gt;
The recycled vehicle ring scored 0 of 2. The same Hyundai Creta claimed the same front-left damage three times under different owners. In the replay both repeat claims scored 45, which means "standard review", not "refer to SIU". Nothing about that car had ever been confirmed as fraud, and our scoring rubric says one repeated identifier without a fraud finding is a review, not a referral. In the live app with the full history, the third claim scores 78 and is referred. We decided we'd rather under-refer than accuse.&lt;br&gt;
Early ring claims are invisible by design. If you need to catch the first claim of a brand-new ring, memory alone won't do it. That's a job for document checks and field investigation. Memory is what stops the tenth claim.&lt;br&gt;
What we'd tell anyone evaluating agent memory&lt;br&gt;
Run an ablation, not a demo. Same model, same prompt, memory on and off. Anything else measures something else.&lt;br&gt;
Replay time. If your memory can see the future, your numbers are fiction.&lt;br&gt;
Pin the model. Fallbacks and retries quietly change what you're measuring.&lt;br&gt;
Report false positives next to recall. For an SIU a wrongly accused customer costs more than a missed claim.&lt;br&gt;
Show the curve. A learning agent should get better after each outcome it's told about. If your curve is flat, your memory isn't doing anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The replay script is in the repo (&lt;code&gt;scripts/replay\_eval.py&lt;/code&gt;). It's also the pilot plan for a real insurer: run it on their closed claims and measure the same with/without difference on their data.
&lt;/h2&gt;

&lt;p&gt;ClaimLens is built on Hindsight (docs). If you're new to the idea, Vectorize has a good explainer on what agent memory is. Code: &lt;a href="https://github.com/rishighosal/claimlens" rel="noopener noreferrer"&gt;https://github.com/rishighosal/claimlens&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>python</category>
    </item>
  </channel>
</rss>
