Building an AI that flags fraud is easy. Make the threshold low enough and it will flag everything. The hard part, the part that decides whether anyone should ever trust it, is not accusing honest people.
This is the story of how our fraud-triage agent went from scoring an innocent driver 80 out of 100 to 20 out of 100, without making it any worse at catching the actual fraud ring he happened to share a garage with.
Watch the 3-minute demo: https://youtu.be/NUwMfwLhgdA · Try it in your browser: https://rishighosal.github.io/claimlens/ · Code: https://github.com/rishighosal/claimlens
The setup
ClaimLens is a triage agent for an insurance Special Investigation Unit (SIU). It uses Hindsight as long-term memory: every past claim and every investigator decision is retained, and every new claim is checked against all of it. The same model scores each claim twice, once alone and once with memory, so you can see exactly what memory changed.
Our test data hides four fraud rings in 215 claims from Hyderabad. One ring runs through a garage in Moosapet, Sri Balaji Auto Works, with the same surveyor signing off staged "unknown vehicle hit me from behind" accidents.
Here's the catch we built in on purpose: Sri Balaji Auto Works has 20 claims in the data. 13 are the ring. 7 are honest customers who simply got their cars fixed there.
The claim that broke our first version
Claim #10614. A driver's car was rear-ended, the other driver's registration is on record, it's being repaired at Sri Balaji, and it was inspected by a different surveyor, T. Lavanya.
Our early version handed the model the history as one undifferentiated list. When a rate limit pushed the agent onto a smaller backup model, it looked at that list and saw: same garage as a confirmed fraud ring, similar rear-end story, similar damage. It scored the claim 80: refer to SIU. Our simple rule-based safety net did exactly the same.
That's a disaster for a real insurer. An honest customer gets their claim delayed, maybe investigated, because of where they took their car. Worse, it's exactly what a dumb rule does ("flag every Sri Balaji claim"), and we were supposed to be better than a rule.
The model wasn't being stupid. We had handed it a pile of "related" past claims with no sense of which relations matter.
Fix 1: grade the evidence before the model sees it
Fraud rings are combinations. A shared phone number between two different claimants is a big deal. A shared garage is not. So before the history reaches the model, every linked past claim gets a grade, computed in plain code:
# simplified from backend/claimlens/agent.py
def grade(ln: Link) -> tuple[str, str]:
"""Label a link STRONG or WEAK, so the model weighs combinations, not topics."""
personal = sorted(ln.reasons & STRONG_PERSONAL) # phone, bank account, vehicle
fraud = _is_fraud(ln)
if personal:
return "STRONG", "a personal identifier appears on a different claim/policy" + (
" that SIU confirmed as fraud" if fraud else "")
if fraud and ln.reasons & {"surveyor", "doctor"}:
return "STRONG", "same surveyor/doctor as a claim SIU confirmed as fraud"
return "WEAK", "shares only context (garage, story); no personal identifier in common"
Each history entry in the prompt now carries its label, and the prompt gives an explicit rubric:
- no STRONG links → normally fast-track
- one STRONG link that isn't confirmed fraud → standard review
- one STRONG link to confirmed fraud, or two or more STRONG links → refer to SIU
The model still does the reasoning and writes the explanation. But it can no longer treat "same garage" like "same bank account".
Fix 2: memory remembers the good outcomes too
This is where Hindsight earns its place. We don't only retain claims; we retain investigator outcomes as their own memories, dated and tagged with the same identifiers:
- "SIU CONFIRMED FRAUD on CLM-2026-10572…"
- "Approved and paid; surveyor assessment accepted…"
- "Cleared after field verification…"
So when #10614 asks memory "what did investigators decide about claims at this garage?", the answer isn't only "fraud". It's also "six other claims here, with other surveyors, were approved without issue". Memory gives the agent the mitigating context a rule never has.
We also wrote the SIU's rules of evidence into the Hindsight bank itself, as directives that reflect must follow:
- A pattern in history is a lead for investigation, never proof.
- High claim volume at a garage or hospital is not suspicious by itself.
- If investigators previously cleared a similar claim, say so and lower suspicion.
Fix 3: the model can't cite what memory didn't return
An agent that invents evidence is worse than one with none. Every red flag the model writes must cite claim IDs, and a validator strips any ID that memory didn't actually return:
kept = [i for i in ids if i in allowed_ids] # only claims Hindsight actually recalled
removed += len(ids) - len(kept)
...
"band": band_for(score), # the band is always derived from the score, never trusted blindly
The UI shows the investigator every recall Hindsight made, live, so nothing is hidden.
The result
Same model (gpt-oss-120b), same claim, from our recorded live run:
| Claim | Without memory | With memory |
|---|---|---|
| #10614: honest driver, ring's garage, different surveyor | 12, fast-track | 20, fast-track |
| #10595: ring claim, shared phone + bank account | 15, fast-track | 88, refer to SIU |
The ring claim right next to it still gets caught, because it shares a phone number and a payee bank account with claims investigators had already repudiated:
We also planted a second decoy: a busy, completely clean car dealership with lots of claims. A volume rule flags it. ClaimLens doesn't.
Across the full nine-month replay (70 claims scored in date order, memory only ever seeing the past), memory took fraud caught from 0 to 14 of 33, and 11 of 12 in August–September. Genuine claims wrongly flagged: zero, with and without memory.
What we learned
- "Related" is not "suspicious". Semantic similarity will happily connect every rear-end collision to every other one. Decide in code which links count.
- Store good outcomes, not just bad ones. "Approved", "cleared" and "surveyor accepted" are the memories that protect honest customers.
-
Put the rules where the reasoning happens. Hindsight directives and disposition (we used skepticism 4, literalism 4, empathy 2) shape every
reflectanswer, not just one prompt. - Make citations checkable. If the agent can't point to a real claim ID, the point gets removed.
- The agent recommends; people decide. ClaimLens never repudiates anything. It tells an investigator where to look and why.
The limitation we accept: because the agent is conservative, the first claims of a new ring get through until investigators confirm one. We think that's the right trade. A false accusation is a real person's money and reputation.
ClaimLens is built on Hindsight (docs). New to agent memory? Start with What is agent memory? Code: https://github.com/rishighosal/claimlens


Top comments (0)