DEV Community

Cover image for Why insurance fraud needs an AI that remembers
Subhankar Nandi
Subhankar Nandi

Posted on

Why insurance fraud needs an AI that remembers

Suresh, Priya and Sneha filed three car-insurance claims in Hyderabad between July and September. Different people, different cars, different parts of the city. Each one said the same thing: an unknown vehicle hit me at night and drove off. Each one, read on its own, looks like a routine claim that any insurer would pay.
Suresh's phone number is Sneha's phone number. Suresh wants his money paid into the same bank account as Priya. The same surveyor signed off all three. And investigators had already rejected Priya's and Sneha's claims as fraud.
None of that is written in Suresh's claim. It lives in the history. That gap, between what one claim says and what the history knows, is the problem we built ClaimLens to close.

Watch the 3-minute demo: https://youtu.be/NUwMfwLhgdA · Try it in your browser: https://rishighosal.github.io/claimlens/ · Code: https://github.com/rishighosal/claimlens

Why this problem matters
Insurance fraud isn't a niche issue. A November 2025 BCG and Medi Assist report estimates Indian insurers lose ₹8,000–10,000 crore a year to fraud, waste and abuse, around 8–10% of claim payouts. That money ends up in the premiums everyone pays.
Regulators have noticed. IRDAI's Insurance Fraud Monitoring Framework guidelines came into force on 1 April 2026, requiring every insurer to run a Fraud Monitoring Unit and keep fraud records. The people in those units, Special Investigation Units (SIUs), are experienced but outnumbered. They can't personally remember thousands of claims across branches and years.
And the most damaging fraud is organised: rings that reuse phone numbers, bank accounts, vehicles and "friendly" surveyors or doctors, while every single claim looks clean.
Why the obvious solutions don't work
Rules over-flag. "Flag every claim from this garage" catches the ring, and also delays the seven honest customers who got their cars repaired there. Genuine claimants pay the price.
Chatbots forget. A typical AI triage bot reads one claim, scores it and forgets it. It can be very smart and still blind, because the evidence isn't in the claim. In our tests, that kind of stateless AI scored every one of the three claims above 15 out of 100: fast-track, pay it.
What an investigator actually needs is a colleague who has read every claim the company ever handled, remembers what the team decided about each one, and can say "this phone number was on a claim we repudiated last week, here's the claim ID."
What we built
ClaimLens is that colleague. It's a triage workbench for an SIU, with Hindsight as its long-term memory.
For each new claim it:
Asks memory targeted questions: who else used this phone, this bank account, this vehicle? What did investigators decide about this surveyor before? Have we heard this exact story?
Scores the claim twice with the same AI model, once alone and once with what memory found, and shows both side by side so the investigator sees exactly what memory changed.
Learns from every decision. When an investigator records an outcome, it's written straight into memory:

# simplified from backend/claimlens/api.py
@app.post("/api/claims/{claim\_id}/decision")
async def decide(claim\_id: str, body: DecisionIn):
    c = \_need(claim\_id)
    verdict = {"decision": body.decision, "notes": body.notes.strip(),
               "investigator": body.investigator.strip() or "SIU desk",
               "closed\_on": date.today().isoformat()}
    await ctx.memory.retain\_verdict(c, verdict)     # one Hindsight retain: the agent just learned something
    return {"ok": True, "verdict": verdict}
Enter fullscreen mode Exit fullscreen mode

That one call is the difference between a tool and a colleague. In our demo, an investigator confirms fraud on Suresh's claim, opens the next claim from the same ring nine days later, and ClaimLens now cites the claim that was closed a minute ago.

Designing for the investigator, not the demo
We made a few product decisions that we think matter more than any model choice:
Show the "without memory" answer too. Putting the stateless score next to the memory score builds trust. The investigator sees that the AI didn't get smarter by magic; it got the history.
Every red flag cites a real claim ID. If the AI mentions a claim that memory didn't actually return, the citation is removed automatically. Investigators can click any claim ID to open it.
Show memory working, live. A side panel lists every question the agent asks Hindsight while it works. No black box.
It recommends; people decide. ClaimLens never rejects a claim. The outcome buttons (approve, clear, refer, confirm fraud) belong to the investigator.
Protect honest customers. An honest driver at the same garage as the fraud ring, with a different surveyor, stays at 20/100, fast-track, because memory also remembers the approved and cleared claims.

Does it work?
We replayed nine months of claims in date order, so memory only ever knew the past, and scored each claim with and without memory using the same model and prompt.
Without memory: 0 fraud claims caught all year.
With Hindsight: 11 of 12 fraud claims caught in August and September, after investigators had confirmed the first cases.
Honest customers wrongly flagged: 0.
Fraud value flagged: ₹36.7 lakh, against ₹0 without memory.

How an insurer could actually adopt this
We thought about this from day one:
Pilot in four weeks: run our replay script on the insurer's own closed claims and measure the same with/without difference on their data.
Shadow mode: score new claims alongside the existing process, without changing any decision.
Production: one memory bank per insurer. Hindsight is open source and can be self-hosted inside the insurer's own environment, which matters for personal data. Account numbers are already masked in claim files.
It also maps neatly onto what IRDAI asks for: red-flag indicators, a fraud incident record and a monitoring unit that learns from its own cases.
A small thing we're proud of
We wanted anyone to be able to try ClaimLens without installing Python or holding API keys. So we built a recorded demo site: the app replays a real run on Hindsight Cloud, every score and recall exactly as it happened, from a static page on GitHub Pages. No server, no keys exposed, same interface.
What we learned
The data is the product. We spent real effort making the synthetic claims realistic (Hyderabad areas, police stations, IFSC-style accounts) and adding decoys to punish lazy rules.
Trust comes from showing your work. Side-by-side scores, cited claim IDs and a live memory panel did more for credibility than any accuracy number.
The limitation is real: the first claims of a brand-new ring get through, because there's nothing to remember yet. Memory stops the tenth claim, not the first.

Fraud rings hide in history. ClaimLens remembers it.

Built on Hindsight by Vectorize (docs). New to the idea? Read What is agent memory? Sources: BCG × Medi Assist report on insurance fraud, waste and abuse (Nov 2025); IRDAI Insurance Fraud Monitoring Framework Guidelines, 2025. Code: https://github.com/rishighosal/claimlens

Top comments (0)