Most agent-memory tutorials go like this: embed the conversation, store it, run a similarity search on the next turn. That works for a chatbot remembering your favourite colour. It fails completely for the problem we worked on, because fraud rings don't share topics. They share identifiers.
This is a practical walkthrough of how we wired Hindsight into ClaimLens, an insurance fraud-triage agent, and the design decisions that took it from "finds similar claims" to "finds the ring".
Watch the 3-minute demo: https://youtu.be/NUwMfwLhgdA · Try it in your browser: https://rishighosal.github.io/claimlens/ · Code: https://github.com/rishighosal/claimlens
The problem memory has to solve
An insurance Special Investigation Unit (SIU) receives a claim: a hit-and-run at night, rear damage, no police report. Alone, it's routine. But in our data, the claimant's phone number already appeared on a claim investigators repudiated, the payee bank account appeared on another, and the same surveyor signed off all of them.
A stateless LLM scores that claim 15/100, fast-track. With Hindsight, the same model with the same prompt scores it 88/100, refer to SIU, and cites every claim it relied on.
Getting there took four Hindsight features working together: a configured bank, structured retains, scoped recalls and reflect with directives.
1. Configure the bank like a team, not a database
A Hindsight bank has missions that guide how facts are extracted and how the bank reasons. We set it up once:
await client.acreate_bank(
bank_id,
retain_mission=RETAIN_MISSION, # "capture claim ID, phone, payee account, vehicle, surveyor... exactly as written"
reflect_mission=REFLECT_MISSION, # "connect a new claim to everything the unit has seen; cite claim IDs"
enable_observations=True,
background="Special Investigation Unit, motor and health claims, Hyderabad region.",
)
await client.aupdate_bank_config(
bank_id,
disposition_skepticism=4, # investigators should be sceptical...
disposition_literalism=4, # ...precise about identifiers...
disposition_empathy=2, # ...but not dismissive of genuine claimants
)
for name, content, priority in DIRECTIVES:
await client.acreate_directive(bank_id, name=name, content=content, priority=priority)
The four directives are the SIU's rules of evidence: a pattern is a lead, never proof; cite claim IDs; volume is not fraud; respect cleared outcomes. They apply to every reflect call, so the investigator briefing and the "ask memory" box follow them without us repeating them in prompts.
2. Design writes for the reads you need
This was the most important decision. Each claim is retained with its identifiers passed twice: as typed entities (so Hindsight's graph can connect them) and as tags (so we can filter on them exactly).
def claim_item(claim):
ents = C.entities(claim) # phone:, account:, vehicle:, surveyor:, garage:, hospital:, doctor:, agent:...
return {
"content": C.render(claim), # a readable claim file
"timestamp": C.as_datetime(claim["intimation_date"]),
"document_id": claim["claim_id"],
"metadata": {"claim_id": claim["claim_id"], "kind": "claim"},
"entities": [{"text": e.label, "type": e.kind} for e in ents],
"tags": ["claim", f"line:{claim['line']}", *[e.tag for e in ents]],
"update_mode": "replace",
}
Investigator outcomes are retained as separate documents, dated on the day they were decided, carrying the same identifier tags plus a verdict tag:
"document_id": f"verdict:{claim['claim_id']}",
"timestamp": C.as_datetime(verdict["closed_on"]),
"tags": ["verdict", f"decision:{verdict['decision']}", *[e.tag for e in ents]],
Why separate? Because "what someone claimed" and "what turned out to be true" are different facts with different dates. Keeping them apart lets the agent ask about each one directly, and it's what makes the agent learn: one aretain_batch call when an investigator clicks Confirm fraud, and the next claim from that ring is judged with it. No retraining, no re-indexing.
3. Many narrow recalls beat one big one
Our first version did one semantic recall with the claim text. It linked every rear-end collision to every other rear-end collision. Useless.
Now the agent plans about 15–17 probes per claim. For every identifier it asks two questions: who else has this? and what did SIU decide about it?
for e in C.entities(claim):
probes.append(Probe(label=f"Who else has {e.label}?",
query=f"Past claims and investigator outcomes involving {e.label}",
tags=[e.tag]))
probes.append(Probe(label=f"What did SIU decide about {e.label}?",
query=f"Investigator outcome, repudiation or approval for claims involving {e.label}",
tags=["verdict", e.tag], tags_match="all_strict"))
The all_strict verdict probe was a late fix. With one recall per identifier, a recent verdict could get crowded out by many older claim facts about the same busy garage. A dedicated probe that requires both the verdict tag and the identifier tag guarantees outcomes are always seen.
Then come the probes tags can't do:
-
Narrative: "Have we heard this story before?" Hindsight runs semantic, keyword, graph and temporal retrieval in parallel and reranks. We added
min_scores={"reranker": 0.55}, because semantic search always returns something, and without a floor every claim had a "similar story". -
Pattern: recall over consolidated observations (
types=["observation"]), the patterns Hindsight distils as memories accumulate. -
Timeline: what happened around this date, using
query_timestampand atemporal_window.
Every recall appears live in the UI, so an investigator can see exactly what the agent asked:
4. Map results back to claims, then let the model weigh them
Each recalled fact is mapped back to a claim through its document_id and metadata, grouped per past claim, and graded STRONG (shared phone, bank account or vehicle; or the same surveyor/doctor as confirmed fraud) or WEAK (only a garage or a similar story). That graph is what the investigator sees:
5. reflect and a mental model nobody wrote
Two features we didn't expect to love:
-
reflectwrites the investigator briefing and answers free-text questions like "What do the Lifeline Multispeciality claims have in common?" It reasons over the whole bank and follows the directives, including citing claim IDs. -
A mental model called SIU fraud playbook, created with
acreate_mental_model(..., trigger={"refresh_after_consolidation": True}). As outcomes accumulate, Hindsight rewrites it into a readable summary of confirmed patterns and of busy entities that were cleared.
Did it work?
We replayed nine months of claims in date order through a fresh bank, scoring each claim with and without memory before retaining it. Same model (gpt-oss-120b), same prompt. In August and September: 0 of 12 fraud claims caught without memory, 11 of 12 with Hindsight, 0 honest claims flagged.
Lessons, in one list
- Pass identifiers as entities and tags. Entities build the graph; tags make exact lookups trivial.
-
Outcomes are memories. Give them their own
document_idand the date they were decided. - Ask narrow questions. One recall per identifier explains why two claims are linked.
- Floor your semantic recall. A reranker threshold stops "similar story" from matching everything.
-
Put the rules in the bank. Directives and disposition travel with every
reflect.
Our honest limitation: the first claims of a brand-new ring get through, because there is nothing to remember yet. Memory is what stops the tenth claim, not the first.
Built on Hindsight. Docs: https://hindsight.vectorize.io/ · Background: What is agent memory? · Code: https://github.com/rishighosal/claimlens




Top comments (0)