DEV Community

Karunya Muddana
Karunya Muddana

Posted on AI-assisted

One Hindsight Setting Decided Whether My Agent Could Learn Judges

Litigators know things that no case file records.

Judge Murthy lets the other side have one adjournment. The second time, he says “last opportunity.” The third time, he closes the evidence and puts costs on them.

Nobody writes that rule down. It lives in a lawyer’s head, spread across a dozen hearings in three unrelated cases, and it’s the kind of thing a junior learns only after getting burned.

I wanted my agent to learn it the same way, from the notes, without anyone typing it in.

For a long time it couldn’t, and the reason turned out to be a single field on the items I was storing.

What Tareekh Is

Tareekh (Urdu for “date”, and in Indian courts, the next hearing date) is a memory for a litigator’s practice.

The lawyer uploads what they already produce: phone photos of pocket-diary pages, certified copies of court order sheets, deeds, legal notices, typed notes.

Tareekh reads each upload, splits it into one entry per hearing, files each entry under the right case and stores it in Hindsight.

When the lawyer asks something, an agent answers from that memory and cites the page it came from.

It remembers. It does not give legal advice.

That line is written into the memory bank itself as a directive, and I’ll come back to why that matters.

The Architecture

The stack is a FastAPI backend, a Next.js front end, Gemini on Vertex AI for OCR and reasoning, and Hindsight running locally in Docker.

There’s one Hindsight bank per lawyer. Every case lives in that one bank, and that decision is the whole story here, because a judge’s habits only show up when you look across cases.

The ingestion pipeline works like this:

  1. The lawyer uploads diary photos, order sheets, deeds, or notes.
  2. The documents go through OCR, segmentation, repair, and review.
  3. Each hearing becomes a separate memory entry.
  4. The entries are retained in Hindsight.
  5. The agent uses recall and reflect to answer questions, with citations pointing back to the source documents.

The case registry, backed by SQLite, helps resolve cases and identify the relevant judge and opposing counsel.

Figure 1: Tareekh’s architecture. Uploads go through ingestion into one Hindsight bank per lawyer; the agent reads it back with recall and reflect.

The Pattern I Wanted It to Find

The practice I test against belongs to one lawyer, Adv. Aditya Varma, and his junior Divya: five cases in front of two judges.

Three of those cases sit before Murthy sir.

His behavior is scattered across them.

Greenfield School Recovery Suit

Opposing counsel was “unwell” at one hearing.

At the next, he was “unwell” again and the judge said “LAST opportunity.”

On 22 June 2026, a third request was rejected, cross-examination was closed as nil, and Rs 3,000 costs were imposed.

Seabreeze Specific-Performance Suit

The other side dumped 14 documents on the table without an index on 9 July 2025.

The judge refused them and put Rs 2,000 costs on the company.

Seabreeze Injunction

He threw back a memo that hadn’t been served on us first.

No single case shows the rule.

Put together, the rule is obvious:

  • Second request: a warning.
  • Third request: costs.
  • Irregular

Figure 2: Aditya’s diary note from 22 June 2026, found in the knowledge graph by searching “Murthy sir adjournment costs”. The pattern is only visible when notes from different cases sit side by side.

Hindsight Observations, and Why Mine Were Stuck Inside One Case

Hindsight does more than store chunks and run vector search.

After you retain items, it runs consolidation in the background and forms observations, which are durable beliefs built up from many memories.

You can steer what it looks for with an observations_mission on the bank.

Mine reads like this:

OBSERVATIONS_MISSION = (
"Build durable knowledge about recurring behaviour: how each judge conducts hearings (what they ask for first, "
"how they treat repeated adjournments, costs, undertakings, synopses, mediation); how each opposing counsel "
"operates (reasons they give for adjournments, how often, which tactics); and what the lawyer himself routinely "
"does. Count occurrences and name the cases and dates behind each pattern. Ignore one-off events."
)

Every hearing I store carries tags:

case:C5
judge:J1
counsel:OC4
client:CL3
type:note
author:aditya

Tags are how I filter recall later (“only this case”, “only this judge”), so each item gets all of them.

Here’s what I missed.

By default, Hindsight groups consolidation by the item’s full tag set.

A Greenfield note tagged case:C5 judge:J1 counsel:OC4 and a Seabreeze note tagged case:C2 judge:J1 counsel:OC2 have different tag sets.

So they land in different buckets.

In practice that meant one scope per case, and a pattern that only exists across cases had nowhere to form.

The Default Behavior

The observations that could form were accurate and useless, along the lines of:

“In the Greenfield suit, costs were imposed after repeated adjournments.”

True, and already known, because it’s what the note says.

The Fix: observation_scopes

Hindsight lets each retained item say which scopes it should consolidate under.

I added one line to the function that builds every memory:

memory.py, build_item() (trimmed)

return {
"content": header + "\n" + entry["text"],
"context": "lawyer's hearing note",
"timestamp": f"{date}T10:30:00+05:30",
"tags": registry.case_tags(cid) + [
f"type:{entry.get('doc_type', 'other')}",
f"author:{author}"
],
"metadata": {
"source_file": source_file,
"upload_id": upload_id,
"case_id": cid,
"hearing_date": date
},
"document_id": f"{upload_id}:{cid}:{date}",
# Consolidate observations per judge, per counsel and per case,
# so cross-case patterns can form.
"observation_scopes": [
[f"judge:{c['judge_id']}"],
[f"counsel:{c['opposing_counsel_id']}"],
[f"case:{cid}"]
],
}

Now each hearing feeds three buckets.

The judge bucket sees every hearing in front of that judge, whatever the case.

The counsel bucket does the same for opposing counsel.

The case bucket still gives the per-case summary I had before.

One advocate’s “illness” is a good example: the diary note from the day of his second sick-day adjournment says he was seen in the bar association canteen at 1.30 pm.

That belongs in a picture of the counsel, not of one case.

That was the whole change.

It’s also the reason I now read the retain API reference before I design tags, and not after.

Mental Models Turn Observations Into Something the Agent Can Use

Observations are raw material.

What the agent actually reads before answering are Hindsight mental models: named, query-defined summaries that refresh themselves after consolidation.

I create one per judge, one per opposing counsel, one for open commitments and one for how the lawyer works.

Here’s the judge mental model configuration:

trig = {
"refresh_after_consolidation": True,
"mode": "delta"
}
for j in db.rows("SELECT id, name FROM judges"):
specs.append({
"id": memory.mental_model_id("judge", j["id"]),
"name": f"Judge: {j['name']}",
"tags": [f"judge:{j['id']}"],
"source_query": (
f"How does {j['name']} conduct hearings? "
"Habits, what they ask for first, how they "
"treat adjournment requests (especially repeated ones), "
"costs, undertakings, synopses and mediation. "
"Name the cases and dates behind each pattern."
),
"trigger": trig,
"max_tokens": 700
})

There’s a choice in there I care about.

Onboarding asks for the lawyer’s cases, judges, courts and opposing counsel, but it deliberately drops anything about temperament or tactics.

The seed data describes Murthy sir as “impatient, hates anything irregular.”

None of that goes into the bank.

If the “Judge: Murthy” model says he’s strict about adjournments, it’s because the notes said so, not because I told it.

Otherwise “the agent learns over time” would just be me reading my own config back.

How the Agent Uses These Models

When a question comes in, a context builder written in plain code (no LLM call) finds the cases mentioned, looks up their judge and counsel, and pulls those mental models into the prompt.

The tool loop then aims recall and reflect with tag filters such as:

judge_id="J1"

Before and After

The question:

“What does Murthy sir do when the other side keeps asking for time?”

Before the scopes change, all the agent had to work with was individual hearings.

It could find the Greenfield costs order, which is a correct fact, but that doesn’t answer the question.

The lawyer is really asking what to expect next week.

After the change, the answer my test set expects, and the one I now check for, is:

He says “last opportunity” on the second request and imposes costs on the third (Greenfield: Rs 3,000 on 22.06.2026, cross closed) [1]. He also refuses irregular filings (Seabreeze SP 09.07.2025: Rs 2,000 costs [2]; memo without copy rejected 11.03.2026 [3]).

Each [n] resolves to the note or order sheet behind it, so the lawyer can tap through to the diary photo.

One Bank Per Lawyer Also Helps Find Contradictions

Keeping every case in one bank pays off in a second way.

A party, Srinivas, deposed in one suit that he and his co-owner were “EQUAL owners, 50-50”, then claimed 70% in a partition suit before a different judge.

Two cases, two courts, one person.

A search for “what share does Srinivas claim” pulls both statements up together.

With a bank per case, you’d have to already suspect the contradiction to go looking for it.

Figure 3: Searching “what share does Srinivas claim” returns his 70% claim in the partition suit next to his “equal owners” deposition from a different case.

What I’d Tell Someone Starting Out

  1. Design scopes before tags.

Tags answer “what can I filter by?”

Scopes answer “what should be learned together?”

They’re different questions, and the default treats them as one.

  1. Don’t seed what you want learned.

If you pre-fill the judge’s temperament, you’ll never know whether the memory works.

I left it empty on purpose and treated the first correct judge model as the real test.

  1. Ask for counts and citations in the observations mission.

“Count occurrences and name the cases and dates” changed the observations from vague adjectives (“the judge is strict”) into claims I could verify.

  1. Consolidation takes time, and your UI should say so.

Mental models refresh after consolidation, and with Gemini 3.x thinking before it answers, I had to raise Hindsight’s LLM timeout to 300 seconds before consolidation stopped timing out.

A judge model built from three hearings is thin.

From ten, it’s genuinely useful.

I don’t have a better answer for the cold start than being honest about it in the interface.

  1. Put the guardrails in the bank, not only in the prompt.

“Cite sources”, “No legal advice” and “Admit gaps” are Hindsight directives on the bank, so they apply to reflect as well as to my own agent prompt.

A tool that says “the judge will impose costs” as a prediction would be dangerous.

One that says “he imposed costs the last two times, here are the orders” is useful.

The Takeaway

If you’re building anything where the valuable knowledge is a pattern across records rather than a fact inside one, look hard at how your memory layer groups things before it summarizes them.

The Hindsight documentation covers observation_scopes and mental models, and Vectorize has a good primer on what agent memory actually is if you’re deciding whether you need more than a vector store.

For me, one field was the difference between an agent that repeats my notes and one that knows the judge.

Top comments (0)