I started this project thinking the hard part would be getting an LLM to understand legal documents.
It wasn't.
The harder problem was getting the system to remember what those documents had already established.
A lawyer working on an ongoing case isn't dealing with one question and one document. Every new filing has to fit with everything that came before it: previous arguments, witness statements, factual claims, earlier motions and relevant cases.
That made us ask a different question:
What happens when an AI agent doesn't just retrieve documents, but remembers the history of a case?
That question became the foundation of our Legal Hindsight Agent.
A case is more than a collection of documents
Our prototype uses a completely fictional case: Riverside Logistics v. Halden Manufacturing.
We created six documents for it, including a complaint, deposition summary, motion, witness statement and two fictional past cases. The case revolves around a delayed delivery.
The complaint says that Halden failed to deliver 5,000 units by March 15. The delay stopped Riverside's production line for twelve days and caused $240,000 in lost production. Other documents add details about the shutdown, communications and the contractual dispute.
Now imagine writing a new filing several weeks later. The new paragraph says:
"The delivery delay was only a minor inconvenience and resulted in negligible losses."
Read by itself, that paragraph sounds perfectly normal.
Read alongside the case history, it is a problem. The existing documents describe a twelve-day shutdown and $240,000 in losses.
The challenge isn't finding the words "delay" or "loss." The challenge is understanding the relationship between the new statement and what the case already says.
That's where search starts to look different from memory.
Our approach: Retain, Recall, Reflect
We built the memory layer using Hindsight.
The workflow has three important operations:
Retain -- store incoming information.
Recall -- retrieve memories relevant to a question or draft.
Reflect -- reason over those memories.
Our case documents are retained into one memory bank. When a lawyer submits a new paragraph, the system recalls the earlier information related to that paragraph and then reasons over it.
The backend is written in Python with FastAPI. The frontend uses Next.js and TypeScript. Hindsight handles the persistent case memory. We also use a plain Groq LLM as a comparison point without access to the case memory.
That comparison became important.
The memory shouldn't disappear behind the interface
One design decision we made early was to expose the memory.
The frontend doesn't just show an AI answer. It has two columns:
Without Memory and With Hindsight Memory
Below them is another panel: What the agent remembered.
That panel shows the memories that contributed to the answer.
This is important because otherwise "our AI has memory" is just a claim. Showing the remembered information makes the behavior inspectable.
What happens when the draft contradicts the case?
The contradictory sample is deliberately simple. The new draft minimizes the delay and losses.
The Hindsight-powered system recalls the earlier facts and identifies the inconsistency. In testing, it specifically connected the new language to the complaint, deposition and motion and provided a confidence level.
We also tested a consistent paragraph. The system did not flag it.
That was important. A system that simply raises warnings whenever it sees unfamiliar wording isn't useful. The goal is to identify meaningful contradictions while leaving consistent statements alone.
Why this changed our thinking
Initially, contradiction detection looked like the main feature.
After building it, we started seeing it differently.
Contradiction detection is really the visible proof that the memory is useful. The deeper capability is maintaining context across time. The lawyer shouldn't have to reconstruct the case from scratch every time a new document is written.
That's the problem we're actually trying to solve.
Search still has a role
This doesn't mean traditional retrieval is useless. Search is extremely useful when you know what you're looking for.
But a long-running agent needs more than retrieval. It needs continuity. A new statement can conflict with something that happened earlier, even when the wording is completely different.
That is where persistent memory becomes interesting.
What I learned
The biggest lesson from this project was:
Retrieval gives an agent information. Memory gives it continuity.
Our prototype is not a validated legal tool and does not replace legal judgment. Everything in the demonstration is synthetic.
But the architecture gave us a useful way to explore what happens when an agent can carry case context forward instead of starting over with every query.
And that changed the way I think about agent memory. The interesting question isn't always:
"What document should the model retrieve?"
Sometimes it is:
"What does the model need to remember before it answers?"
Built using Hindsight, an open-source persistent memory system for AI agents.
Hindsight GitHub repository: https://github.com/vectorize-io/hindsight
Hindsight documentation: https://hindsight.vectorize.io/
What is agent memory (Vectorize): https://vectorize.io/what-is-agent-memory
Our project's code: https://github.com/Priyanka5N6/legal-hindsight-agent




Top comments (0)