How We Built Vera’s Memory With Hindsight
We called it Vera.
Our team, Certified Failures, went through quite a few names before settling on it. We wanted something that matched our style and the character of the project—simple, memorable, and aesthetic rather than overly technical. We chose Vera because of its meaning—faith and truth—and because the name just felt right for the project.
That idea of remembering became the heart of Vera.
The Problem
User feedback comes from everywhere: support tickets, in-app feedback, reviews, and imported datasets. The difficult part isn't collecting those messages. It's understanding how they change over time.
A complaint that looks like a new problem today might actually be something users have been reporting for months.
We wanted Vera to answer questions like:
“How has the PDF upload problem changed over time?”
or:
“Did user complaints improve after a product update?”
To answer questions like these, Vera needed more than the latest batch of feedback. It needed memory.
Giving Vera a Memory
We used Hindsight as Vera's long-term memory layer.
The core flow is simple:
User Feedback
↓
Hindsight RETAIN
↓
Long-Term Memory
↓
User Question
↓
Hindsight RECALL / REFLECT
↓
Historical Evidence
↓
Insight
When feedback enters Vera, we retain the original text along with information such as its product area, source, and timestamp.
A simplified version of our implementation is:
result = await self.client.aretain(
bank_id=self.bank_id,
content=text,
context=context,
timestamp=date_str,
document_id=document_id,
metadata=metadata,
tags=tags
)
The timestamp is important because feedback isn't just text. When something happened can change what that feedback means.
When a user asks a question, Vera decides how deeply it needs to search its memory. For questions involving changes over time or before-and-after comparisons, it can use Hindsight's areflect():
reflect_response = await self.client.areflect(
bank_id=self.bank_id,
query=question,
budget=budget,
max_tokens=4096,
)
The retrieved memories then provide the historical context Vera uses to construct its response.
You can learn more through the link
Hindsight GitHub: github.com/vectorize-io/hindsight.
The Before-and-After Moment
The clearest difference appeared when we tested Vera with our four-month feedback dataset.
With only recent April data, Vera identified PDF uploads as a serious problem and interpreted the failures as a recent regression.
After retaining the historical January–April feedback, Vera connected the timeline:
January: users reported slow PDF uploads.
February: stalled upload progress became more common.
March: users began reporting 504 timeout errors after a release.
April: PDF uploads had become a critical blocker.
The important insight wasn't simply that PDF uploads were failing.
The problem had been developing for months before it became a visible failure.
Vera also surfaced a contrasting result: login complaints dropped after the SSO rollout. The same historical memory could therefore reveal both an unresolved problem and a successful product change.
That was the behavior we wanted from an agent with memory—not just finding similar feedback, but connecting events across time.
What We Learned
Memory needs to be part of the architecture. Simply adding previous messages to a prompt isn't enough for long-term understanding.
Time matters. For feedback analysis, knowing when something happened can completely change its meaning.
Evidence matters. Vera can connect its conclusions to dates and original feedback instead of presenting unexplained summaries.
Most importantly, historical context changes the questions we can ask. Without memory, we can ask, “What's broken?” With memory, we can ask, “How did it become broken, and did anything we changed actually fix it?”
That was our biggest takeaway from building Vera:
User feedback becomes much more useful when an agent can remember its history.
With love,
— Certified Failures
Top comments (0)