I built an incident-response agent on top of Hindsight, an agent memory system. This post is about why I chose it over the obvious alternative of a vector store, how I used its two retrieval calls, and, importantly, what I never tested.
I'll say the last part up front: I did not benchmark Hindsight against plain vector search. So this isn't a "memory beats vector search" post. It's a description of what the architecture gave me and where I saw it help.
The Choice:
The project takes an incident description and returns a diagnosis, using 104 real postmortems from Slack, Cloudflare, GitHub, AWS, Datadog, CircleCI, and LaunchDarkly. The point is to warn against "trap actions," meaning fixes that made past outages worse.
A plain vector store would return the postmortem paragraphs nearest to the query. That's useful, but I wanted the agent to reason across incidents, for example noticing that several unrelated outages share a pattern where rollback made things worse. That pattern doesn't live in any single paragraph.
Hindsight doesn't only store text. When you retain a postmortem, it extracts facts and derives observations and links between them. Here's what my bank looked like after retaining 104 incidents:
•759 world facts
•5 experiences
•182 observations
•946 memories total
•7,135 links
That structure is the reason I chose it. For the concepts behind it, Vectorize's agent memory primer is a good read, and the docs cover the API.
Using it: three calls
The integration surface is small:
python
client = Hindsight(
base_url=os.getenv("HINDSIGHT_API_URL"),
api_key=os.getenv("HINDSIGHT_API_KEY"),
)
def store_incident(content: str):
client.retain(bank_id=BANK_ID, content=content)
def recall_incidents(query: str):
return client.recall(bank_id=BANK_ID, query=query)
def reflect_on_incidents(query: str):
return client.reflect(bank_id=BANK_ID, query=query)
retain stores an incident. recall returns matching memories. reflect returns a reasoned answer synthesized across memory.
reflect vs. recall
I call both on every request, because they answer different questions.
python
reflection = reflect_on_incidents(enriched_query) # synthesized answer
scored_memories = recall_with_scores(enriched_query) # raw memories, ranked
reflect is where cross-incident statements come from: "this looks like a dependency-capacity failure, and in similar cases rollback made things worse." I paste that into the prompt as a reasoned starting point.
recall gives me the underlying memories, which I re-rank with a small keyword heuristic (term overlap plus a bonus for memories mentioning a trap) and append as evidence the model can cite.
My reasoning was that a synthesis without evidence is something the model has to trust, and evidence without synthesis is something the model has to assemble on its own. I used both. What I didn't do is run reflect-only and recall-only versions, so I can't tell you how much each contributes.
What I saw
The demo scenario: a checkout service returning 500s on roughly 12% of requests after a 06:31 deploy. I ran the same model with and without the memory block.
Without memory, the model fabricated a NullPointerException, "112 occurrences" of it from a kubectl logs command it never ran, and a Helm revision that doesn't exist. It recommended an immediate rollback.
With memory, it diagnosed a likely dependency-capacity issue (connection-pool exhaustion on Redis or the database) and warned DO NOT roll back, on the grounds that the memory base flags rollback as a trap for this failure class. The answer referenced similar past pool-exhaustion incidents and told me what log messages to look for.
Two honest notes. The memory-backed answer also suggested checking BGP and systemd-networkd changes, which looks like noise from unrelated incidents in the bank. And I can't say whether a plain vector store feeding the same prompt would have produced the same answer, because I didn't try it.
The held-out test
I held out 10 of 114 incidents, kept 104 in memory, and wrote symptom-only queries. With memory, 9 of 10 root causes matched the true category (a tenth run hit a rate limit and counts as a miss). Without memory, 0 of 10 were fully correct: 4 partial and 6 hallucinated.
That comparison is memory vs. no memory. It is not Hindsight vs. anything else.
What I can't claim
•No vector-search comparison. I never built one. Any statement that this architecture is better than plain retrieval would be a guess.
•No reflect/recall ablation. I don't know which call is doing the work.
•n = 10, graded by me against the dataset's true_category label, at the category level.
•Held-out isn't unrelated. Outages cluster into recurring failure classes, so a held-out incident can resemble retained ones, which also flatters any retrieval method.
•I don't know what's happening inside the memory layer in detail. I can report the counts and behavior I observed. For how facts, observations, and links are produced, read the docs and source instead of trusting my summary.
•Retrieval noise. A mixed bank brings unrelated suggestions with it.
If you want to settle the vector-search question, the experiment is straightforward: same held-out set, same model and prompt, one arm with plain vector retrieval and one with Hindsight, plus reflect-only and recall-only arms. I'd like to see that result myself.
Takeaways
1.Separate synthesis from evidence. Give the model a reasoned answer and the raw memories behind it.
2.Structure the memory for the questions you'll ask. I wanted cross-incident patterns, so I wanted more than nearest-paragraph retrieval.
3.Keep the integration thin. Three wrapper functions were enough to swap memory in and out, which made the baseline comparison easy.
4.Don't claim comparisons you didn't run. I tested memory vs. no memory, not architecture vs. architecture.
5.Write down the experiment you'd run next. It turns a limitation into a plan.

Top comments (0)