I've lost count of how many times I've been paged into an incident that felt strangely familiar.
You search Slack. You dig through old tickets. You open a few post-mortems. Twenty minutes later, you find it: the exact same root cause happened four months ago.
The fix was already known. It just wasn't available to the person who needed it when they needed it.
So I built a small incident-response agent that can remember past incidents.
Not "remember" in the vague marketing sense. It stores past incidents, recalls relevant ones when a new alert comes in, and can reason over those memories before suggesting what to investigate.
The memory layer is Hindsight, and this is what building on it actually looked like.
What it does
The agent has three operations, all backed by Hindsight:
-
retain()— store a past incident, resolution, or other useful outcome -
recall()— retrieve incidents relevant to the current situation -
reflect()— reason over the recalled information and produce an answer, with citations back to the original memories
That's essentially the interface.
Everything around it is the agent logic, prompting, and formatting.
At the core, the triage flow looks like this:
def triage(incident: str):
memories = client.recall(
bank_id=BANK_ID,
query=incident
)
answer = client.reflect(
bank_id=BANK_ID,
query=(
"New incident: " + incident + "\n"
"Using past incidents, what is the most likely root cause, "
"which runbook should we run first, and which past incidents "
"does this resemble?"
),
)
return answer
The interesting part wasn't wiring up these API calls.
It was deciding what "memory" should actually mean for an on-call agent.
Recall vs. plain retrieval
My first instinct was simple: take every past post-mortem, put it into a vector store, and use similarity search whenever a new incident arrives.
That gets you documents that sound like the current incident.
But that's not necessarily what an on-call engineer needs.
The useful question is often:
"What happened the last time we saw this, what did we try, and what actually fixed it?"
That distinction matters.
One of my seeded incidents involved a Redis eviction storm. The first thing the team tried was restarting the pods.
It didn't work.
The actual fix was a configuration change.
A basic similarity search could easily surface "restart pods" because that phrase appears prominently in the incident. But the fact that the action was attempted and failed is much more important than the fact that it was mentioned.
This is where the reflect() step becomes interesting.
Instead of simply returning the nearest matching chunks, the idea is to give the model the recalled incidents and ask it to reason over them.
That gives the agent a chance to distinguish:
- something that was tried and worked
- something that was tried and failed
- the actual root cause
- the runbook that resolved the incident
That's much closer to the kind of memory I want an on-call agent to have.
Before and after
Here's the same incident prompt run in two different ways.
Baseline: no incident history
First, I ran a plain LLM call with no access to previous incidents:
python agent.py baseline "checkout-service is throwing 502s right after today's deploy"
The response was essentially generic troubleshooting advice:
Check the deploy diff, inspect error logs, and consider rolling back.
That's reasonable advice.
But it could apply to almost any service on any day.
The model has no idea that we've seen this exact pattern before.
With incident memory
Now I ran the same incident through the agent with Hindsight:
python agent.py triage "checkout-service is throwing 502s right after today's deploy"
Here's the recall output, unedited:
1. Any 502 burst right after a deploy should prompt checking DB pool metrics first.
3. Incident INC-101: checkout-service returned 502s after a deploy; root cause was
DB connection pool exhaustion due to a new ORM call opening a transaction and
never closing it; resolved by rolling back the deploy, increasing pool size
from 20 to 50, adding a query timeout; runbook 'db-connection-pool-exhaustion'
was used to resolve the incident.
4. Incident INC-139 in June 2026: checkout-service experienced 502 errors after a
deploy, similar to INC-101; resolved by rolling back the deployment and
patching the missing close in the ORM call.
11. Runbook 'restart pods' was ineffective for Incident INC-114.
This is where the difference becomes useful.
The agent isn't just seeing that "checkout-service" and "502" appeared in an old document.
It found two previous incidents with the same combination of symptoms and a similar underlying cause: a database connection being left open.
It also surfaced the runbook that was used to resolve the earlier incidents:
db-connection-pool-exhaustion
And there's another useful piece of information:
Runbook 'restart pods' was ineffective for Incident INC-114.
That negative result matters.
A fresh on-call engineer might see "restart pods" in an old incident and try it again. A memory system that stores outcomes can tell you that it was already attempted and didn't help.
That's the kind of context I wanted the agent to preserve.
One limitation I ran into
There's an important caveat.
My reflect() step, which is supposed to turn the recalled memories into one synthesized response, hit a key-configuration error on my server during testing.
So the recall layer above is what actually ran successfully in this version.
I haven't treated the reflection output as a successful part of the demo yet. Fixing that configuration issue is the next step.
I'd rather call that out than pretend the complete pipeline worked perfectly.
The learning loop
The other part I like about this architecture is what happens after an incident is resolved.
The agent can store the outcome:
def resolved(summary: str):
client.retain(
bank_id=BANK_ID,
content=summary
)
That means the resolution becomes part of the memory available to future incidents.
For example:
Incident → Investigation → Resolution → Retain
↓
Future incidents
↓
Recall
Over time, this can turn individual incident histories into a reusable operational memory.
I'm careful not to claim that simply adding more incidents automatically makes the agent better. The quality of the retained information, retrieval, and reasoning still matters.
But the important architectural property is there: resolved incidents can become useful context for future incidents.
That's different from treating every incident as an isolated ticket.
What I learned
1. Rate limits are a real constraint
While seeding five incidents back-to-back, I hit Groq's free-tier token limit surprisingly quickly.
I ended up adding a deliberate delay between retain() calls.
If you're building a prototype on a free API tier, rate limits aren't something to worry about later. They're part of the development environment from the beginning.
2. Retrieval alone isn't enough
Retrieval gets you relevant information.
Reasoning over that information is what can turn it into something actionable.
For an incident-response agent, "here are five similar incidents" is less useful than:
"Two previous incidents had the same symptoms. Both were caused by connection-pool exhaustion, and this runbook resolved them."
That's the difference I'm interested in exploring with reflect().
3. Negative results are valuable memories
"We tried X and it didn't work" can be just as important as "we tried Y and it fixed the problem."
Incident reports often preserve the final resolution while losing the failed attempts.
For an on-call agent, those failed attempts are useful because they prevent the same investigation paths from being repeated unnecessarily.
4. The interesting design decisions are in the query
The API calls themselves aren't particularly complicated.
The harder question is what you ask the system to do with the retrieved information.
For example, my reflection prompt explicitly asks for:
- likely root cause
- first runbook to try
- similar historical incidents
That framing determines what the resulting answer is useful for.
In other words, the memory system is only part of the design. The questions you ask it matter just as much.
5. Start with real incidents when you can
For the initial prototype, I used realistic but fabricated incidents so I could iterate quickly.
That was useful for getting the system working.
But real post-mortems contain the kind of specific operational details that make retrieval much more useful: unusual symptoms, failed fixes, exact configuration changes, service dependencies, and environment-specific behavior.
That's where I expect this kind of system to become much more valuable.
What's next?
The immediate next step is fixing the configuration issue with reflect() and completing the full recall → reasoning → response flow.
After that, I'd like to experiment with a few things:
- comparing memory-based triage against a plain LLM baseline
- testing retrieval quality with a larger set of incidents
- measuring whether negative outcomes are retrieved reliably
- adding structured incident metadata
- connecting the agent to actual post-mortems and runbooks
- evaluating how often the suggested root cause matches the eventual resolution
The goal isn't to replace the on-call engineer.
It's to reduce the time spent rediscovering things the team has already learned.
Final thoughts
The most interesting thing I took away from this experiment is that "AI memory" becomes much more concrete when you apply it to a problem like incident response.
A context window can give a model more information.
A memory system can give it access to information from before the current conversation.
But the useful part isn't simply remembering that an incident happened.
It's remembering:
- what happened
- what was tried
- what failed
- what worked
- why it worked
- and which previous incidents looked similar
That's the kind of memory I'd want available the next time I'm paged at 2 AM.
If you're building something similar, the Hindsight documentation is a good place to start. The retain() / recall() / reflect() split is a useful abstraction once you've experienced the difference between "documents that match" and "context I can actually act on."

Top comments (0)