Most incident-response tools are good at helping engineers deal with what is happening right now.
The harder problem is what happens the next time.
An organisation may have already seen the same failure several times. Engineers may have documented what happened, what they tried, what failed, and what eventually fixed it. But when a similar incident happens again, that knowledge is usually buried in old incident reports or scattered across documentation.
I wanted to see what would happen if an incident-response agent could actually remember that experience and use it the next time.
That became Memento.
Memento is an incident-response agent built around persistent organisational memory. It takes a current incident, recalls relevant experience from previous incidents, uses that experience to shape the investigation, and then retains the outcome so future investigations can benefit from it.
The interesting part isn't simply that it remembers.
It's that remembering changes what it does.
The problem
Take a payment service returning HTTP 503 errors.
A runbook might say:
Restart payment-service first.
That's a reasonable instruction in isolation.
But imagine that several previous incidents showed that restarting the service wasn't actually solving the underlying problem. The real causes were repeatedly connected to infrastructure capacity, upstream dependencies, connection handling, or surrounding services.
That information already exists.
The problem is that the next engineer dealing with the incident may not find it at the moment they need it.
I wanted the agent to bring that experience into the current investigation automatically.
The core loop
Memento's workflow is:
Incident
↓
Recall relevant historical experience
↓
Generate investigation plan
↓
Investigate and resolve
↓
Record outcome
↓
Retain the new learning
↓
Use it during future incidents
Hindsight is the persistent memory layer behind this loop.
I use it to retain incident experiences, recall relevant historical information, and reflect over accumulated memory when the system needs a more grounded conclusion.
The application itself is split into a Next.js frontend and a FastAPI backend. The backend handles the incident workflow and communicates with Hindsight and the language model.
I deliberately kept the architecture focused around Hindsight rather than adding another separate vector database just for retrieval.
The part that matters: memory changes the investigation
Here's where the difference becomes visible.
Without historical memory, an incident agent might produce something like:
- Restart payment-service
- Check application logs
- Verify pod health
After recalling relevant organisational experience, Memento can produce a different investigation order:
- Inspect infrastructure and connection metrics
- Check upstream dependencies
- Review circuit-breaker behaviour
- Verify external dependency health
- Review recent deployments
- Check pod and resource health
- Restart payment-service only as a last resort
The important difference isn't that the second answer is longer.
It's that the priority changed.
Previous incidents changed the order in which the current incident should be investigated.
That's the behaviour I wanted persistent agent memory to produce.
A small look at the implementation
When an incident is created, Memento stores the experience in Hindsight:
await hindsight.aretain(
bank_id="memento-demo-final",
content=incident_report,
context="incident-report",
document_id=incident_id,
)
When a new incident arrives, the system recalls relevant experience:
memories = await hindsight.arecall(
bank_id="memento-demo-final",
query=incident_query,
types=["world", "experience", "observation"],
)
That recalled experience then becomes part of the reasoning process.
This distinction mattered to me while building the system.
A search system can tell you that similar incidents existed.
An agent using persistent memory should be able to do something useful with what it found.
Knowledge Drift
The other problem I wanted to solve was knowledge drift.
Runbooks don't automatically stay correct.
Infrastructure changes. Dependencies change. Systems evolve. Engineers discover better ways of troubleshooting problems.
That means documented procedures can gradually diverge from what the organisation has actually learned through experience.
Memento compares current runbook guidance against accumulated incident experience and surfaces those conflicts.
For example:
RUNBOOK
Restart payment-service first
HISTORICAL EXPERIENCE
Previous incidents repeatedly showed that restarting the
service did not resolve the underlying failure.
DRIFT
Conflict detected
The system can distinguish between situations where historical experience conflicts with the runbook and situations where the two remain aligned.
That turns incident memory into something more useful than a collection of old reports.
It starts answering a practical engineering question:
Is our current documentation still consistent with what we have actually learned?
Why I didn't want memory hidden in the backend
One thing I changed while building Memento was how much of the memory process was visible in the interface.
It would be easy to let the system use Hindsight silently and only show the final recommendation.
I didn't think that would make a very convincing product.
When an agent changes an investigation plan, the engineer should be able to understand why.
So the interface exposes:
historical incident evidence
learned observations
incident history
knowledge drift
the reason a recommendation changed
I'm not trying to expose every internal detail of the memory system.
I'm trying to make the relationship between memory and behaviour understandable.
The learning loop doesn't stop when the incident ends
There's another important part of the system.
Suppose an engineer investigates the payment incident and discovers that an HPA configuration caused the capacity problem.
That outcome can become another memory.
A future incident can then retrieve it.
So the loop becomes:
Incident A
↓
Investigation
↓
Outcome
↓
New memory
Incident B
↓
Recall Incident A
↓
Different investigation
That feedback loop is what makes Memento more interesting than a static incident knowledge base.
The memory isn't just historical context.
It can influence future behaviour.
One lesson I learned: similarity isn't enough
The hardest part wasn't getting the model to generate an investigation plan.
The harder problem was deciding how much trust to put in historical similarity.
Two incidents can look very similar and still have completely different root causes.
A payment incident involving Stripe, for example, shouldn't automatically inherit every recommendation from another payment incident just because the words look similar.
That led to an important principle for the system:
Historical evidence should influence reasoning, not replace it.
The agent still needs to consider the current incident.
Memory should give it experience, not blind certainty.
Another lesson: memory quality matters
Persistent memory is only useful when the things being retained are useful.
If every interaction becomes memory without any thought about what should be retained, retrieval eventually becomes noisy.
For Memento, that meant treating incident outcomes and observations as meaningful experiences rather than simply storing every response verbatim.
The quality of tomorrow's investigation depends partly on what the system learned today.
What I wanted Memento to become
The larger idea is simple.
Production incidents shouldn't just become closed tickets.
They should become organisational experience.
An incident tells you what happened.
A good post-incident process tells you what was learned.
Persistent agent memory gives you a way to carry that learning into the next incident.
That's the part of agent systems I find most interesting.
The useful question isn't:
Can the agent remember?
It's:
What changes because it remembered?
Memento is my attempt to build around that question.
Learn more
Hindsight GitHub
Hindsight Documentation
Vectorize — What is Agent Memory?
Top comments (0)