What happens when the same production problem occurs twice?
The first time, engineers investigate everything from scratch. They check logs, deployments, database connections, services, and infrastructure. Eventually, they find the root cause, fix the issue, and write a postmortem.
But when a similar incident happens months later, how much of that experience is actually available at the moment it matters?
That question led us to build **IncidentMind for the Hindsight Hackathon, with @code.in
IncidentMind is an AI-powered incident-response system designed around one simple idea:
What if an AI agent could remember what happened during previous incidents and use that experience during the next one?
From Incident History to AI Memory
Traditional incident management systems are good at recording what happened.
But recording an incident and actually learning from it are two different things.
A postmortem might contain the exact information an engineer needs during a future outage: the root cause, what fixed it, which investigation steps worked, and which ones didn't.
The challenge is finding that knowledge at the right moment.
That's where Hindsight becomes a central part of IncidentMind.
Instead of treating previous incidents as static historical records, we use Hindsight as a persistent memory layer for the AI agent.
IncidentMind can retain useful information from incidents and postmortems, then recall relevant memories when a new incident occurs.
This gives the agent something beyond general LLM knowledge: the team's own operational experience.
A Real Example: Payment API 503 Errors
We tested IncidentMind using a Payment API incident.
Imagine a new incident appears:
Payment API experiencing repeated 503 errors.
The description indicates that database connection pool exhaustion may be involved.
An ordinary AI assistant could explain several possible causes of a 503 error.
IncidentMind can go one step further.
It can ask Hindsight:
Have we seen anything similar before?
During our testing, Hindsight recalled relevant historical memories connected to Payment API failures and database connection pool exhaustion.
Among the retrieved memories was a previous P1 incident where database connection pool exhaustion affected the Payment API. The previous incident was resolved by restarting the payment service.
Other recalled information included investigation steps and an approach that had not been effective previously.
Now the AI has something much more useful to reason with.
It isn't simply saying:
"Database connection pool exhaustion is one possible cause."
It can say, in context:
"A similar Payment API incident previously involved database connection pool exhaustion. That incident was resolved by restarting the payment service. Consider checking current connection pool utilization and recent deployment changes as part of the investigation."
The historical incident doesn't prove that the current incident has the same root cause.
Instead, it gives the engineer relevant evidence and a starting point.
That distinction is important.
Making Postmortems Part of the Learning Process
One of the most interesting parts of IncidentMind is what happens after an incident is resolved.
Normally, a postmortem feels like the final step.
For us, it's actually the beginning of the next learning cycle.
IncidentMind stores the postmortem in Supabase with information such as:
- Root cause
- What worked
- What failed
- Lessons learned
- Resolution information
That information is then structured and retained in Hindsight.
So the next time a similar incident occurs, the agent has a chance to retrieve what was learned from the previous one.
This creates a continuous learning mechanism where every resolved incident can potentially become useful experience for future incidents.
The Most Important Part: Learning From Experience
One of the biggest things this project taught us is that memory isn't valuable simply because information is stored.
The real value comes from retrieving the right experience at the right time.
An AI model can already reason about a database connection pool.
But knowing that your team experienced the same problem before, knowing how it was resolved, and knowing what didn't work during that investigation provides a completely different kind of context.
That's what we're exploring with IncidentMind.
The goal isn't to create an AI that magically knows everything.
It's to create an AI agent that can say:
"I've seen something similar before. Here's what happened, here's what we learned, and here's what you may want to investigate."
And after the current incident is resolved, the system doesn't simply forget about it.
The new experience becomes part of its memory.
That creates something we find particularly exciting: an incident-response agent that can become more useful as the organization accumulates experience.
What's Next?
We're continuing to improve IncidentMind toward a more capable and actionable investigation agent.
Some of the areas we're exploring include better presentation of historical evidence, clearer connections between Hindsight memories and AI reasoning, more structured investigation guidance, and making the system's newly acquired knowledge more visible after each incident.
Ultimately, we want IncidentMind to make incident response less about repeatedly starting from zero and more about building on what the team has already learned.
Because the most useful incident-response knowledge isn't always found in a textbook.
Sometimes, it's hidden inside the last outage.
And maybe the most useful AI incident responder isn't the one that knows everything.
Maybe it's the one that remembers what your team already learned.
Top comments (0)