Building an Incident Response Agent That Learns From the Past
An incident starts with a familiar pattern.
A service becomes slow. Requests begin failing. Error rates increase. Engineers open dashboards, check logs, investigate recent changes, and try to understand what went wrong.
Eventually, the root cause is identified and the incident is resolved.
Then the same type of problem appears again a few weeks or months later.
The team may have already solved something similar, but the useful details are often scattered across old postmortems, tickets, documentation, and incident notes.
That made me think about a different approach:
What if an incident-response agent could remember previous incidents and bring that experience into the next investigation?
That question became the foundation for Incident-Memory-Copilot.
The project combines an LLM-based reasoning layer with persistent memory using Hindsight. Instead of asking an AI model to troubleshoot every incident from scratch, the system first searches for relevant previous experiences and then uses those memories as additional context.
The workflow is:
Recall → Reason → Resolve → Retain
The important part isn't simply generating an answer.
The goal is to make organizational incident knowledge available at the moment an engineer needs it.

Incident-Memory-Copilot interface for investigating production incidents with historical memory.
The problem: incident knowledge is easy to lose
Consider an API that starts returning intermittent 502 Bad Gateway errors.
A general-purpose AI assistant might suggest:
- Check application logs.
- Inspect upstream dependencies.
- Review recent deployments.
- Check network connectivity.
- Examine CPU and memory usage.
- Investigate connection pools.
These are reasonable starting points.
But they are still generic.
Now imagine that the same service had experienced a similar problem six months earlier. During that incident, the team discovered that connection-pool exhaustion was responsible, and one attempted configuration change actually made the problem worse.
That historical information could significantly change how the new incident is investigated.
The problem is therefore not just:
"Can an AI generate troubleshooting suggestions?"
It is:
"Can the system find the organization's previous experience that is relevant to this incident?"
That is where persistent memory becomes important.
Designing the Incident-Memory-Copilot
I designed the system around a continuous incident lifecycle:
CURRENT INCIDENT
│
▼
Memory Recall
│
▼
Relevant Past Incidents
│
▼
AI Reasoning
│
▼
Investigation Plan
│
▼
Resolution
│
▼
Postmortem
│
▼
Memory Retention
│
└──────────────► Future Incident
When an engineer submits an incident, the system extracts the important information and searches Hindsight for related memories.
The retrieved information is then provided to the reasoning stage.
After the incident has been resolved, the important lessons from that incident can be stored back into memory.
This creates a continuous learning loop instead of treating every incident as an independent question.
The first step is recall
One of the main architectural decisions was to make memory retrieval happen before the AI generates its investigation.
A simple implementation could send the incident directly to an LLM:
Incident
↓
LLM
↓
Troubleshooting Suggestions
Instead, Incident-Memory-Copilot uses:
Incident
↓
Extract Symptoms
↓
Search Hindsight
↓
Retrieve Related Experiences
↓
Provide Historical Context
↓
LLM Reasoning
↓
Investigation Recommendation
This small architectural difference changes the type of information available to the model.
The model is no longer working only with the current symptoms.
It can also consider what happened during previous incidents with similar characteristics.
Hindsight acts as the persistent memory layer that allows the system to retain and recall this information.

Hindsight recall retrieves relevant historical incident experience before the copilot reasons about the current incident.
What gets stored in memory?
The system doesn't need to remember every piece of text associated with an incident.
The useful information is the experience gained from the incident.
For example:
Service:
Payment API
Symptoms:
Intermittent 502 errors during high traffic
Root Cause:
Connection-pool exhaustion
Successful Resolution:
Increased pool capacity and corrected connection handling
Failed Attempts:
Restarting the service temporarily reduced errors but did not solve the underlying problem
Lesson:
Check connection-pool utilization before repeatedly restarting instances
This kind of information is much more useful for future investigations than simply storing a generic troubleshooting document.
The memory can contain:
- Affected service
- Incident symptoms
- Root cause
- Investigation steps
- Successful actions
- Failed actions
- Resolution
- Lessons learned
- Relevant operational context
After the incident is completed, these details can be retained in Hindsight for future retrieval.
The real test is the second incident
A memory system isn't particularly interesting if we only demonstrate it with the first incident.
The more meaningful scenario is what happens when another incident occurs later.
Suppose the system previously learned about a payment-service outage caused by connection exhaustion.
Several months later, a new incident arrives:
Checkout requests are intermittently failing.
Gateway errors are increasing during
a period of high traffic.
Latency has increased and upstream
connection failures are being observed.
The wording is different.
The incident is not a copy of the previous one.
However, the symptoms may still be related.
The system searches its memory and retrieves the earlier incident experience.
The reasoning process can then combine:
Current Incident
+
Historical Experience
↓
Reasoning
↓
Investigation Plan
This is where persistent memory provides value.
The LLM does not have to somehow remember an incident from a previous conversation. The memory layer retrieves the relevant experience when it is needed.
Memory changes the lifecycle of the agent
A normal chatbot mostly works within the boundaries of a conversation.
An incident-memory agent has a much longer lifecycle:
Incident
↓
Investigation
↓
Resolution
↓
Postmortem
↓
Learning
↓
Memory
↓
Future Incident
↓
Recall
↓
Investigation
The output of one incident becomes potential input for another.
Over time, the system can therefore accumulate operational experience.
This is an important difference between a chatbot that simply generates responses and an agent workflow that maintains persistent memory.
Recall is not the same as reasoning
Another design consideration was separating memory retrieval from reasoning.
The memory system answers:
"What previous experiences might be relevant?"
The reasoning layer answers:
"What does that information mean for the incident happening now?"
The workflow can therefore be represented as:
Current Incident
↓
Recall
↓
Relevant Memories
↓
Reflection
↓
Historical Context
↓
AI Reasoning
↓
Investigation Recommendation
Retrieved memories may contain several pieces of information.
The reasoning stage determines how those pieces relate to the current incident.
This separation also makes the architecture easier to reason about because memory retrieval and AI reasoning have different responsibilities.
What I learned while building it
1. More memory doesn't automatically mean better results
One of the first lessons from designing a memory-based agent is that simply adding more information isn't enough.
The important question is:
Did the system retrieve information that is actually relevant?
If unrelated incidents are retrieved, they can create noise and potentially distract the reasoning process.
That means incident queries and memory structure matter just as much as the LLM prompt.
2. The complete loop is more important than a single prompt
It is relatively straightforward to create a prompt that produces an incident checklist.
The more interesting engineering challenge is building the complete lifecycle:
Recall
↓
Reason
↓
Resolve
↓
Retain
↓
Recall Again
The value of the system appears when information captured during one incident becomes useful during another.
3. Failed troubleshooting attempts are important
Incident documentation often focuses heavily on the successful fix.
But failed actions can also be valuable.
For example, suppose engineers repeatedly restarted a service during an earlier incident. The restart temporarily reduced the errors but did not resolve the underlying issue.
That information can prevent future engineers from immediately repeating the same approach.
So the memory should capture not only:
"What worked?"
but also:
"What did we try that didn't work?"
4. Memory capture should be part of incident resolution
If engineers have to perform a completely separate knowledge-management process after every incident, important information can easily be missed.
Instead, memory retention should naturally follow the incident and postmortem process.
Incident Resolved
↓
Postmortem Created
↓
Important Knowledge Identified
↓
Memory Retained
↓
Available for Future Incidents
This makes memory a part of the operational workflow rather than another administrative task.
Where this project can be useful
Incident-Memory-Copilot is primarily designed for teams that regularly deal with production incidents and need to reuse previous troubleshooting knowledge.
It can be useful for:
- DevOps teams
- SRE teams
- Backend engineering teams
- Platform engineering teams
- Cloud operations teams
- Production support teams
- Organizations with large collections of historical postmortems
The main idea is not to replace engineers.
Instead, the system helps engineers find relevant historical context faster.
An engineer still evaluates the evidence, decides what actions are appropriate, and remains responsible for the actual resolution.
The bigger idea
Production incidents are rarely completely isolated events.
Organizations often encounter related failures involving the same services, dependencies, infrastructure, deployment patterns, or configuration problems.
Every resolved incident can therefore produce knowledge that may be useful later.
The important questions are:
What happened?
Why did it happen?
What fixed it?
What didn't work?
What should we remember?
Traditional documentation can answer these questions, but the information may remain passive until someone manually searches for it.
A memory-enabled agent changes that interaction.
Instead of expecting an engineer to remember every historical incident, the system can retrieve relevant experience when a new incident occurs.
That is the idea behind Incident-Memory-Copilot.
It isn't intended to be an AI that magically knows every production problem.
It isn't just another chatbot generating generic troubleshooting steps.
It is a system designed around a simple lifecycle:
Recall previous experience → reason about the current incident → support the investigation → retain what was learned.
The goal is simple:
The next incident should have the benefit of the incidents that came before it.
Resources
Top comments (0)