I Built an Incident Agent That Learns From What Happened Before
When an incident happens, engineers rarely get the luxury of starting with a completely new problem.
A similar issue may have happened weeks or months ago. Someone may already have found the right fix, tried an approach that failed, or discovered an important detail during the investigation. The problem is that this experience is often buried in logs, tickets, conversations, or simply in someone's memory.
That was the idea behind IncidentRecall.
Instead of treating every incident as a completely new problem, we wanted to build an agent that could remember previous incidents and use those experiences when responding to future ones.
The core idea is simple:
Incident → Recall → Recommendation → Feedback → Memory Update → Future Incident
And Hindsight is what makes the memory part of this loop possible.
The problem with starting from scratch
Imagine an engineer receives an incident:
"The service is experiencing repeated failures after a configuration change."
A normal AI assistant can analyze the current information and suggest possible solutions.
But what if the same team had already experienced a very similar failure?
Maybe the previous incident revealed that:
one particular configuration caused the problem,
one proposed fix didn't work,
another fix successfully resolved it,
and the engineers learned something that wasn't obvious from the initial symptoms.
If the new response ignores all of that, the engineer is effectively solving the same problem again.
That's the problem we wanted to address.
IncidentRecall is designed to bring previous incident experiences into the current response instead of treating every incident independently.
What IncidentRecall does
The system revolves around an incident-response workflow.
When a new incident is created, the system can look for relevant experiences from previous incidents.
Those memories can then be used to provide recommendations for the current situation.
The engineer can review the recommendation, provide feedback, and that feedback becomes part of the learning loop.

This is the part of the project that I found most interesting: the system isn't just using memory to answer a question. The response can become another piece of experience that can be useful later.
Why Hindsight mattered
We used Hindsight as the memory layer for the agent.
Hindsight GitHub repository�
Hindsight documentation�
The important distinction for our project was between simply giving an agent more context and giving it a mechanism for retaining and recalling useful experience.
For an incident-response system, previous incidents aren't just static information.
They represent experiences:
What happened?
What was tried?
What worked?
What didn't work?
What did the engineers learn?
That makes memory particularly useful for this kind of system.
The broader concept of agent memory is also described by Vectorize here:
Vectorize's explanation of agent memory�
The interesting part: memory should affect future behavior
The most important design idea for me wasn't simply storing an incident.
It was making sure that the stored experience could influence what happens next.
Without a feedback loop, memory can easily become just another database.
Our intended cycle is:
Incident
↓
Recall relevant experience
↓
Generate recommendation
↓
Engineer evaluates it
↓
Feedback
↓
Update memory
↓
Use that experience later
This gives the system a way to carry knowledge from one incident into another.
My role: Testing and Demo
My main contribution to the project was testing and demonstration.
Rather than focusing on building one particular application component, I worked on understanding the complete flow and checking whether the system behaved as expected from a user's perspective.
I tested the major parts of the incident-response workflow, including:
Creating an incident
Checking the recall of previous incidents
Reviewing recommendations
Providing engineer feedback
Checking the memory-update flow
Running the complete workflow from incident creation to future recall
One of the important things I focused on was making sure that the different parts of the system worked together rather than testing each screen or feature in isolation.
For example, it wasn't enough for the recommendation screen to display correctly.
I needed to verify the complete chain:
Incident
↓
Recall
↓
Recommendation
↓
Feedback
↓
Memory Update
That helped me understand the project as a complete system rather than as a collection of separate features.
Testing the complete workflow
During testing, I looked at the system from the perspective of someone actually using it.
A typical flow was:
Step 1 — Create an incident
A new incident is introduced into the system.
Step 2 — Recall
The system searches its previous experience for incidents that could be relevant.
Step 3 — Recommendation
The recalled information is used to support the recommendation.
Step 4 — Feedback
The engineer can evaluate the recommendation and provide feedback.
Step 5 — Memory update
The experience becomes part of the system's memory.
Step 6 — Future incident
When a similar incident appears later, the previous experience can become relevant again.
This was also the flow I focused on when preparing the project demonstration.
What I learned from testing an AI system
Testing an AI-based application felt different from testing a simple application with completely deterministic outputs.
With a traditional application, I can often define an expected output and check whether the actual output matches it.
With an AI agent, the behavior can involve context, memory, and generated recommendations.
That made the testing process more about checking whether the overall behavior and workflow made sense, rather than checking only whether a particular sentence was returned.
It also made me pay more attention to the transitions between components.
For example:
Is the incident actually reaching the recall stage?
Is relevant previous experience being surfaced?
Does the recommendation reflect the recalled information?
Does feedback make it through the memory-update stage?
Those questions are just as important as whether the interface looks correct.
The before-and-after idea
The easiest way to understand the value of memory is to compare two situations.
Without persistent memory
New incident
↓
Analyze current context
↓
Generate recommendation
The previous experience isn't part of the reasoning process.
With IncidentRecall
New incident
↓
Recall previous experience
↓
Use relevant memories
↓
Generate recommendation
↓
Engineer feedback
↓
Update memory
The second approach gives the agent an opportunity to use what the system has learned from previous incidents.
That was the central idea behind our project.
What surprised me
The most interesting part of the project wasn't simply connecting an AI model to a memory system.
It was realizing that memory changes the way we think about testing an agent.
Once previous experiences can affect future behavior, testing becomes an ongoing process.
You aren't only asking:
"Does the agent give the correct answer?"
You also need to ask:
"What does the agent remember?"
"What happens when it remembers the wrong thing?"
"Does feedback change what happens later?"
"Does a previous incident actually influence a similar future incident?"
Those are questions I hadn't considered as deeply before working on this project.
Lessons I took away
- Memory is more useful when it is part of a feedback loop Simply storing information isn't enough. The useful part is the connection between past experience and future behavior.
- End-to-end testing matters Individual components can work correctly while the complete workflow still has problems. Testing the entire incident → recall → recommendation → feedback → memory cycle helped expose that difference.
- AI systems need a different testing mindset Generated responses aren't always deterministic, so testing needs to consider behavior, context, and workflow in addition to exact outputs.
- Previous failures can be valuable data An unsuccessful approach isn't necessarily useless. For an incident-response system, knowing what didn't work can be just as useful as knowing what did.
- A good demo should tell the system's story For the project demonstration, I found it much easier to explain the system by following one incident through the complete workflow rather than showing isolated screens. Final thoughts IncidentRecall started with a straightforward question: What if an incident-response agent could actually learn from the incidents it had already seen? The answer isn't just about storing previous incidents. It's about creating a loop where experiences can be recalled, recommendations can be evaluated, feedback can be captured, and future incidents can benefit from what happened before. For me, working on the testing and demonstration side of this project was a valuable way to understand that entire loop. It gave me hands-on experience with testing an AI-based system, thinking about memory-driven behavior, finding issues across an end-to-end workflow, and explaining a technical system through a live demonstration. The bigger takeaway I have is simple: An agent becomes much more interesting when its past can change what it does next.
Top comments (0)