I Built an Incident Response Agent That Remembers What Happened Before
A database timeout shouldn't have to be solved from scratch every time it happens.
That was the idea that got me interested in building this project.
When an incident happens in a real system, there is usually some history behind it. Maybe the same service had a similar problem a few weeks ago. Maybe someone already found the root cause and fixed it. The problem is that this information isn't always available when another incident occurs.
So I started thinking about a simple question:
What if an AI incident-response agent could remember how previous incidents were solved?
That became the idea behind my project, AI Incident Response with Persistent Memory Using Hindsight.
I built it using Gemini for reasoning, Hindsight for persistent memory, and Streamlit for the interface.
The main idea is:
Gemini reasons. Hindsight remembers.
GitHub: https://github.com/2300032753/incident_response_agent
Live Demo: https://ai-response-agent.streamlit.app/
Why I wanted to build this
I chose incident response because it is a good example of where previous experience can actually matter.
Imagine that an Order Service starts showing database connection timeouts during a traffic spike.
An engineer could start checking the database, network, application configuration, server resources, and so on.
But suppose the same team had already faced this problem before and discovered that the real cause was connection pool exhaustion.
That previous investigation is valuable.
The question is: how can an AI agent make use of it?
I didn't want the agent to simply generate an answer based on the current error. I wanted it to first ask whether something similar had happened before.
My first incident
I created a sample incident called INC-001 — Database Connection Timeout.
The Order Service was experiencing database connection timeouts during a traffic spike.
After investigation, the root cause was:
Database connection pool exhaustion.
The solution was to increase the database connection pool from 20 to 50 and restart the affected service.
I also recorded the lesson from the incident:
When database timeouts occur during high traffic, check connection pool utilization before investigating unrelated network failures.
This was the first experience I wanted the agent to remember.
I stored more than just the error message. The memory included the incident ID, service, severity, description, root cause, resolution, lesson learned, and date.
That turned out to be important because a raw error message isn't enough to understand what actually happened.
Adding Hindsight
The next part was connecting the incident to Hindsight.
I use Hindsight to retain the incident after it has been resolved:
client.retain(
bank_id=BANK_ID,
content=memory_text,
context="Production incident investigation and resolution",
metadata={
"incident_id": incident["id"],
"service": incident["service"],
"severity": incident["severity"]
}
)
One thing I wanted to make clear while building this was that Hindsight isn't retraining Gemini.
I'm not training the model again whenever a new incident happens.
Instead, I'm storing useful experiences so they can be retrieved later.
That distinction is important to how I designed the project.
The part I found most interesting: recall
Storing an incident is useful, but the real test is whether the agent can find it later.
For a new incident, I use Hindsight's recall functionality:
result = client.recall(
bank_id=BANK_ID,
query=query,
types=["experience", "observation"],
max_tokens=3000,
budget="mid"
)
Now imagine another database timeout happens.
The agent can search its previous experiences and find INC-001.
It can retrieve information such as:
Previous Incident: INC-001
Root Cause:
Database connection pool exhaustion
Resolution:
Pool increased from 20 to 50
Lesson:
Check connection pool utilization during
high-traffic database timeouts.
That information can then be given to Gemini as context for the current investigation.
This is the part of the project that made the idea feel useful to me.
The agent isn't just answering:
"What could be wrong?"
It can also consider:
"Have we dealt with something similar before?"
Building the Streamlit application
I used Streamlit for the interface because I wanted to focus more on the agent and memory workflow rather than spend a lot of time building a separate frontend.
The application has a few main sections.
Dashboard gives an overview of the incident information.
Report Incident lets the user enter a new incident, including the service, severity, error, and description.
AI Investigation runs the investigation workflow.
Memory lets the user see the stored incident experiences.
And Agent Chat allows questions such as:
- Have we seen this problem before?
- How did we solve the previous database timeout?
- What did we learn from previous incidents?
- What should I investigate first?
The goal was to make the memory something an engineer can actually interact with instead of something hidden in the backend.
How I structured the project
I kept the code separated into a few simple files:
incident-response-agent/
│
├── app.py
├── agent.py
├── hindsight_memory.py
│
├── data/
│ └── incidents.json
│
├── requirements.txt
├── runtime.txt
└── .env.example
app.py handles the Streamlit application.
agent.py handles Gemini.
hindsight_memory.py contains the Hindsight memory operations.
The sample incident data is kept under data/.
Separating the memory code from the UI helped me test the Hindsight part independently before putting everything together.
One thing I learned during the project
The biggest thing I learned is that an AI agent having reasoning doesn't automatically mean it has useful memory.
At first, I was thinking mostly about the model's ability to answer questions.
While working on this project, I started thinking more about what information should actually be remembered.
For an incident, I found these pieces particularly useful:
What happened?
Why did it happen?
How was it fixed?
What should we remember for next time?
That is much more useful than saving only:
"Database connection timeout."
I also realized that retrieved memory shouldn't simply be treated as the answer.
Two incidents can look similar but have completely different causes.
For example, one database timeout could be caused by connection pool exhaustion, while another could be caused by an actual database outage.
So the previous incident should provide context, not automatically determine the solution.
That is why the combination of Hindsight and Gemini makes sense in this project.
Hindsight provides the previous experience. Gemini reasons about the current situation.
What the project doesn't do yet
This is still a prototype, and I don't want to present it as a complete production incident-management system.
The current version uses structured incident data and demonstrates the memory workflow.
A real production version would need integrations with things like:
- Application logs
- Monitoring systems
- Alerting platforms
- Incident-management tools
- Slack or Microsoft Teams
It could also automatically save the important information from a resolved incident instead of requiring someone to enter it manually.
There is also more work needed around deciding when two incidents are actually similar enough to use as previous context.
What I would build next
If I continue developing this, I'd like to connect it to real incident and monitoring data.
Some of the features I'd like to add are:
- Automatic incident detection
- Better incident similarity matching
- Automatic severity classification
- Incident timelines
- Automated post-incident reports
- Recommendations based on previous incidents
- Slack or Teams notifications
The most interesting part for me would be seeing the memory grow over time.
A resolved incident wouldn't just become an old ticket. It could become useful context for a future investigation.
Final thoughts
This project started with a simple question:
Can an AI incident-response agent remember what a team has already learned?
My approach was to separate memory from reasoning.
Hindsight keeps the previous experiences.
Gemini uses those experiences as context while reasoning about a new incident.
Streamlit gives the user a simple interface to interact with the system.
The workflow becomes:
Investigate
↓
Resolve
↓
Remember
↓
Recall
↓
Investigate with context
For me, that's the most interesting part of the project.
The goal isn't to make the AI magically know everything.
It's to make sure useful experience isn't lost after an incident is resolved.
Gemini reasons. Hindsight remembers.
Project Links
GitHub:
https://github.com/2300032753/incident_response_agent
Live Demo:
https://ai-response-agent.streamlit.app/
Hindsight GitHub:
https://github.com/vectorize-io/hindsight
Hindsight Documentation:
https://hindsight.vectorize.io/
What is Agent Memory:
https://vectorize.io/what-is-agent-memory
Top comments (0)