<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Akshaya Simhadri</title>
    <description>The latest articles on DEV Community by Akshaya Simhadri (@akshaya_simhadri_e6e05bef).</description>
    <link>https://dev.to/akshaya_simhadri_e6e05bef</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146974%2F635c77dc-c95b-44c1-8969-09fd49202fab.png</url>
      <title>DEV Community: Akshaya Simhadri</title>
      <link>https://dev.to/akshaya_simhadri_e6e05bef</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/akshaya_simhadri_e6e05bef"/>
    <language>en</language>
    <item>
      <title>AI Incident Response with Persistent Memory Using Hindsight</title>
      <dc:creator>Akshaya Simhadri</dc:creator>
      <pubDate>Mon, 28 Sep 2026 18:08:05 +0000</pubDate>
      <link>https://dev.to/akshaya_simhadri_e6e05bef/ai-incident-response-with-persistent-memory-using-hindsight-3896</link>
      <guid>https://dev.to/akshaya_simhadri_e6e05bef/ai-incident-response-with-persistent-memory-using-hindsight-3896</guid>
      <description>&lt;h1&gt;
  
  
  I Built an Incident Response Agent That Remembers What Happened Before
&lt;/h1&gt;

&lt;p&gt;A database timeout shouldn't have to be solved from scratch every time it happens.&lt;/p&gt;

&lt;p&gt;That was the idea that got me interested in building this project.&lt;/p&gt;

&lt;p&gt;When an incident happens in a real system, there is usually some history behind it. Maybe the same service had a similar problem a few weeks ago. Maybe someone already found the root cause and fixed it. The problem is that this information isn't always available when another incident occurs.&lt;/p&gt;

&lt;p&gt;So I started thinking about a simple question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if an AI incident-response agent could remember how previous incidents were solved?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That became the idea behind my project, &lt;strong&gt;AI Incident Response with Persistent Memory Using Hindsight&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I built it using &lt;strong&gt;Gemini for reasoning, Hindsight for persistent memory, and Streamlit for the interface&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The main idea is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Gemini reasons. Hindsight remembers.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/2300032753/incident_response_agent" rel="noopener noreferrer"&gt;https://github.com/2300032753/incident_response_agent&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Live Demo:&lt;/strong&gt; &lt;a href="https://ai-response-agent.streamlit.app/" rel="noopener noreferrer"&gt;https://ai-response-agent.streamlit.app/&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  Why I wanted to build this
&lt;/h2&gt;

&lt;p&gt;I chose incident response because it is a good example of where previous experience can actually matter.&lt;/p&gt;

&lt;p&gt;Imagine that an Order Service starts showing database connection timeouts during a traffic spike.&lt;/p&gt;

&lt;p&gt;An engineer could start checking the database, network, application configuration, server resources, and so on.&lt;/p&gt;

&lt;p&gt;But suppose the same team had already faced this problem before and discovered that the real cause was &lt;strong&gt;connection pool exhaustion&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That previous investigation is valuable.&lt;/p&gt;

&lt;p&gt;The question is: &lt;strong&gt;how can an AI agent make use of it?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I didn't want the agent to simply generate an answer based on the current error. I wanted it to first ask whether something similar had happened before.&lt;/p&gt;


&lt;h2&gt;
  
  
  My first incident
&lt;/h2&gt;

&lt;p&gt;I created a sample incident called &lt;strong&gt;INC-001 — Database Connection Timeout&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The Order Service was experiencing database connection timeouts during a traffic spike.&lt;/p&gt;

&lt;p&gt;After investigation, the root cause was:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Database connection pool exhaustion.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The solution was to increase the database connection pool from &lt;strong&gt;20 to 50&lt;/strong&gt; and restart the affected service.&lt;/p&gt;

&lt;p&gt;I also recorded the lesson from the incident:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;When database timeouts occur during high traffic, check connection pool utilization before investigating unrelated network failures.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This was the first experience I wanted the agent to remember.&lt;/p&gt;

&lt;p&gt;I stored more than just the error message. The memory included the incident ID, service, severity, description, root cause, resolution, lesson learned, and date.&lt;/p&gt;

&lt;p&gt;That turned out to be important because a raw error message isn't enough to understand what actually happened.&lt;/p&gt;


&lt;h2&gt;
  
  
  Adding Hindsight
&lt;/h2&gt;

&lt;p&gt;The next part was connecting the incident to Hindsight.&lt;/p&gt;

&lt;p&gt;I use Hindsight to retain the incident after it has been resolved:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;memory_text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Production incident investigation and resolution&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;incident_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;service&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;service&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;severity&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;severity&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One thing I wanted to make clear while building this was that Hindsight isn't retraining Gemini.&lt;/p&gt;

&lt;p&gt;I'm not training the model again whenever a new incident happens.&lt;/p&gt;

&lt;p&gt;Instead, I'm storing useful experiences so they can be retrieved later.&lt;/p&gt;

&lt;p&gt;That distinction is important to how I designed the project.&lt;/p&gt;




&lt;h2&gt;
  
  
  The part I found most interesting: recall
&lt;/h2&gt;

&lt;p&gt;Storing an incident is useful, but the real test is whether the agent can find it later.&lt;/p&gt;

&lt;p&gt;For a new incident, I use Hindsight's recall functionality:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;types&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;experience&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;observation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;budget&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now imagine another database timeout happens.&lt;/p&gt;

&lt;p&gt;The agent can search its previous experiences and find INC-001.&lt;/p&gt;

&lt;p&gt;It can retrieve information such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Previous Incident: INC-001

Root Cause:
Database connection pool exhaustion

Resolution:
Pool increased from 20 to 50

Lesson:
Check connection pool utilization during
high-traffic database timeouts.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That information can then be given to Gemini as context for the current investigation.&lt;/p&gt;

&lt;p&gt;This is the part of the project that made the idea feel useful to me.&lt;/p&gt;

&lt;p&gt;The agent isn't just answering:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What could be wrong?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It can also consider:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Have we dealt with something similar before?"&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Building the Streamlit application
&lt;/h2&gt;

&lt;p&gt;I used Streamlit for the interface because I wanted to focus more on the agent and memory workflow rather than spend a lot of time building a separate frontend.&lt;/p&gt;

&lt;p&gt;The application has a few main sections.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dashboard&lt;/strong&gt; gives an overview of the incident information.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Report Incident&lt;/strong&gt; lets the user enter a new incident, including the service, severity, error, and description.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI Investigation&lt;/strong&gt; runs the investigation workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory&lt;/strong&gt; lets the user see the stored incident experiences.&lt;/p&gt;

&lt;p&gt;And &lt;strong&gt;Agent Chat&lt;/strong&gt; allows questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Have we seen this problem before?&lt;/li&gt;
&lt;li&gt;How did we solve the previous database timeout?&lt;/li&gt;
&lt;li&gt;What did we learn from previous incidents?&lt;/li&gt;
&lt;li&gt;What should I investigate first?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal was to make the memory something an engineer can actually interact with instead of something hidden in the backend.&lt;/p&gt;




&lt;h2&gt;
  
  
  How I structured the project
&lt;/h2&gt;

&lt;p&gt;I kept the code separated into a few simple files:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;incident-response-agent/
│
├── app.py
├── agent.py
├── hindsight_memory.py
│
├── data/
│   └── incidents.json
│
├── requirements.txt
├── runtime.txt
└── .env.example
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;app.py&lt;/code&gt; handles the Streamlit application.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;agent.py&lt;/code&gt; handles Gemini.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;hindsight_memory.py&lt;/code&gt; contains the Hindsight memory operations.&lt;/p&gt;

&lt;p&gt;The sample incident data is kept under &lt;code&gt;data/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Separating the memory code from the UI helped me test the Hindsight part independently before putting everything together.&lt;/p&gt;




&lt;h2&gt;
  
  
  One thing I learned during the project
&lt;/h2&gt;

&lt;p&gt;The biggest thing I learned is that &lt;strong&gt;an AI agent having reasoning doesn't automatically mean it has useful memory&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;At first, I was thinking mostly about the model's ability to answer questions.&lt;/p&gt;

&lt;p&gt;While working on this project, I started thinking more about what information should actually be remembered.&lt;/p&gt;

&lt;p&gt;For an incident, I found these pieces particularly useful:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happened?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why did it happen?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How was it fixed?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What should we remember for next time?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is much more useful than saving only:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Database connection timeout."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I also realized that retrieved memory shouldn't simply be treated as the answer.&lt;/p&gt;

&lt;p&gt;Two incidents can look similar but have completely different causes.&lt;/p&gt;

&lt;p&gt;For example, one database timeout could be caused by connection pool exhaustion, while another could be caused by an actual database outage.&lt;/p&gt;

&lt;p&gt;So the previous incident should provide &lt;strong&gt;context&lt;/strong&gt;, not automatically determine the solution.&lt;/p&gt;

&lt;p&gt;That is why the combination of Hindsight and Gemini makes sense in this project.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hindsight provides the previous experience. Gemini reasons about the current situation.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What the project doesn't do yet
&lt;/h2&gt;

&lt;p&gt;This is still a prototype, and I don't want to present it as a complete production incident-management system.&lt;/p&gt;

&lt;p&gt;The current version uses structured incident data and demonstrates the memory workflow.&lt;/p&gt;

&lt;p&gt;A real production version would need integrations with things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Application logs&lt;/li&gt;
&lt;li&gt;Monitoring systems&lt;/li&gt;
&lt;li&gt;Alerting platforms&lt;/li&gt;
&lt;li&gt;Incident-management tools&lt;/li&gt;
&lt;li&gt;Slack or Microsoft Teams&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It could also automatically save the important information from a resolved incident instead of requiring someone to enter it manually.&lt;/p&gt;

&lt;p&gt;There is also more work needed around deciding when two incidents are actually similar enough to use as previous context.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I would build next
&lt;/h2&gt;

&lt;p&gt;If I continue developing this, I'd like to connect it to real incident and monitoring data.&lt;/p&gt;

&lt;p&gt;Some of the features I'd like to add are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Automatic incident detection&lt;/li&gt;
&lt;li&gt;Better incident similarity matching&lt;/li&gt;
&lt;li&gt;Automatic severity classification&lt;/li&gt;
&lt;li&gt;Incident timelines&lt;/li&gt;
&lt;li&gt;Automated post-incident reports&lt;/li&gt;
&lt;li&gt;Recommendations based on previous incidents&lt;/li&gt;
&lt;li&gt;Slack or Teams notifications&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The most interesting part for me would be seeing the memory grow over time.&lt;/p&gt;

&lt;p&gt;A resolved incident wouldn't just become an old ticket. It could become useful context for a future investigation.&lt;/p&gt;




&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;This project started with a simple question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can an AI incident-response agent remember what a team has already learned?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My approach was to separate memory from reasoning.&lt;/p&gt;

&lt;p&gt;Hindsight keeps the previous experiences.&lt;/p&gt;

&lt;p&gt;Gemini uses those experiences as context while reasoning about a new incident.&lt;/p&gt;

&lt;p&gt;Streamlit gives the user a simple interface to interact with the system.&lt;/p&gt;

&lt;p&gt;The workflow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Investigate
    ↓
Resolve
    ↓
Remember
    ↓
Recall
    ↓
Investigate with context
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For me, that's the most interesting part of the project.&lt;/p&gt;

&lt;p&gt;The goal isn't to make the AI magically know everything.&lt;/p&gt;

&lt;p&gt;It's to make sure useful experience isn't lost after an incident is resolved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gemini reasons. Hindsight remembers.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Project Links
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/2300032753/incident_response_agent" rel="noopener noreferrer"&gt;https://github.com/2300032753/incident_response_agent&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Live Demo:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://ai-response-agent.streamlit.app/" rel="noopener noreferrer"&gt;https://ai-response-agent.streamlit.app/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hindsight GitHub:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;https://github.com/vectorize-io/hindsight&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hindsight Documentation:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;https://hindsight.vectorize.io/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is Agent Memory:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;https://vectorize.io/what-is-agent-memory&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>sre</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
