<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Anjali devi Muppidi</title>
    <description>The latest articles on DEV Community by Anjali devi Muppidi (@anjali_01037).</description>
    <link>https://dev.to/anjali_01037</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150555%2F566ed626-e7c9-40ed-80ab-e4a420d0fc2a.png</url>
      <title>DEV Community: Anjali devi Muppidi</title>
      <link>https://dev.to/anjali_01037</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/anjali_01037"/>
    <language>en</language>
    <item>
      <title>I Built an Incident-Response Agent That Remembers What Worked and What Didn’t</title>
      <dc:creator>Anjali devi Muppidi</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:14:20 +0000</pubDate>
      <link>https://dev.to/anjali_01037/i-built-an-ai-incident-agent-with-hindsight-that-remembers-what-failed-g8p</link>
      <guid>https://dev.to/anjali_01037/i-built-an-ai-incident-agent-with-hindsight-that-remembers-what-failed-g8p</guid>
      <description>&lt;p&gt;I Built an Incident-Response Agent That Remembers What Worked and What Didn’t&lt;/p&gt;

&lt;p&gt;When something goes wrong in a software system, engineers usually start by checking the logs and trying to understand what caused the problem.&lt;/p&gt;

&lt;p&gt;But there’s another question that can be just as important:&lt;/p&gt;

&lt;p&gt;Has this happened before, and what did we try last time?&lt;/p&gt;

&lt;p&gt;A team might have already dealt with a similar incident a few days or weeks ago. Maybe one fix worked, while another didn't. If that information isn't available when the next incident happens, someone might end up repeating the same failed approach.&lt;/p&gt;

&lt;p&gt;That was the main idea behind our project.&lt;/p&gt;

&lt;p&gt;We wanted to build an incident-response agent that doesn't just look at the current error and generate a response. Instead, it can also look at previous incidents and use that experience when dealing with a new one.&lt;/p&gt;

&lt;p&gt;For this, we used Hindsight as the memory layer, Groq for the language model, and Streamlit for the interface.&lt;/p&gt;

&lt;p&gt;The overall idea is pretty simple:&lt;/p&gt;

&lt;p&gt;New incident → Check previous incidents → Generate a response → Record the result → Use it in the future&lt;/p&gt;

&lt;p&gt;Why we wanted the agent to have memory&lt;/p&gt;

&lt;p&gt;Imagine that an application suddenly starts showing:&lt;/p&gt;

&lt;p&gt;Database connection timeout under heavy load&lt;/p&gt;

&lt;p&gt;An AI model can obviously suggest a few possible solutions. It might recommend checking the database, increasing the connection pool, restarting a service, or looking at network issues.&lt;/p&gt;

&lt;p&gt;But what if the team has already dealt with this?&lt;/p&gt;

&lt;p&gt;Suppose they previously found that increasing the connection pool solved the issue, while restarting the service didn't make any difference.&lt;/p&gt;

&lt;p&gt;A normal AI conversation doesn't automatically know about that previous experience.&lt;/p&gt;

&lt;p&gt;That's where we thought memory could make the agent more useful.&lt;/p&gt;

&lt;p&gt;Instead of treating every incident as a completely new problem, the agent can search through previous incidents and bring relevant ones into the current conversation.&lt;/p&gt;

&lt;p&gt;Our Incident Command Center&lt;/p&gt;

&lt;p&gt;We built a Streamlit interface called Incident Command Center.&lt;/p&gt;

&lt;p&gt;The user can enter an incident or error log and provide information such as its severity and the service or component involved.&lt;/p&gt;

&lt;p&gt;The system then analyzes the incident and looks for relevant information from its memory.&lt;/p&gt;

&lt;p&gt;[ADD YOUR INCIDENT COMMAND CENTER SCREENSHOT HERE]&lt;/p&gt;

&lt;p&gt;Our Incident Command Center where a new incident can be entered and analyzed.&lt;/p&gt;

&lt;p&gt;The interface also shows whether the AI engine and memory engine are ready.&lt;/p&gt;

&lt;p&gt;The three main parts of the application are:&lt;/p&gt;

&lt;p&gt;Streamlit for the interface&lt;br&gt;
Hindsight for remembering previous incidents&lt;br&gt;
Groq for generating the AI response&lt;br&gt;
What we wanted the agent to remember&lt;/p&gt;

&lt;p&gt;We didn't want to save only the error message.&lt;/p&gt;

&lt;p&gt;For each incident, the useful information includes the error itself, what we thought caused it, what fix we tried, and what happened after trying that fix.&lt;/p&gt;

&lt;p&gt;That last part is particularly useful.&lt;/p&gt;

&lt;p&gt;If a fix didn't work, that's still useful information. The next time a similar incident happens, the agent can know that the same approach was already tried.&lt;/p&gt;

&lt;p&gt;The code for saving an incident looks like this:&lt;/p&gt;

&lt;p&gt;def save_incident(log, root_cause, fix, outcome):&lt;br&gt;
    text = (&lt;br&gt;
        f"Error: {log}\n"&lt;br&gt;
        f"Root cause: {root_cause}\n"&lt;br&gt;
        f"Fix: {fix}\n"&lt;br&gt;
        f"Outcome: {outcome}"&lt;br&gt;
    )&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;response = requests.post(
    f"{BASE_URL}/spaces/{SPACE_ID}/retain",
    headers={"Authorization": f"Bearer {HINDSIGHT_KEY}"},
    json={"content": text}
)

return response.json()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Finding something similar from the past&lt;/p&gt;

&lt;p&gt;When a new incident comes in, we can ask Hindsight to find relevant memories.&lt;/p&gt;

&lt;p&gt;The basic idea is that the current error becomes the search query, and Hindsight returns related incidents that the agent can use.&lt;/p&gt;

&lt;p&gt;def find_similar(new_log):&lt;br&gt;
    response = requests.post(&lt;br&gt;
        f"{BASE_URL}/spaces/{SPACE_ID}/recall",&lt;br&gt;
        headers={"Authorization": f"Bearer {HINDSIGHT_KEY}"},&lt;br&gt;
        json={"query": new_log}&lt;br&gt;
    )&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;return response.json()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;This gives the agent some history to work with instead of making it start from scratch every time.&lt;/p&gt;

&lt;p&gt;Giving the AI that history&lt;/p&gt;

&lt;p&gt;This is where the memory becomes part of the actual response.&lt;/p&gt;

&lt;p&gt;After finding similar incidents, we include them along with the new incident when creating the prompt for the language model.&lt;/p&gt;

&lt;p&gt;prompt = f"""&lt;br&gt;
You are an on-call engineering assistant.&lt;/p&gt;

&lt;p&gt;New incident: {log}&lt;/p&gt;

&lt;p&gt;Similar past incidents from memory:&lt;/p&gt;

&lt;p&gt;{similar}&lt;/p&gt;

&lt;p&gt;Based on past incidents, suggest the likely root cause and fix.&lt;/p&gt;

&lt;p&gt;If a past fix failed, do not suggest it again.&lt;br&gt;
"""&lt;/p&gt;

&lt;p&gt;So rather than simply asking the model what it thinks the solution should be, we're giving it some context about what happened previously.&lt;/p&gt;

&lt;p&gt;That doesn't mean the previous solution will always be correct for the new incident. It simply gives the model more information to consider.&lt;/p&gt;

&lt;p&gt;What changes when the agent has memory?&lt;/p&gt;

&lt;p&gt;Here's a simple example.&lt;/p&gt;

&lt;p&gt;Without memory&lt;/p&gt;

&lt;p&gt;An engineer enters:&lt;/p&gt;

&lt;p&gt;Database connection timeout under heavy load&lt;/p&gt;

&lt;p&gt;The AI looks at the current error and suggests a possible solution.&lt;/p&gt;

&lt;p&gt;It has no information about what the team tried during previous incidents.&lt;/p&gt;

&lt;p&gt;With memory&lt;/p&gt;

&lt;p&gt;The same type of incident happens again.&lt;/p&gt;

&lt;p&gt;This time, the agent can retrieve something like:&lt;/p&gt;

&lt;p&gt;Error: Database connection timeout after 500 requests&lt;br&gt;
Root cause: Connection pool exhausted&lt;br&gt;
Fix: Increased pool size from 10 to 50&lt;br&gt;
Outcome: Worked&lt;/p&gt;

&lt;p&gt;Now the response has some real history behind it.&lt;/p&gt;

&lt;p&gt;It might also retrieve an earlier attempt that failed:&lt;/p&gt;

&lt;p&gt;Fix: Restarted database service&lt;br&gt;
Outcome: Failed&lt;/p&gt;

&lt;p&gt;That gives the agent another useful piece of information: something was already tried and didn't solve the problem.&lt;/p&gt;

&lt;p&gt;That's the main difference we wanted to show.&lt;/p&gt;

&lt;p&gt;Why failed fixes are useful too&lt;/p&gt;

&lt;p&gt;We usually think of remembering successful solutions as the useful part.&lt;/p&gt;

&lt;p&gt;But while working on this idea, one thing that stood out was that failed attempts can be just as valuable.&lt;/p&gt;

&lt;p&gt;If an engineer already tried restarting a service and it didn't solve the issue, there isn't much value in suggesting exactly the same thing again without a reason.&lt;/p&gt;

&lt;p&gt;So instead of only remembering:&lt;/p&gt;

&lt;p&gt;"This fix worked."&lt;/p&gt;

&lt;p&gt;the agent can also remember:&lt;/p&gt;

&lt;p&gt;"This fix was tried, and it didn't work."&lt;/p&gt;

&lt;p&gt;That creates a simple feedback loop:&lt;/p&gt;

&lt;p&gt;Incident → Suggested fix → Worked/Failed → Memory → Future incident&lt;/p&gt;

&lt;p&gt;Over time, the memory becomes a record of previous troubleshooting experiences.&lt;/p&gt;

&lt;p&gt;Giving the memory some initial data&lt;/p&gt;

&lt;p&gt;Of course, if the memory starts completely empty, there isn't much for the agent to recall.&lt;/p&gt;

&lt;p&gt;So for testing, we use example incidents covering different types of problems, such as database issues, API timeouts, memory leaks, deployment failures, and third-party service outages.&lt;/p&gt;

&lt;p&gt;This gives the agent some initial history.&lt;/p&gt;

&lt;p&gt;Then we can enter a new incident through the Streamlit interface and see whether relevant previous incidents are retrieved.&lt;/p&gt;

&lt;p&gt;The basic workflow is:&lt;/p&gt;

&lt;p&gt;Enter the incident.&lt;br&gt;
Select the severity and service information.&lt;br&gt;
Analyze the incident.&lt;br&gt;
Retrieve relevant memories.&lt;br&gt;
Generate a suggested response.&lt;br&gt;
Record whether the suggested fix worked or failed.&lt;br&gt;
What I learned from building this&lt;br&gt;
Memory needs context&lt;/p&gt;

&lt;p&gt;One thing that became clear is that simply saving an error message isn't enough.&lt;/p&gt;

&lt;p&gt;The error becomes much more useful when it's connected to the root cause, the attempted fix, and the outcome.&lt;/p&gt;

&lt;p&gt;That gives the agent more information to work with when it encounters something similar later.&lt;/p&gt;

&lt;p&gt;A failed solution isn't wasted information&lt;/p&gt;

&lt;p&gt;It's easy to think that only successful fixes are worth remembering.&lt;/p&gt;

&lt;p&gt;But during incident response, knowing what didn't work can save time too.&lt;/p&gt;

&lt;p&gt;It can stop the same approach from being blindly repeated.&lt;/p&gt;

&lt;p&gt;Memory works better when it's part of the workflow&lt;/p&gt;

&lt;p&gt;We didn't want memory to be something that users had to manually search through whenever an incident happened.&lt;/p&gt;

&lt;p&gt;Instead, it is part of the response process:&lt;/p&gt;

&lt;p&gt;Recall → Analyze → Respond → Record&lt;/p&gt;

&lt;p&gt;That makes the memory useful at the point where the agent is actually making a suggestion.&lt;/p&gt;

&lt;p&gt;Similar incidents aren't necessarily identical&lt;/p&gt;

&lt;p&gt;There's also an important limitation.&lt;/p&gt;

&lt;p&gt;Two incidents can look similar but have completely different causes.&lt;/p&gt;

&lt;p&gt;So previous incidents should be treated as context, not as a guaranteed answer.&lt;/p&gt;

&lt;p&gt;The suggested fix still needs to be checked against the actual system.&lt;/p&gt;

&lt;p&gt;The main idea&lt;/p&gt;

&lt;p&gt;For me, the interesting part of this project isn't just getting an AI model to explain an error.&lt;/p&gt;

&lt;p&gt;It's giving the agent some experience to work with.&lt;/p&gt;

&lt;p&gt;An incident happens.&lt;br&gt;
The agent looks at what happened before.&lt;br&gt;
It suggests a response.&lt;br&gt;
The result gets recorded.&lt;br&gt;
And that experience can be useful the next time something similar happens.&lt;/p&gt;

&lt;p&gt;That's what we wanted to explore with Hindsight.&lt;/p&gt;

&lt;p&gt;Instead of an incident-response agent that starts from zero every time, we wanted one that can remember:&lt;/p&gt;

&lt;p&gt;what happened, what was tried, and what actually worked.&lt;/p&gt;

&lt;p&gt;That small change makes the idea of an AI incident-response agent much more interesting&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flr1dxcf94rjhu1juclfm.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flr1dxcf94rjhu1juclfm.jpeg" alt=" " width="799" height="372"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fao9b2n641xvkvtfipsbo.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fao9b2n641xvkvtfipsbo.jpeg" alt=" " width="800" height="512"&gt;&lt;/a&gt;# I Built an AI Incident Agent with Hindsight That Remembers What Failed&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
