<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Shiva Kumar Kanneboina</title>
    <description>The latest articles on DEV Community by Shiva Kumar Kanneboina (@shivakumar_k).</description>
    <link>https://dev.to/shivakumar_k</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150204%2F87a41d89-48b8-43f9-a138-abb7dbc7f1dc.png</url>
      <title>DEV Community: Shiva Kumar Kanneboina</title>
      <link>https://dev.to/shivakumar_k</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shivakumar_k"/>
    <language>en</language>
    <item>
      <title>I Built an Incident-Response Agent That Remembers What Failed</title>
      <dc:creator>Shiva Kumar Kanneboina</dc:creator>
      <pubDate>Tue, 29 Sep 2026 16:07:26 +0000</pubDate>
      <link>https://dev.to/shivakumar_k/i-built-an-incident-response-agent-that-remembers-what-failed-3ong</link>
      <guid>https://dev.to/shivakumar_k/i-built-an-incident-response-agent-that-remembers-what-failed-3ong</guid>
      <description>&lt;h1&gt;
  
  
  I Built an Incident-Response Agent That Remembers What Failed
&lt;/h1&gt;

&lt;p&gt;Production incidents have a frustrating habit of repeating themselves.&lt;/p&gt;

&lt;p&gt;A database gets overloaded. An API starts timing out. A cache runs out of memory. Engineers investigate, try a few fixes, eventually resolve the problem, and write everything down.&lt;/p&gt;

&lt;p&gt;Then, weeks or months later, something similar happens again.&lt;/p&gt;

&lt;p&gt;The problem is not always a lack of documentation. The problem is that the useful knowledge from the previous incident is usually scattered across tickets, postmortems, dashboards, Slack conversations, and the memories of the engineers who handled it.&lt;/p&gt;

&lt;p&gt;So when the next incident arrives, the investigation often starts from scratch.&lt;/p&gt;

&lt;p&gt;I wanted to build something different: an incident-response agent that could remember what happened before and use that experience when investigating the next incident.&lt;/p&gt;

&lt;p&gt;That is how I built &lt;strong&gt;OpsMind&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The idea behind OpsMind
&lt;/h2&gt;

&lt;p&gt;OpsMind is an AI incident-intelligence agent designed to help engineers investigate production incidents using two sources of information:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What is happening right now&lt;/li&gt;
&lt;li&gt;What happened during similar incidents in the past&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The second part is what makes the system interesting.&lt;/p&gt;

&lt;p&gt;A typical AI assistant can look at the current symptoms and suggest possible causes. OpsMind can also ask a different question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Have we seen something like this before?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If the system remembers a previous incident involving connection-pool exhaustion, database pressure, failed mitigation attempts, and a successful resolution, that experience becomes part of the investigation.&lt;/p&gt;

&lt;p&gt;The agent is no longer starting with an empty context.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5gydbzdlcxl4osm5jn0m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5gydbzdlcxl4osm5jn0m.png" alt=" " width="800" height="374"&gt;&lt;/a&gt;&lt;br&gt;
When I first thought about incident memory, the obvious approach was to store summaries of previous incidents.&lt;/p&gt;

&lt;p&gt;But a summary like this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Database was overloaded. Increased the connection pool size. Incident resolved."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;doesn't contain enough information to be useful during the next incident.&lt;/p&gt;

&lt;p&gt;What caused the overload?&lt;/p&gt;

&lt;p&gt;What symptoms appeared first?&lt;/p&gt;

&lt;p&gt;Which service was involved?&lt;/p&gt;

&lt;p&gt;What did the engineer try?&lt;/p&gt;

&lt;p&gt;Which attempts failed?&lt;/p&gt;

&lt;p&gt;What actually fixed the problem?&lt;/p&gt;

&lt;p&gt;And perhaps most importantly, what should the next engineer avoid doing?&lt;/p&gt;

&lt;p&gt;That last question turned out to be particularly important.&lt;/p&gt;

&lt;p&gt;During an incident, a failed mitigation is still valuable information.&lt;/p&gt;

&lt;p&gt;If restarting a service did not solve the problem, that is useful.&lt;/p&gt;

&lt;p&gt;If adding more replicas made the underlying database pressure worse, that is useful.&lt;/p&gt;

&lt;p&gt;If clearing a cache increased database load, that is useful too.&lt;/p&gt;

&lt;p&gt;These are not just historical details. They are operational lessons.&lt;/p&gt;

&lt;p&gt;That led me to a simple principle for OpsMind:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Remember not only what worked, but also what failed and why.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Using Hindsight as the memory layer
&lt;/h2&gt;

&lt;p&gt;I integrated Hindsight as the persistent memory layer for OpsMind.&lt;/p&gt;

&lt;p&gt;The goal was to move beyond storing previous conversations or plain incident summaries. The agent needs to retain useful information from an incident and later recall relevant experiences when another incident occurs.&lt;/p&gt;

&lt;p&gt;The memory can capture information about facts, entities, relationships, temporal context, experiences, and lessons.&lt;/p&gt;

&lt;p&gt;For OpsMind, this means an incident can become more than a document sitting in a database. It can become part of a connected body of operational knowledge.&lt;/p&gt;

&lt;p&gt;The basic learning loop is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Incident → Recall → Investigate → Recommend → Retain the outcome&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;During an incident, OpsMind can recall relevant historical experiences.&lt;/p&gt;

&lt;p&gt;After the incident is resolved, the outcome can be retained so that the experience becomes useful the next time a similar problem appears.&lt;/p&gt;

&lt;p&gt;That creates a feedback loop rather than a stateless assistant.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2wb2w9qvvnvds8sb31vk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2wb2w9qvvnvds8sb31vk.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  A real example from OpsMind
&lt;/h2&gt;

&lt;p&gt;One of the incidents in my dataset involved a Payments API experiencing connection-pool saturation.&lt;/p&gt;

&lt;p&gt;The underlying problem was a leaked client caused by a missing &lt;code&gt;client.release()&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;What made this incident useful as training data was that it contained more than the final solution.&lt;/p&gt;

&lt;p&gt;A restart was attempted, but it did not address the underlying problem.&lt;/p&gt;

&lt;p&gt;Replica scaling was also attempted, but it was not the right solution. In this situation, increasing capacity could actually make database exhaustion worse.&lt;/p&gt;

&lt;p&gt;The eventual resolution was a rollback.&lt;/p&gt;

&lt;p&gt;So instead of remembering only:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Payments API had connection-pool saturation and was fixed by a rollback."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;OpsMind can retain the broader experience:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Symptoms
  ↓
Connection pool saturation
  ↓
Investigation
  ↓
Client leak identified
  ↓
Restart attempted
  ↓
Unsuccessful
  ↓
Replica scaling attempted
  ↓
Unsuccessful / potentially counterproductive
  ↓
Rollback
  ↓
Lesson learned
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That difference matters.&lt;/p&gt;

&lt;p&gt;If a similar incident happens again, the system has more context than just the final answer.&lt;/p&gt;

&lt;p&gt;It has the path that led to the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Similar symptoms do not always mean the same problem
&lt;/h2&gt;

&lt;p&gt;Another important part of the system is that historical memory should not be treated as a guaranteed answer.&lt;/p&gt;

&lt;p&gt;For example, connection-pool pressure can happen for very different reasons.&lt;/p&gt;

&lt;p&gt;In the OpsMind dataset, one incident involved a client leak. Another involved a database connection-pool configuration problem. Another involved a slow, unindexed database query that caused long-running operations and eventually contributed to pool starvation.&lt;/p&gt;

&lt;p&gt;The symptoms can look similar while the root causes are different.&lt;/p&gt;

&lt;p&gt;That is why I don't want OpsMind to simply retrieve an old incident and copy its solution.&lt;/p&gt;

&lt;p&gt;Instead, historical incidents are additional evidence.&lt;/p&gt;

&lt;p&gt;The agent can compare the current symptoms with previous experiences, look for similarities and differences, and then use that context when forming a recommendation.&lt;/p&gt;

&lt;p&gt;The goal is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Use the past to improve the investigation, not to replace the evidence from the present.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That distinction is important for any system that is going to be used around production infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The value of remembering failed fixes
&lt;/h2&gt;

&lt;p&gt;While building the incident dataset, I noticed a pattern that became one of my favorite parts of the project.&lt;/p&gt;

&lt;p&gt;A lot of operational knowledge is hidden inside failed attempts.&lt;/p&gt;

&lt;p&gt;For example, the dataset contains incidents where:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Restarting a service did not fix an underlying resource leak.&lt;/li&gt;
&lt;li&gt;Increasing capacity could delay or worsen database exhaustion.&lt;/li&gt;
&lt;li&gt;Killing idle database sessions did not address the underlying query behavior.&lt;/li&gt;
&lt;li&gt;Adding Kafka consumers could make a rebalance problem worse.&lt;/li&gt;
&lt;li&gt;Clearing a cache could increase load on the database.&lt;/li&gt;
&lt;li&gt;Increasing memory could delay an out-of-memory problem without fixing the underlying issue.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These lessons are easy to lose.&lt;/p&gt;

&lt;p&gt;A traditional incident summary tends to focus on the final resolution. But during a real incident, knowing what has already been tried can prevent engineers from repeating the same mistakes.&lt;/p&gt;

&lt;p&gt;That gives the memory system another useful question to answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What did we try last time that did not work?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For incident response, that can be just as important as asking what worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned while building OpsMind
&lt;/h2&gt;

&lt;p&gt;The first thing I learned is that &lt;strong&gt;memory quality matters more than memory quantity&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It is easy to store a large amount of incident data.&lt;/p&gt;

&lt;p&gt;It is much harder to make that information useful.&lt;/p&gt;

&lt;p&gt;The memory needs to preserve the parts of an incident that actually matter during an investigation: symptoms, root causes, actions, outcomes, failed attempts, and lessons.&lt;/p&gt;

&lt;p&gt;The second lesson was that &lt;strong&gt;retrieval needs to stay grounded in the current incident&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Historical experience is useful context, but it should not become unquestioned truth.&lt;/p&gt;

&lt;p&gt;If current logs, metrics, or traces contradict something that happened in an older incident, the current evidence should take priority.&lt;/p&gt;

&lt;p&gt;The third lesson was that &lt;strong&gt;visualizing memory is surprisingly useful during development&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Looking at the Hindsight memory graph made it easier to understand how incident information was being connected and whether the retained data contained meaningful relationships.&lt;/p&gt;

&lt;p&gt;It also exposed data-quality issues.&lt;/p&gt;

&lt;p&gt;For example, duplicate or test memories can introduce unnecessary noise. Cleaning those up became part of preparing the system for a more realistic demonstration.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the learning loop works
&lt;/h2&gt;

&lt;p&gt;I think about OpsMind's memory system in four stages.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Retain
&lt;/h3&gt;

&lt;p&gt;After an incident, the system retains the useful operational knowledge.&lt;/p&gt;

&lt;p&gt;That includes what happened, what was discovered, what actions were attempted, what worked, what failed, and what was learned.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Recall
&lt;/h3&gt;

&lt;p&gt;When a new incident arrives, OpsMind looks for relevant historical experiences.&lt;/p&gt;

&lt;p&gt;The goal isn't to retrieve everything the system knows.&lt;/p&gt;

&lt;p&gt;The goal is to find experiences that can actually help with the current investigation.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Reflect
&lt;/h3&gt;

&lt;p&gt;The recalled experiences become context for reasoning.&lt;/p&gt;

&lt;p&gt;The agent can compare the current symptoms with previous incidents and consider both the similarities and the differences.&lt;/p&gt;

&lt;p&gt;This is where memory becomes more than search.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Retain again
&lt;/h3&gt;

&lt;p&gt;Once the incident is resolved, the new outcome becomes another experience.&lt;/p&gt;

&lt;p&gt;Over time, the system can build a growing collection of operational knowledge.&lt;/p&gt;

&lt;p&gt;That is the part that turns the architecture from a chatbot into a learning incident assistant.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would improve next
&lt;/h2&gt;

&lt;p&gt;OpsMind is still a prototype, and there are several areas I would improve before treating it as a production incident-response system.&lt;/p&gt;

&lt;p&gt;The first is evaluation.&lt;/p&gt;

&lt;p&gt;It is not enough to demonstrate that the agent can retrieve a related incident. I want to measure whether the retrieved memory actually improves investigation quality, helps identify likely root causes faster, and reduces repeated failed mitigations.&lt;/p&gt;

&lt;p&gt;The second is better evidence handling.&lt;/p&gt;

&lt;p&gt;A recommendation should clearly distinguish between:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Current incident evidence&lt;/li&gt;
&lt;li&gt;Historical experience&lt;/li&gt;
&lt;li&gt;Agent inference&lt;/li&gt;
&lt;li&gt;Recommended action&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That makes it easier for an engineer to understand why the agent reached a particular conclusion.&lt;/p&gt;

&lt;p&gt;The third is deeper integration with real operational systems.&lt;/p&gt;

&lt;p&gt;A production version could connect directly to logs, metrics, traces, deployment history, incident-management systems, and runbooks.&lt;/p&gt;

&lt;p&gt;I would also keep human approval in the loop for actions that can change production infrastructure.&lt;/p&gt;

&lt;p&gt;The goal isn't to build an agent that blindly executes changes.&lt;/p&gt;

&lt;p&gt;The goal is to build an agent that gives engineers better context when they need to make decisions quickly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why persistent memory matters
&lt;/h2&gt;

&lt;p&gt;The more I worked on OpsMind, the more I realized that incident response is fundamentally a knowledge problem.&lt;/p&gt;

&lt;p&gt;Organizations accumulate years of operational experience.&lt;/p&gt;

&lt;p&gt;Engineers learn which failures happen repeatedly.&lt;/p&gt;

&lt;p&gt;They learn which fixes are safe.&lt;/p&gt;

&lt;p&gt;They learn which seemingly obvious fixes can make a situation worse.&lt;/p&gt;

&lt;p&gt;They learn patterns that never quite make it into a formal runbook.&lt;/p&gt;

&lt;p&gt;But a stateless AI agent cannot reliably carry that experience from one incident to the next.&lt;/p&gt;

&lt;p&gt;Persistent memory changes that.&lt;/p&gt;

&lt;p&gt;With Hindsight, OpsMind can retain operational experiences and bring relevant ones back into a future investigation.&lt;/p&gt;

&lt;p&gt;The result is not simply an AI that answers:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What should I do?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It can also ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What happened the last time we saw something like this?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And, perhaps more importantly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What did we try last time, what failed, and what did we learn?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the behavior I wanted to build.&lt;/p&gt;

&lt;p&gt;Because when the next production incident happens, the investigation should not have to start from a blank context window.&lt;/p&gt;

&lt;p&gt;It should start with everything the system has already learned.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Project:&lt;/strong&gt; OpsMind&lt;br&gt;
&lt;strong&gt;Memory layer:&lt;/strong&gt; Hindsight&lt;br&gt;
&lt;strong&gt;Core loop:&lt;/strong&gt; Retain → Recall → Reflect → Retain&lt;br&gt;
&lt;strong&gt;Focus:&lt;/strong&gt; Persistent incident intelligence and learning from operational experience&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>hindsight</category>
    </item>
  </channel>
</rss>
