<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dharani Kamsali</title>
    <description>The latest articles on DEV Community by Dharani Kamsali (@dharani_kamsali_d42811d68).</description>
    <link>https://dev.to/dharani_kamsali_d42811d68</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4148947%2Fbc5ff52a-3735-4ec0-92d7-85bd5b5acc80.png</url>
      <title>DEV Community: Dharani Kamsali</title>
      <link>https://dev.to/dharani_kamsali_d42811d68</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dharani_kamsali_d42811d68"/>
    <language>en</language>
    <item>
      <title>How Hindsight Changed My Approach to Incident Retrieval</title>
      <dc:creator>Dharani Kamsali</dc:creator>
      <pubDate>Tue, 29 Sep 2026 09:32:37 +0000</pubDate>
      <link>https://dev.to/dharani_kamsali_d42811d68/how-hindsight-changed-my-approach-to-incident-retrieval-2fca</link>
      <guid>https://dev.to/dharani_kamsali_d42811d68/how-hindsight-changed-my-approach-to-incident-retrieval-2fca</guid>
      <description>&lt;h1&gt;
  
  
  Building an Incident Response Agent That Learns From the Past
&lt;/h1&gt;

&lt;p&gt;An incident starts with a familiar pattern.&lt;/p&gt;

&lt;p&gt;A service becomes slow. Requests begin failing. Error rates increase. Engineers open dashboards, check logs, investigate recent changes, and try to understand what went wrong.&lt;/p&gt;

&lt;p&gt;Eventually, the root cause is identified and the incident is resolved.&lt;/p&gt;

&lt;p&gt;Then the same type of problem appears again a few weeks or months later.&lt;/p&gt;

&lt;p&gt;The team may have already solved something similar, but the useful details are often scattered across old postmortems, tickets, documentation, and incident notes.&lt;/p&gt;

&lt;p&gt;That made me think about a different approach:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if an incident-response agent could remember previous incidents and bring that experience into the next investigation?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That question became the foundation for &lt;strong&gt;Incident-Memory-Copilot&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The project combines an LLM-based reasoning layer with persistent memory using &lt;strong&gt;Hindsight&lt;/strong&gt;. Instead of asking an AI model to troubleshoot every incident from scratch, the system first searches for relevant previous experiences and then uses those memories as additional context.&lt;/p&gt;

&lt;p&gt;The workflow is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recall → Reason → Resolve → Retain&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The important part isn't simply generating an answer.&lt;/p&gt;

&lt;p&gt;The goal is to make organizational incident knowledge available at the moment an engineer needs it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw53ok1238elrjsnd0lbf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw53ok1238elrjsnd0lbf.png" alt=" " width="800" height="350"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Incident-Memory-Copilot interface for investigating production incidents with historical memory&lt;/em&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  The problem: incident knowledge is easy to lose
&lt;/h2&gt;

&lt;p&gt;Consider an API that starts returning intermittent &lt;code&gt;502 Bad Gateway&lt;/code&gt; errors.&lt;/p&gt;

&lt;p&gt;A general-purpose AI assistant might suggest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Check application logs.&lt;/li&gt;
&lt;li&gt;Inspect upstream dependencies.&lt;/li&gt;
&lt;li&gt;Review recent deployments.&lt;/li&gt;
&lt;li&gt;Check network connectivity.&lt;/li&gt;
&lt;li&gt;Examine CPU and memory usage.&lt;/li&gt;
&lt;li&gt;Investigate connection pools.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are reasonable starting points.&lt;/p&gt;

&lt;p&gt;But they are still generic.&lt;/p&gt;

&lt;p&gt;Now imagine that the same service had experienced a similar problem six months earlier. During that incident, the team discovered that connection-pool exhaustion was responsible, and one attempted configuration change actually made the problem worse.&lt;/p&gt;

&lt;p&gt;That historical information could significantly change how the new incident is investigated.&lt;/p&gt;

&lt;p&gt;The problem is therefore not just:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Can an AI generate troubleshooting suggestions?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Can the system find the organization's previous experience that is relevant to this incident?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is where persistent memory becomes important.&lt;/p&gt;
&lt;h2&gt;
  
  
  Designing the Incident-Memory-Copilot
&lt;/h2&gt;

&lt;p&gt;I designed the system around a continuous incident lifecycle:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             CURRENT INCIDENT
                    │
                    ▼
             Memory Recall
                    │
                    ▼
        Relevant Past Incidents
                    │
                    ▼
              AI Reasoning
                    │
                    ▼
          Investigation Plan
                    │
                    ▼
               Resolution
                    │
                    ▼
              Postmortem
                    │
                    ▼
             Memory Retention
                    │
                    └──────────────► Future Incident
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When an engineer submits an incident, the system extracts the important information and searches Hindsight for related memories.&lt;/p&gt;

&lt;p&gt;The retrieved information is then provided to the reasoning stage.&lt;/p&gt;

&lt;p&gt;After the incident has been resolved, the important lessons from that incident can be stored back into memory.&lt;/p&gt;

&lt;p&gt;This creates a continuous learning loop instead of treating every incident as an independent question.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first step is recall
&lt;/h2&gt;

&lt;p&gt;One of the main architectural decisions was to make memory retrieval happen &lt;strong&gt;before&lt;/strong&gt; the AI generates its investigation.&lt;/p&gt;

&lt;p&gt;A simple implementation could send the incident directly to an LLM:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident
   ↓
LLM
   ↓
Troubleshooting Suggestions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Instead, Incident-Memory-Copilot uses:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident
   ↓
Extract Symptoms
   ↓
Search Hindsight
   ↓
Retrieve Related Experiences
   ↓
Provide Historical Context
   ↓
LLM Reasoning
   ↓
Investigation Recommendation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This small architectural difference changes the type of information available to the model.&lt;/p&gt;

&lt;p&gt;The model is no longer working only with the current symptoms.&lt;/p&gt;

&lt;p&gt;It can also consider what happened during previous incidents with similar characteristics.&lt;/p&gt;

&lt;p&gt;Hindsight acts as the persistent memory layer that allows the system to retain and recall this information.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj8h7zz6d11qv3ako5bdw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj8h7zz6d11qv3ako5bdw.png" alt=" " width="800" height="347"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Hindsight recall retrieves relevant historical incident experience before the copilot reasons about the current incident.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  What gets stored in memory?
&lt;/h2&gt;

&lt;p&gt;The system doesn't need to remember every piece of text associated with an incident.&lt;/p&gt;

&lt;p&gt;The useful information is the experience gained from the incident.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Service:
Payment API

Symptoms:
Intermittent 502 errors during high traffic

Root Cause:
Connection-pool exhaustion

Successful Resolution:
Increased pool capacity and corrected connection handling

Failed Attempts:
Restarting the service temporarily reduced errors but did not solve the underlying problem

Lesson:
Check connection-pool utilization before repeatedly restarting instances
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This kind of information is much more useful for future investigations than simply storing a generic troubleshooting document.&lt;/p&gt;

&lt;p&gt;The memory can contain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Affected service&lt;/li&gt;
&lt;li&gt;Incident symptoms&lt;/li&gt;
&lt;li&gt;Root cause&lt;/li&gt;
&lt;li&gt;Investigation steps&lt;/li&gt;
&lt;li&gt;Successful actions&lt;/li&gt;
&lt;li&gt;Failed actions&lt;/li&gt;
&lt;li&gt;Resolution&lt;/li&gt;
&lt;li&gt;Lessons learned&lt;/li&gt;
&lt;li&gt;Relevant operational context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After the incident is completed, these details can be retained in Hindsight for future retrieval.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real test is the second incident
&lt;/h2&gt;

&lt;p&gt;A memory system isn't particularly interesting if we only demonstrate it with the first incident.&lt;/p&gt;

&lt;p&gt;The more meaningful scenario is what happens when another incident occurs later.&lt;/p&gt;

&lt;p&gt;Suppose the system previously learned about a payment-service outage caused by connection exhaustion.&lt;/p&gt;

&lt;p&gt;Several months later, a new incident arrives:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Checkout requests are intermittently failing.

Gateway errors are increasing during
a period of high traffic.

Latency has increased and upstream
connection failures are being observed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The wording is different.&lt;/p&gt;

&lt;p&gt;The incident is not a copy of the previous one.&lt;/p&gt;

&lt;p&gt;However, the symptoms may still be related.&lt;/p&gt;

&lt;p&gt;The system searches its memory and retrieves the earlier incident experience.&lt;/p&gt;

&lt;p&gt;The reasoning process can then combine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Current Incident
       +
Historical Experience
       ↓
    Reasoning
       ↓
Investigation Plan
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is where persistent memory provides value.&lt;/p&gt;

&lt;p&gt;The LLM does not have to somehow remember an incident from a previous conversation. The memory layer retrieves the relevant experience when it is needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory changes the lifecycle of the agent
&lt;/h2&gt;

&lt;p&gt;A normal chatbot mostly works within the boundaries of a conversation.&lt;/p&gt;

&lt;p&gt;An incident-memory agent has a much longer lifecycle:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident
   ↓
Investigation
   ↓
Resolution
   ↓
Postmortem
   ↓
Learning
   ↓
Memory
   ↓
Future Incident
   ↓
Recall
   ↓
Investigation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The output of one incident becomes potential input for another.&lt;/p&gt;

&lt;p&gt;Over time, the system can therefore accumulate operational experience.&lt;/p&gt;

&lt;p&gt;This is an important difference between a chatbot that simply generates responses and an agent workflow that maintains persistent memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recall is not the same as reasoning
&lt;/h2&gt;

&lt;p&gt;Another design consideration was separating memory retrieval from reasoning.&lt;/p&gt;

&lt;p&gt;The memory system answers:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What previous experiences might be relevant?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The reasoning layer answers:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What does that information mean for the incident happening now?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The workflow can therefore be represented as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Current Incident
       ↓
     Recall
       ↓
Relevant Memories
       ↓
   Reflection
       ↓
Historical Context
       ↓
  AI Reasoning
       ↓
Investigation Recommendation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Retrieved memories may contain several pieces of information.&lt;/p&gt;

&lt;p&gt;The reasoning stage determines how those pieces relate to the current incident.&lt;/p&gt;

&lt;p&gt;This separation also makes the architecture easier to reason about because memory retrieval and AI reasoning have different responsibilities.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned while building it
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. More memory doesn't automatically mean better results
&lt;/h3&gt;

&lt;p&gt;One of the first lessons from designing a memory-based agent is that simply adding more information isn't enough.&lt;/p&gt;

&lt;p&gt;The important question is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Did the system retrieve information that is actually relevant?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If unrelated incidents are retrieved, they can create noise and potentially distract the reasoning process.&lt;/p&gt;

&lt;p&gt;That means incident queries and memory structure matter just as much as the LLM prompt.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. The complete loop is more important than a single prompt
&lt;/h3&gt;

&lt;p&gt;It is relatively straightforward to create a prompt that produces an incident checklist.&lt;/p&gt;

&lt;p&gt;The more interesting engineering challenge is building the complete lifecycle:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Recall
  ↓
Reason
  ↓
Resolve
  ↓
Retain
  ↓
Recall Again
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The value of the system appears when information captured during one incident becomes useful during another.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Failed troubleshooting attempts are important
&lt;/h3&gt;

&lt;p&gt;Incident documentation often focuses heavily on the successful fix.&lt;/p&gt;

&lt;p&gt;But failed actions can also be valuable.&lt;/p&gt;

&lt;p&gt;For example, suppose engineers repeatedly restarted a service during an earlier incident. The restart temporarily reduced the errors but did not resolve the underlying issue.&lt;/p&gt;

&lt;p&gt;That information can prevent future engineers from immediately repeating the same approach.&lt;/p&gt;

&lt;p&gt;So the memory should capture not only:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What worked?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;but also:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What did we try that didn't work?"&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Memory capture should be part of incident resolution
&lt;/h3&gt;

&lt;p&gt;If engineers have to perform a completely separate knowledge-management process after every incident, important information can easily be missed.&lt;/p&gt;

&lt;p&gt;Instead, memory retention should naturally follow the incident and postmortem process.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident Resolved
       ↓
Postmortem Created
       ↓
Important Knowledge Identified
       ↓
Memory Retained
       ↓
Available for Future Incidents
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes memory a part of the operational workflow rather than another administrative task.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this project can be useful
&lt;/h2&gt;

&lt;p&gt;Incident-Memory-Copilot is primarily designed for teams that regularly deal with production incidents and need to reuse previous troubleshooting knowledge.&lt;/p&gt;

&lt;p&gt;It can be useful for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DevOps teams&lt;/li&gt;
&lt;li&gt;SRE teams&lt;/li&gt;
&lt;li&gt;Backend engineering teams&lt;/li&gt;
&lt;li&gt;Platform engineering teams&lt;/li&gt;
&lt;li&gt;Cloud operations teams&lt;/li&gt;
&lt;li&gt;Production support teams&lt;/li&gt;
&lt;li&gt;Organizations with large collections of historical postmortems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The main idea is not to replace engineers.&lt;/p&gt;

&lt;p&gt;Instead, the system helps engineers find relevant historical context faster.&lt;/p&gt;

&lt;p&gt;An engineer still evaluates the evidence, decides what actions are appropriate, and remains responsible for the actual resolution.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bigger idea
&lt;/h2&gt;

&lt;p&gt;Production incidents are rarely completely isolated events.&lt;/p&gt;

&lt;p&gt;Organizations often encounter related failures involving the same services, dependencies, infrastructure, deployment patterns, or configuration problems.&lt;/p&gt;

&lt;p&gt;Every resolved incident can therefore produce knowledge that may be useful later.&lt;/p&gt;

&lt;p&gt;The important questions are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What happened?

Why did it happen?

What fixed it?

What didn't work?

What should we remember?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Traditional documentation can answer these questions, but the information may remain passive until someone manually searches for it.&lt;/p&gt;

&lt;p&gt;A memory-enabled agent changes that interaction.&lt;/p&gt;

&lt;p&gt;Instead of expecting an engineer to remember every historical incident, the system can retrieve relevant experience when a new incident occurs.&lt;/p&gt;

&lt;p&gt;That is the idea behind &lt;strong&gt;Incident-Memory-Copilot&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It isn't intended to be an AI that magically knows every production problem.&lt;/p&gt;

&lt;p&gt;It isn't just another chatbot generating generic troubleshooting steps.&lt;/p&gt;

&lt;p&gt;It is a system designed around a simple lifecycle:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recall previous experience → reason about the current incident → support the investigation → retain what was learned.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The goal is simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The next incident should have the benefit of the incidents that came before it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Resources&lt;/p&gt;

&lt;p&gt;-&lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight GitHub&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;-&lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight Documentation&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;-&lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;Vectorize Agent Memory&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;-&lt;a href="https://github.com/kamsalidharani/incident-memory-copilot.git" rel="noopener noreferrer"&gt;Incident-Memory-Copilot&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>llm</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
