<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dikshitha Kasoju</title>
    <description>The latest articles on DEV Community by Dikshitha Kasoju (@dikshithakasoju).</description>
    <link>https://dev.to/dikshithakasoju</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147776%2F4b785dbe-ee4e-4592-b41a-69436bdeec5d.png</url>
      <title>DEV Community: Dikshitha Kasoju</title>
      <link>https://dev.to/dikshithakasoju</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dikshithakasoju"/>
    <language>en</language>
    <item>
      <title>What Changes When an Incident Agent Can Remember?</title>
      <dc:creator>Dikshitha Kasoju</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:09:17 +0000</pubDate>
      <link>https://dev.to/dikshithakasoju/what-changes-when-an-incident-agent-can-remember-1426</link>
      <guid>https://dev.to/dikshithakasoju/what-changes-when-an-incident-agent-can-remember-1426</guid>
      <description>&lt;p&gt;An AI assistant can analyze logs. It can summarize an incident. It can even suggest a possible root cause.&lt;/p&gt;

&lt;p&gt;But there is a problem with starting every investigation from zero.&lt;/p&gt;

&lt;p&gt;If the same class of failure happens again next month, the assistant may have no idea that the team already solved something similar.&lt;/p&gt;

&lt;p&gt;That was the part of incident response I found most interesting while working on &lt;strong&gt;IncidentMind&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The goal wasn't just to build an AI assistant that could investigate production incidents. The more interesting goal was to give the investigation process a memory layer so that a resolved incident could become useful context for a future one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Incident → Investigation → Resolution → Memory → Future Investigation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And &lt;strong&gt;Hindsight&lt;/strong&gt; is what makes the last two steps possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem With Starting From Zero
&lt;/h2&gt;

&lt;p&gt;Production incidents rarely arrive as clean, isolated problems.&lt;/p&gt;

&lt;p&gt;An authentication failure might involve:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a recent deployment&lt;/li&gt;
&lt;li&gt;configuration changes&lt;/li&gt;
&lt;li&gt;multiple service instances&lt;/li&gt;
&lt;li&gt;logs and traces&lt;/li&gt;
&lt;li&gt;external dependencies&lt;/li&gt;
&lt;li&gt;user-facing symptoms&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An engineer has to connect all of those pieces before deciding what is actually happening.&lt;/p&gt;

&lt;p&gt;AI can help organize that information, but there is still a missing piece:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What happened the last time we saw something similar?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A traditional application database can tell us that an incident existed.&lt;/p&gt;

&lt;p&gt;It can tell us its ID, status, service, severity, and other application data.&lt;/p&gt;

&lt;p&gt;But operational experience is different.&lt;/p&gt;

&lt;p&gt;The useful information might be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This type of authentication failure was previously caused by inconsistent signing-key versions across service instances.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is not just application state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is experience.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That distinction became the central design idea behind IncidentMind.&lt;/p&gt;

&lt;h2&gt;
  
  
  What IncidentMind Actually Remembers
&lt;/h2&gt;

&lt;p&gt;IncidentMind has two different persistence responsibilities.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SQLite&lt;/strong&gt; handles application state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hindsight&lt;/strong&gt; handles retained incident experience.&lt;/p&gt;

&lt;p&gt;That means the application can keep normal records such as incident IDs and statuses in SQLite while using Hindsight to retain information that could help with future investigations.&lt;/p&gt;

&lt;p&gt;A resolved incident can contribute information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Confirmed root cause&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resolution steps&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Runbook&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lessons learned&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prevention steps&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Service and severity context&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important part is that memory is created from a &lt;strong&gt;resolved incident&lt;/strong&gt; rather than treating every piece of incoming information as something worth remembering.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7lbp2bp969z6427hdb5y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7lbp2bp969z6427hdb5y.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;IncidentMind separates current incident state from retained incident experience.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A Production-Style Failure
&lt;/h2&gt;

&lt;p&gt;To test the workflow, we used an authentication incident.&lt;/p&gt;

&lt;p&gt;The service was an &lt;strong&gt;Authentication API&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A deployment introduced JWT validation middleware and changed the authentication service's signing-key configuration to use a centralized secrets provider.&lt;/p&gt;

&lt;p&gt;After the deployment, previously authenticated users were intermittently logged out, while new authentication attempts returned HTTP 401 errors.&lt;/p&gt;

&lt;p&gt;The logs contained a particularly useful clue:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ERROR auth-api
JWT validation failed:
InvalidSignatureError: signature verification failed

WARN auth-api
POST /auth/refresh returned status=401

WARN auth-api
Instance=auth-api-7c8d9 signing_key_version=v2

WARN auth-api
Instance=auth-api-5f2a1 signing_key_version=v1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The interesting detail is that the database itself was healthy.&lt;/p&gt;

&lt;p&gt;The problem was distributed across the authentication instances.&lt;/p&gt;

&lt;p&gt;One instance was using one signing-key version while another was using a different version.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The incident logs contain the clues needed to connect the authentication failures with inconsistent signing-key configuration.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Giving the Agent the Full Incident Context
&lt;/h2&gt;

&lt;p&gt;The next step is submitting the incident to IncidentMind.&lt;/p&gt;

&lt;p&gt;Instead of sending the model only an error message, the workflow provides broader context:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;affected service&lt;/li&gt;
&lt;li&gt;incident ID&lt;/li&gt;
&lt;li&gt;severity&lt;/li&gt;
&lt;li&gt;recent deployment or configuration changes&lt;/li&gt;
&lt;li&gt;symptoms&lt;/li&gt;
&lt;li&gt;error logs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That gives the investigation agent more than a single error string to reason about.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftvn3cpup6u2d6nrs5w4r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftvn3cpup6u2d6nrs5w4r.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;IncidentMind collects the surrounding context before triggering the investigation.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The agent then produces a structured investigation.&lt;/p&gt;

&lt;p&gt;In this example, the investigation connected the authentication failures with the inconsistent signing-key versions.&lt;/p&gt;

&lt;p&gt;The important point is that the model isn't replacing the engineer.&lt;/p&gt;

&lt;p&gt;The engineer still needs to verify the diagnosis.&lt;/p&gt;

&lt;p&gt;The agent is helping organize the available evidence and surface a likely explanation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhbekgrbd6ne7p1txsjgv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhbekgrbd6ne7p1txsjgv.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The investigation view combines the incident context with the AI-generated investigation.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Part That Changes Everything: The Resolution Becomes Memory
&lt;/h2&gt;

&lt;p&gt;Finding a root cause is not the end of the workflow.&lt;/p&gt;

&lt;p&gt;Once the cause was confirmed, the affected authentication instances were brought onto the same signing-key version, the affected instances were restarted, and authentication flows were verified again.&lt;/p&gt;

&lt;p&gt;We also recorded the operational knowledge around the fix.&lt;/p&gt;

&lt;p&gt;The runbook was:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RB-AUTH-KEY-08 — JWT Signing-Key Synchronization Failure&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The lesson was that authentication deployments should verify signing-key consistency across instances and include automated configuration checks and controlled key rotation.&lt;/p&gt;

&lt;p&gt;This is the information that should survive beyond the individual incident.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3rsczyxhnhkege7xb1j4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3rsczyxhnhkege7xb1j4.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The resolved incident is converted into reusable operational experience.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How Hindsight Fits Into the Workflow
&lt;/h2&gt;

&lt;p&gt;The Hindsight integration is deliberately small.&lt;/p&gt;

&lt;p&gt;When the incident has been resolved, IncidentMind retains the completed incident knowledge:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;bank&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;memory_payload&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;severity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;incident_resolution&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important detail here isn't the number of lines of code.&lt;/p&gt;

&lt;p&gt;It is &lt;strong&gt;what those lines represent&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The application is taking an outcome that was previously trapped inside one incident and making it available to future investigations.&lt;/p&gt;

&lt;p&gt;Later, when another incident occurs, IncidentMind can query Hindsight for relevant historical experience:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;recall_resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;bank&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;budget&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That creates a very different investigation flow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Without memory:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Current incident → AI investigation&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With memory:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Current incident + relevant past experience → AI investigation&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5l6ai26r47rc5u1p5q6t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5l6ai26r47rc5u1p5q6t.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The retain and recall operations form the memory layer of the investigation workflow.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Application State and Memory Are Different
&lt;/h2&gt;

&lt;p&gt;One design decision I found particularly useful was keeping application state and operational memory separate.&lt;/p&gt;

&lt;p&gt;The application database can answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Which incidents exist?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hindsight can help answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Have we seen something like this before, and what did we learn?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those questions sound similar, but they serve different purposes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SQLite&lt;/strong&gt; gives the application predictable persistence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hindsight&lt;/strong&gt; gives the agent a mechanism for retaining and recalling experience.&lt;/p&gt;

&lt;p&gt;This separation also makes the architecture easier to reason about.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbgnkqi10qvqww2wwizsi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbgnkqi10qvqww2wwizsi.png" alt=" " width="755" height="485"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;IncidentMind separates application state in SQLite from operational memory in Hindsight.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The important architectural boundary is between &lt;strong&gt;application state&lt;/strong&gt; and &lt;strong&gt;operational memory&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making Memory Visible
&lt;/h2&gt;

&lt;p&gt;There was another lesson that wasn't obvious at first.&lt;/p&gt;

&lt;p&gt;A memory operation can succeed in the backend without being obvious to the person using the application.&lt;/p&gt;

&lt;p&gt;That makes debugging difficult.&lt;/p&gt;

&lt;p&gt;If an engineer clicks a button to retain a resolved incident, they should be able to see that the operation actually contributed to the application's memory state.&lt;/p&gt;

&lt;p&gt;We therefore made the retained-memory state visible on the dashboard.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ts4bjyr89e0waubfv8m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ts4bjyr89e0waubfv8m.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The dashboard after successful retention shows the updated retained-memory count.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This sounds like a small UI detail, but it matters for systems involving memory.&lt;/p&gt;

&lt;p&gt;When persistence is part of the product's behavior, &lt;strong&gt;it should be observable&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Otherwise, it becomes difficult to tell whether the system actually remembered something or merely appeared to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Looking at the Code Behind the Memory Layer
&lt;/h2&gt;

&lt;p&gt;The Hindsight client is initialized separately from the rest of the application logic.&lt;/p&gt;

&lt;p&gt;The project also has a local fallback path when Hindsight credentials aren't configured, while the configured environment can use Hindsight Cloud.&lt;/p&gt;

&lt;p&gt;That made development easier because the rest of the application didn't have to completely stop working when the external memory service wasn't available.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe7he86j5bmhegx7yssjl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe7he86j5bmhegx7yssjl.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The memory module handles the Hindsight integration and the application's fallback behavior.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The important architectural boundary is that the rest of IncidentMind doesn't need to know every detail about how memory is stored.&lt;/p&gt;

&lt;p&gt;It can work with two simple operations:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;retain this experience&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;and&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;recall relevant experience&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That keeps the memory mechanism relatively isolated from the incident workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing the Workflow
&lt;/h2&gt;

&lt;p&gt;Another thing I didn't want was a system that only looked convincing during a manual demo.&lt;/p&gt;

&lt;p&gt;The project includes automated tests around important application behavior, including incident handling and the retained-memory dashboard counter.&lt;/p&gt;

&lt;p&gt;The final test suite passed with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;7 passed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Automated tests cover important application behavior, including the retained-memory flow.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This became particularly useful while changing the retention workflow.&lt;/p&gt;

&lt;p&gt;A memory-related change can affect both backend behavior and what the dashboard reports, so having tests around that behavior provides a safety net.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Learned About Useful Agent Memory
&lt;/h2&gt;

&lt;p&gt;The biggest lesson for me was that &lt;strong&gt;adding memory isn't simply adding a &lt;code&gt;retain()&lt;/code&gt; call.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There are at least three separate questions.&lt;/p&gt;

&lt;h3&gt;
  
  
  What should the system remember?
&lt;/h3&gt;

&lt;p&gt;For incident response, storing every raw log isn't necessarily the most useful form of memory.&lt;/p&gt;

&lt;p&gt;A confirmed root cause, the actual resolution, the runbook, and the lesson learned are much more useful operational knowledge.&lt;/p&gt;

&lt;h3&gt;
  
  
  When should something become memory?
&lt;/h3&gt;

&lt;p&gt;A submitted incident isn't necessarily a lesson.&lt;/p&gt;

&lt;p&gt;The incident becomes much more valuable after the root cause has been confirmed and the resolution is known.&lt;/p&gt;

&lt;p&gt;That is why the retention step belongs after resolution.&lt;/p&gt;

&lt;h3&gt;
  
  
  How should memory affect a new investigation?
&lt;/h3&gt;

&lt;p&gt;This is probably the most important question.&lt;/p&gt;

&lt;p&gt;A previous incident should provide &lt;strong&gt;context&lt;/strong&gt;, not automatically become the answer.&lt;/p&gt;

&lt;p&gt;If an old incident looks similar to a new one, the agent still needs to consider the evidence from the current incident.&lt;/p&gt;

&lt;p&gt;Historical memory should help the investigation, &lt;strong&gt;not replace it&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That distinction matters if an AI system is going to be useful in an engineering environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Improve Next
&lt;/h2&gt;

&lt;p&gt;There are several things I would explore from here.&lt;/p&gt;

&lt;p&gt;First, I would make recalled incidents more visible during the investigation itself.&lt;/p&gt;

&lt;p&gt;If a previous incident influenced the agent's reasoning, an engineer should be able to see which historical experience was relevant.&lt;/p&gt;

&lt;p&gt;Second, I would build richer relationships between incidents.&lt;/p&gt;

&lt;p&gt;For example, incidents could be connected through:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;service&lt;/li&gt;
&lt;li&gt;deployment&lt;/li&gt;
&lt;li&gt;dependency&lt;/li&gt;
&lt;li&gt;configuration&lt;/li&gt;
&lt;li&gt;failure pattern&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Third, I would measure memory quality instead of only measuring whether something was successfully stored.&lt;/p&gt;

&lt;p&gt;A memory system can retain information perfectly and still retrieve irrelevant information.&lt;/p&gt;

&lt;p&gt;So an important future metric is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Did the recalled experience actually help with this incident?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Finally, I would explore how retained runbooks and prevention steps could move the system beyond diagnosis and toward improving incident-response practices over time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bigger Idea
&lt;/h2&gt;

&lt;p&gt;The interesting part of IncidentMind isn't that an LLM can read an incident.&lt;/p&gt;

&lt;p&gt;The more interesting question is what happens when the system can carry useful experience from one incident into another.&lt;/p&gt;

&lt;p&gt;A normal incident workflow often looks like:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Detect → Investigate → Fix → Close&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The workflow we explored adds another step:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Detect → Investigate → Fix → Learn → Recall&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That extra step changes the role of the AI assistant.&lt;/p&gt;

&lt;p&gt;Instead of being useful only during the current incident, it can become part of a longer learning loop.&lt;/p&gt;

&lt;p&gt;The goal isn't to make the agent automatically correct.&lt;/p&gt;

&lt;p&gt;The goal is to make sure that when the team learns something valuable from an incident, that knowledge doesn't disappear when the incident is closed.&lt;/p&gt;

&lt;p&gt;That's the idea behind using Hindsight with IncidentMind:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Resolve the incident. Retain the lesson. Use it when the next one arrives.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Resources
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Hindsight GitHub:&lt;/strong&gt; &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;https://github.com/vectorize-io/hindsight&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hindsight Documentation:&lt;/strong&gt; &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;https://hindsight.vectorize.io/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What Is Agent Memory?:&lt;/strong&gt; &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;https://vectorize.io/what-is-agent-memory&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;IncidentMind:&lt;/strong&gt; &lt;a href="https://github.com/MRajeshwariReddy/incidentmind" rel="noopener noreferrer"&gt;https://github.com/MRajeshwariReddy/incidentmind&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>softwareengineering</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
