<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sai Madhuri</title>
    <description>The latest articles on DEV Community by Sai Madhuri (@sai_ed6de9697ae4832f5094b).</description>
    <link>https://dev.to/sai_ed6de9697ae4832f5094b</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4151868%2Feaee5d99-a33d-4239-896b-f0bc56558e86.png</url>
      <title>DEV Community: Sai Madhuri</title>
      <link>https://dev.to/sai_ed6de9697ae4832f5094b</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sai_ed6de9697ae4832f5094b"/>
    <language>en</language>
    <item>
      <title>I Built an Agent That Learns From Resolved Incidents</title>
      <dc:creator>Sai Madhuri</dc:creator>
      <pubDate>Thu, 01 Oct 2026 06:34:03 +0000</pubDate>
      <link>https://dev.to/sai_ed6de9697ae4832f5094b/i-built-an-agent-that-learns-from-resolved-incidents-44je</link>
      <guid>https://dev.to/sai_ed6de9697ae4832f5094b/i-built-an-agent-that-learns-from-resolved-incidents-44je</guid>
      <description>&lt;p&gt;I Gave Incident Response a Memory With Hindsight&lt;/p&gt;

&lt;h1&gt;
  
  
  ai #devops #sre #hindsight #agents #incidentresponse
&lt;/h1&gt;

&lt;p&gt;Production incidents rarely happen in isolation.&lt;/p&gt;

&lt;p&gt;A Payment API returns 502s. A database connection pool is exhausted. Someone investigates, finds the fix, and closes the incident. A few months later, something similar happens and another engineer starts the investigation from the beginning.&lt;/p&gt;

&lt;p&gt;I wanted to change that part of the workflow.&lt;/p&gt;

&lt;p&gt;I built Incident-Memory-Copilot, an incident-response agent that uses Hindsight as its persistent memory layer. The idea is simple: when an incident happens, the agent should be able to look at what the organization learned from previous incidents, use that context during investigation, and then retain the new lesson after the incident is resolved.&lt;/p&gt;

&lt;p&gt;The interesting part is not just recalling an old incident.&lt;/p&gt;

&lt;p&gt;It is closing the loop.&lt;/p&gt;

&lt;p&gt;The problem: incident knowledge gets lost&lt;/p&gt;

&lt;p&gt;During an outage, an engineer may have access to:&lt;/p&gt;

&lt;p&gt;Current incident details&lt;/p&gt;

&lt;p&gt;Logs and metrics&lt;/p&gt;

&lt;p&gt;Deployment information&lt;/p&gt;

&lt;p&gt;Runbooks&lt;/p&gt;

&lt;p&gt;Previous incident reports&lt;/p&gt;

&lt;p&gt;Postmortems&lt;/p&gt;

&lt;p&gt;Notes about what worked&lt;/p&gt;

&lt;p&gt;Notes about what failed&lt;/p&gt;

&lt;p&gt;The information exists, but the connection between the current incident and previous experience is usually the hard part.&lt;/p&gt;

&lt;p&gt;A normal LLM can look at an HTTP 502 and give a reasonable list of possible causes.&lt;/p&gt;

&lt;p&gt;But I wanted the agent to answer a more useful question:&lt;/p&gt;

&lt;p&gt;“Have we seen something like this before, and what did we learn from it?”&lt;/p&gt;

&lt;p&gt;That is where Hindsight became the memory layer.&lt;/p&gt;

&lt;p&gt;What I built&lt;/p&gt;

&lt;p&gt;The application is an operations console for incident investigation.&lt;/p&gt;

&lt;p&gt;The overview brings together active incidents, historical memory, memory records, and the Hindsight connection.&lt;/p&gt;

&lt;p&gt;Figure 1 — Incident Operations dashboard showing active incidents, historical memory, memory records, and Hindsight activity.&lt;/p&gt;

&lt;p&gt;The important design choice is that memory is part of the incident lifecycle rather than a separate search feature.&lt;/p&gt;

&lt;p&gt;The flow is:&lt;/p&gt;

&lt;p&gt;Current Incident&lt;br&gt;
      ↓&lt;br&gt;
Recall historical context&lt;br&gt;
      ↓&lt;br&gt;
Reflect across relevant memories&lt;br&gt;
      ↓&lt;br&gt;
Recommended investigation&lt;br&gt;
      ↓&lt;br&gt;
Human review&lt;br&gt;
      ↓&lt;br&gt;
Resolution&lt;br&gt;
      ↓&lt;br&gt;
Postmortem / learning&lt;br&gt;
      ↓&lt;br&gt;
Retain new organizational memory&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdytax9qfu8zissxu15to.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdytax9qfu8zissxu15to.png" alt=" " width="799" height="393"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpsnjvp2fq3tt88rhnhs7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpsnjvp2fq3tt88rhnhs7.png" alt=" " width="800" height="1200"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That gives the system a persistent learning loop instead of a one-shot answer.&lt;/p&gt;

&lt;p&gt;The architecture&lt;/p&gt;

&lt;p&gt;I designed the system around a simple separation of responsibilities.&lt;/p&gt;

&lt;p&gt;The data foundation is intentionally broader than a collection of incident titles.&lt;br&gt;
The system uses:&lt;/p&gt;

&lt;p&gt;Rootly incident/log information&lt;/p&gt;

&lt;p&gt;PagerDuty incident-response knowledge&lt;/p&gt;

&lt;p&gt;PagerDuty postmortem knowledge&lt;/p&gt;

&lt;p&gt;100–150 realistic synthetic incident records&lt;/p&gt;

&lt;p&gt;50–100 realistic runbooks&lt;/p&gt;

&lt;p&gt;50–100 realistic postmortems&lt;/p&gt;

&lt;p&gt;The point is to give the memory layer enough operational context to answer questions about previous experience rather than just retrieve isolated documents.&lt;/p&gt;

&lt;p&gt;The core flow: Retain, Recall, Reflect&lt;/p&gt;

&lt;p&gt;The Hindsight integration maps naturally to the incident lifecycle.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Recall before investigation&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When a new incident arrives, the agent first looks for relevant historical context.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;/p&gt;

&lt;p&gt;memories = client.recall(&lt;br&gt;
    bank_id=BANK_ID,&lt;br&gt;
    query=current_incident&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;The query is based on the current incident rather than a generic search phrase.&lt;/p&gt;

&lt;p&gt;That matters because the useful historical context depends on what is happening now.&lt;/p&gt;

&lt;p&gt;For example, a Payment API connection problem should surface memories around connection pools, database exhaustion, similar payment-service incidents, and related operational lessons.&lt;/p&gt;

&lt;p&gt;The goal is not to blindly reuse the previous fix.&lt;/p&gt;

&lt;p&gt;The goal is to give the engineer more context before making the next decision.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Reflect when several memories matter&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Recall gives the agent relevant memories.&lt;/p&gt;

&lt;p&gt;Reflection helps when the useful answer is distributed across multiple memories.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;/p&gt;

&lt;p&gt;reflection = client.reflect(&lt;br&gt;
    bank_id=BANK_ID,&lt;br&gt;
    query="What patterns and failed fixes should I consider?"&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;This gives the agent a way to move from individual historical incidents toward a broader operational pattern.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Retain the actual lesson&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;After the incident is resolved, the most important information should not disappear.&lt;/p&gt;

&lt;p&gt;The application captures:&lt;/p&gt;

&lt;p&gt;Root cause&lt;/p&gt;

&lt;p&gt;What worked&lt;/p&gt;

&lt;p&gt;What failed&lt;/p&gt;

&lt;p&gt;Lesson learned&lt;/p&gt;

&lt;p&gt;Prevention&lt;/p&gt;

&lt;p&gt;That information can then be taught back to organizational memory.&lt;/p&gt;

&lt;p&gt;client.retain(&lt;br&gt;
    bank_id=BANK_ID,&lt;br&gt;
    content=incident_learning,&lt;br&gt;
    context="resolved production incident"&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;The exact value of retention depends on what we put into memory.&lt;/p&gt;

&lt;p&gt;“Incident closed” is not very useful.&lt;/p&gt;

&lt;p&gt;A much better memory is:&lt;/p&gt;

&lt;p&gt;Symptom&lt;br&gt;
→ Investigation&lt;br&gt;
→ Failed action&lt;br&gt;
→ Successful action&lt;br&gt;
→ Root cause&lt;br&gt;
→ Lesson&lt;br&gt;
→ Prevention&lt;/p&gt;

&lt;p&gt;That is the information another engineer can actually use later.&lt;/p&gt;

&lt;p&gt;A real incident flow&lt;/p&gt;

&lt;p&gt;The most useful screen in the application is the active-incident investigation view.&lt;br&gt;
The example is a Payment API incident returning HTTP 502 errors.&lt;/p&gt;

&lt;p&gt;The screen contains current incident information and then brings in historical context.&lt;/p&gt;

&lt;p&gt;The investigation is separated into:&lt;/p&gt;

&lt;p&gt;Historical matches&lt;/p&gt;

&lt;p&gt;Hindsight reflection&lt;/p&gt;

&lt;p&gt;Recommended investigation&lt;/p&gt;

&lt;p&gt;Actions requiring human review&lt;/p&gt;

&lt;p&gt;This is important because historical similarity should not automatically become the root cause.&lt;/p&gt;

&lt;p&gt;A previous incident is evidence.&lt;/p&gt;

&lt;p&gt;It is not proof.&lt;/p&gt;

&lt;p&gt;If the current environment has changed, blindly repeating an old remediation can make an incident worse.&lt;/p&gt;

&lt;p&gt;The agent therefore uses memory to improve the investigation while leaving the final operational decision with the engineer.&lt;/p&gt;

&lt;p&gt;The part I cared about most: remembering what failed&lt;/p&gt;

&lt;p&gt;When an incident is resolved, the application asks the engineer to capture the complete lesson.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhiyvcoe2eqq157qix0l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhiyvcoe2eqq157qix0l.png" alt=" " width="800" height="478"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2j67byjgsi0s4mv2g9lj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2j67byjgsi0s4mv2g9lj.png" alt=" " width="800" height="357"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F27a1mjrlghje3kk0f3wm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F27a1mjrlghje3kk0f3wm.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;br&gt;
The example shown in the application is a database connection-pool problem.&lt;/p&gt;

&lt;p&gt;The root cause is that connection-pool limits were misconfigured in deployment v2.8.0.&lt;/p&gt;

&lt;p&gt;Rolling back to v2.7.9 restored the connection-pool limits.&lt;/p&gt;

&lt;p&gt;But restarting the API gateway was a failed approach because it caused a traffic spike and made the database problem worse.&lt;/p&gt;

&lt;p&gt;That failed action is worth remembering.&lt;/p&gt;

&lt;p&gt;If the next engineer only sees:&lt;/p&gt;

&lt;p&gt;“Rollback to v2.7.9.”&lt;/p&gt;

&lt;p&gt;they know what worked.&lt;/p&gt;

&lt;p&gt;If they also see:&lt;/p&gt;

&lt;p&gt;“Do not restart the gateway first during this type of database-exhaustion event.”&lt;/p&gt;

&lt;p&gt;they know what to avoid.&lt;/p&gt;

&lt;p&gt;That is a much more useful organizational memory.&lt;/p&gt;

&lt;p&gt;Why this is different from a normal runbook&lt;/p&gt;

&lt;p&gt;Runbooks are excellent for known procedures.&lt;/p&gt;

&lt;p&gt;But incident response also contains experience that is difficult to encode as a fixed procedure.&lt;/p&gt;

&lt;p&gt;A runbook can say:&lt;/p&gt;

&lt;p&gt;Check database connection utilization.&lt;br&gt;
Check application pool limits.&lt;br&gt;
Check recent deployments.&lt;/p&gt;

&lt;p&gt;An incident memory can add:&lt;/p&gt;

&lt;p&gt;This failure previously appeared after deployment v2.8.0.&lt;br&gt;
Rolling back restored the connection-pool limits.&lt;br&gt;
Restarting the gateway increased traffic and worsened the condition.&lt;/p&gt;

&lt;p&gt;The runbook tells you what to check.&lt;/p&gt;

&lt;p&gt;The memory tells you what happened when someone checked it before.&lt;/p&gt;

&lt;p&gt;I wanted both.&lt;/p&gt;

&lt;p&gt;Making memory visible&lt;/p&gt;

&lt;p&gt;I also wanted the system to make its memory activity visible instead of hiding it behind an AI response.&lt;/p&gt;

&lt;p&gt;The dashboard exposes:&lt;/p&gt;

&lt;p&gt;RECALL&lt;br&gt;
8 memories retrieved&lt;/p&gt;

&lt;p&gt;REFLECT&lt;br&gt;
Historical pattern synthesized&lt;/p&gt;

&lt;p&gt;RETAIN&lt;br&gt;
New learning stored&lt;/p&gt;

&lt;p&gt;This makes the lifecycle easier to understand:&lt;/p&gt;

&lt;p&gt;Recall what happened before.&lt;/p&gt;

&lt;p&gt;Reflect on the historical evidence.&lt;/p&gt;

&lt;p&gt;Retain what we learned this time.&lt;/p&gt;

&lt;p&gt;For an incident-response system, that transparency matters.&lt;/p&gt;

&lt;p&gt;An engineer should be able to understand why the agent is recommending a particular investigation path.&lt;/p&gt;

&lt;p&gt;Human review is part of the architecture&lt;/p&gt;

&lt;p&gt;I intentionally kept the engineer in the loop.&lt;/p&gt;

&lt;p&gt;The agent can retrieve historical context and produce recommended actions, but it does not silently execute a production remediation.&lt;/p&gt;

&lt;p&gt;The workflow is:&lt;/p&gt;

&lt;p&gt;Incident&lt;br&gt;
   ↓&lt;br&gt;
Historical memory&lt;br&gt;
   ↓&lt;br&gt;
Agent investigation&lt;br&gt;
   ↓&lt;br&gt;
Evidence-backed recommendation&lt;br&gt;
   ↓&lt;br&gt;
Human review&lt;br&gt;
   ↓&lt;br&gt;
Resolution&lt;br&gt;
   ↓&lt;br&gt;
Postmortem&lt;br&gt;
   ↓&lt;br&gt;
New memory&lt;/p&gt;

&lt;p&gt;That boundary matters because historical information can become stale.&lt;/p&gt;

&lt;p&gt;A service can be upgraded.&lt;/p&gt;

&lt;p&gt;A deployment architecture can change.&lt;/p&gt;

&lt;p&gt;A previous workaround can stop being safe.&lt;/p&gt;

&lt;p&gt;Persistent memory should therefore support engineering judgment, not replace it.&lt;/p&gt;

&lt;p&gt;What I learned&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Memory is more than storage&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Putting incident documents somewhere does not automatically create useful memory.&lt;/p&gt;

&lt;p&gt;The important part is retrieving the right experience in the context of a later decision.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The quality of retained information matters&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If I retain only incident titles and status values, future recall will be shallow.&lt;/p&gt;

&lt;p&gt;Root causes, actions, failed approaches, lessons, and prevention steps are much more valuable.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Failed actions deserve first-class treatment&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The successful fix tells me what to do.&lt;/p&gt;

&lt;p&gt;The failed fix tells me what not to repeat.&lt;/p&gt;

&lt;p&gt;For incident response, both are operational knowledge.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Historical similarity is evidence, not truth&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A previous incident can be highly relevant without having the same root cause.&lt;/p&gt;

&lt;p&gt;That is why current evidence and human review remain part of the workflow.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The real loop is Incident → Memory → Incident&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The most interesting behavior is not one successful recall.&lt;/p&gt;

&lt;p&gt;It is the repeated cycle:&lt;/p&gt;

&lt;p&gt;Incident&lt;br&gt;
   ↓&lt;br&gt;
Investigation&lt;br&gt;
   ↓&lt;br&gt;
Resolution&lt;br&gt;
   ↓&lt;br&gt;
Learning&lt;br&gt;
   ↓&lt;br&gt;
Hindsight Memory&lt;br&gt;
   ↓&lt;br&gt;
Next Incident&lt;/p&gt;

&lt;p&gt;The system gets a chance to make previous operational experience useful again.&lt;/p&gt;

&lt;p&gt;Why I chose Hindsight&lt;/p&gt;

&lt;p&gt;I used Hindsight because its memory model fits this workflow naturally.&lt;/p&gt;

&lt;p&gt;The Hindsight GitHub repository and Hindsight documentation describe the retain, recall, and reflect model that the system is built around.&lt;/p&gt;

&lt;p&gt;The broader Vectorize guide to agent memory is also useful for understanding the difference between temporary conversational context and persistent agent memory.&lt;/p&gt;

&lt;p&gt;For this project, the division is straightforward:&lt;/p&gt;

&lt;p&gt;Retain validated incident learning.&lt;/p&gt;

&lt;p&gt;Recall relevant historical experience.&lt;/p&gt;

&lt;p&gt;Reflect when several memories need to be synthesized.&lt;/p&gt;

&lt;p&gt;That gives the incident-response agent a persistent source of organizational context.&lt;/p&gt;

&lt;p&gt;Where I would take it next&lt;/p&gt;

&lt;p&gt;The next step is not simply adding more UI.&lt;/p&gt;

&lt;p&gt;I would connect the system more deeply with operational systems:&lt;/p&gt;

&lt;p&gt;Monitoring and observability platforms&lt;/p&gt;

&lt;p&gt;Log-management systems&lt;/p&gt;

&lt;p&gt;Incident-management platforms&lt;/p&gt;

&lt;p&gt;Deployment and CI/CD systems&lt;/p&gt;

&lt;p&gt;Engineering communication channels&lt;/p&gt;

&lt;p&gt;Automated postmortem generation&lt;/p&gt;

&lt;p&gt;Runbook repositories&lt;/p&gt;

&lt;p&gt;I would also evaluate the memory layer itself.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;How often does recall surface a genuinely relevant incident?&lt;/p&gt;

&lt;p&gt;Which retained fields contribute most to useful recommendations?&lt;/p&gt;

&lt;p&gt;How often does the historical context agree with the eventual root cause?&lt;/p&gt;

&lt;p&gt;Which failed actions are successfully avoided later?&lt;/p&gt;

&lt;p&gt;How does recall change as the memory bank grows?&lt;/p&gt;

&lt;p&gt;Those questions would tell me whether the memory layer is actually improving incident investigation.&lt;/p&gt;

&lt;p&gt;Closing the loop&lt;/p&gt;

&lt;p&gt;The architecture ultimately comes down to one idea:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             ┌──────────────┐
             │   INCIDENT   │
             └──────┬───────┘
                    ↓
             ┌──────────────┐
             │ INVESTIGATE  │
             └──────┬───────┘
                    ↓
             ┌──────────────┐
             │   RESOLVE    │
             └──────┬───────┘
                    ↓
             ┌──────────────┐
             │    LEARN     │
             └──────┬───────┘
                    ↓
             ┌──────────────┐
             │    RETAIN    │
             └──────┬───────┘
                    │
                    └──────────────→ NEXT INCIDENT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;I did not want to build an assistant that gives an answer and forgets it.&lt;/p&gt;

&lt;p&gt;I wanted an incident-response system where the resolution of one outage can become useful context for the next one.&lt;/p&gt;

&lt;p&gt;That is the idea behind Incident-Memory-Copilot:&lt;/p&gt;

&lt;p&gt;turn incident history into organizational memory, and organizational memory into better-informed incident response.&lt;/p&gt;

&lt;p&gt;Project&lt;/p&gt;

&lt;p&gt;Incident-Memory-Copilot on GitHub&lt;/p&gt;

&lt;p&gt;Hindsight on GitHub&lt;/p&gt;

&lt;p&gt;Hindsight Documentation&lt;/p&gt;

&lt;p&gt;Vectorize — What Is Agent Memory?&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>devops</category>
      <category>sre</category>
    </item>
    <item>
      <title>I Designed Incident Response Around Persistent Memory</title>
      <dc:creator>Sai Madhuri</dc:creator>
      <pubDate>Wed, 30 Sep 2026 09:44:49 +0000</pubDate>
      <link>https://dev.to/sai_ed6de9697ae4832f5094b/i-designed-incident-response-around-persistent-memory-bib</link>
      <guid>https://dev.to/sai_ed6de9697ae4832f5094b/i-designed-incident-response-around-persistent-memory-bib</guid>
      <description>&lt;p&gt;I Gave Incident Response a Memory With Hindsight&lt;/p&gt;

&lt;p&gt;The first useful question during an outage is often not “what could be wrong?” but “have we seen this before?”&lt;/p&gt;

&lt;p&gt;I built Incident-Memory-Copilot around that question. The system combines an incident-response workflow with persistent organizational memory so that a new incident can be investigated using what the organization learned from previous incidents.&lt;/p&gt;

&lt;p&gt;What the system does&lt;/p&gt;

&lt;p&gt;The application is an incident operations console. It gives an engineer a place to inspect active incidents, search historical memory, review runbooks and postmortems, investigate a new incident, and explicitly teach the system what was learned after resolution.&lt;/p&gt;

&lt;p&gt;The important architectural decision is that Hindsight is not treated as a secondary search box. It sits inside the incident lifecycle.&lt;/p&gt;

&lt;p&gt;At a high level, the flow is:&lt;/p&gt;

&lt;p&gt;DATA SOURCES&lt;br&gt;
                     │&lt;br&gt;
       ┌─────────────┼──────────────┐&lt;br&gt;
       ▼             ▼              ▼&lt;br&gt;
    Rootly       PagerDuty       PagerDuty&lt;br&gt;
     Logs       Incident Docs   Postmortems&lt;br&gt;
       │             │              │&lt;br&gt;
       └─────────────┼──────────────┘&lt;br&gt;
                     │&lt;br&gt;
                     ▼&lt;br&gt;
          SYNTHETIC ORGANIZATIONAL DATA&lt;br&gt;
                     │&lt;br&gt;
       ┌─────────────┼──────────────┐&lt;br&gt;
       ▼             ▼              ▼&lt;br&gt;
   100–150         50–100         50–100&lt;br&gt;
   Incidents       Runbooks      Postmortems&lt;br&gt;
       │             │              │&lt;br&gt;
       └─────────────┼──────────────┘&lt;br&gt;
                     ▼&lt;br&gt;
              HINDSIGHT CLOUD&lt;br&gt;
                     │&lt;br&gt;
          ┌──────────┼──────────┐&lt;br&gt;
          ▼          ▼          ▼&lt;br&gt;
        RETAIN     RECALL     REFLECT&lt;br&gt;
          │          │          │&lt;br&gt;
          └──────────┼──────────┘&lt;br&gt;
                     ▼&lt;br&gt;
          INCIDENT RESPONSE AGENT&lt;br&gt;
                     │&lt;br&gt;
                     ▼&lt;br&gt;
           EVIDENCE-BACKED ACTIONS&lt;br&gt;
                     │&lt;br&gt;
                     ▼&lt;br&gt;
               HUMAN ENGINEER&lt;br&gt;
                     │&lt;br&gt;
                     ▼&lt;br&gt;
                POSTMORTEM&lt;br&gt;
                     │&lt;br&gt;
                     └──────→ RETAIN&lt;/p&gt;

&lt;p&gt;The data foundation combines operational incident information, incident-response knowledge, postmortems, realistic synthetic incidents, runbooks, and postmortems. Hindsight becomes the layer that turns that accumulated information into reusable memory.&lt;/p&gt;

&lt;p&gt;The application then has a simple loop:&lt;/p&gt;

&lt;p&gt;Recall → Investigate → Human decision → Resolve → Retain&lt;/p&gt;

&lt;p&gt;That loop is more important to me than any individual UI screen.&lt;/p&gt;

&lt;p&gt;Why I wanted memory in the incident workflow&lt;/p&gt;

&lt;p&gt;An LLM can already explain an HTTP 502, list possible database problems, or suggest checking a deployment.&lt;/p&gt;

&lt;p&gt;That is not the difficult part.&lt;/p&gt;

&lt;p&gt;The difficult part is knowing what happened in this environment before.&lt;/p&gt;

&lt;p&gt;Suppose a payment service is returning HTTP 502 responses. A generic assistant might recommend checking the gateway, application health, upstream dependencies, database connectivity, and recent deployments.&lt;/p&gt;

&lt;p&gt;Those are reasonable checks.&lt;/p&gt;

&lt;p&gt;But suppose the organization previously had a related incident where a deployment changed database connection-pool limits. The team discovered that restarting the API gateway made the situation worse because it created a traffic spike, while rolling back the deployment restored the correct connection-pool configuration.&lt;/p&gt;

&lt;p&gt;That historical experience is much more useful than a generic list of possible causes.&lt;/p&gt;

&lt;p&gt;The incident console captures exactly this kind of information.&lt;/p&gt;

&lt;p&gt;A concrete incident flow&lt;/p&gt;

&lt;p&gt;The current incident screen represents an example around a Payment API returning HTTP 502 errors.&lt;/p&gt;

&lt;p&gt;The incident contains operational context such as the service, severity, error, current CPU and memory utilization, recent deployment, and explanatory description.&lt;/p&gt;

&lt;p&gt;The investigation then moves through three memory-oriented stages:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;         CURRENT INCIDENT
                │
                ▼
          HINDSIGHT RECALL
                │
                ▼
      HISTORICAL INCIDENTS
                │
                ▼
      HINDSIGHT REFLECTION
                │
                ▼
      RECOMMENDED INVESTIGATION
                │
                ▼
         HUMAN REVIEW
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The key point is that historical information is presented alongside the current evidence.&lt;/p&gt;

&lt;p&gt;The agent is not simply saying:&lt;/p&gt;

&lt;p&gt;“This happened before, so do the same thing.”&lt;/p&gt;

&lt;p&gt;Instead, the previous incident becomes evidence that an engineer can compare against the current situation.&lt;/p&gt;

&lt;p&gt;That distinction matters in production systems because infrastructure changes over time. The same symptom can have a different cause.&lt;/p&gt;

&lt;p&gt;Retain, Recall, and Reflect&lt;/p&gt;

&lt;p&gt;The integration with Hindsight follows three operations.&lt;/p&gt;

&lt;p&gt;Retain&lt;/p&gt;

&lt;p&gt;When an incident has been resolved and the engineer knows the actual root cause, what worked, what failed, and what should be prevented can be retained as organizational memory.&lt;/p&gt;

&lt;p&gt;Conceptually, the operation looks like:&lt;/p&gt;

&lt;p&gt;client.retain(&lt;br&gt;
    bank_id=BANK_ID,&lt;br&gt;
    content=incident_learning,&lt;br&gt;
    context="resolved production incident"&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;The important part is not the API call itself. It is the quality of what gets retained.&lt;/p&gt;

&lt;p&gt;A useful memory record should capture the operational lesson rather than merely saying that an incident was closed.&lt;/p&gt;

&lt;p&gt;Recall&lt;/p&gt;

&lt;p&gt;When a new incident arrives, the current incident becomes the basis for retrieving relevant historical knowledge:&lt;/p&gt;

&lt;p&gt;memories = client.recall(&lt;br&gt;
    bank_id=BANK_ID,&lt;br&gt;
    query=current_incident&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;This is where the system moves beyond a stateless interaction.&lt;/p&gt;

&lt;p&gt;The engineer is no longer asking the model to reason only from the current error message. The model can receive historical experiences that are relevant to the current investigation.&lt;/p&gt;

&lt;p&gt;Hindsight's recall operation is designed to combine multiple retrieval signals rather than relying only on one simple similarity lookup. The official documentation describes semantic, keyword, graph, and temporal retrieval being combined and reranked. citeturn0search0turn0search2&lt;/p&gt;

&lt;p&gt;Reflect&lt;/p&gt;

&lt;p&gt;Recall gives the agent relevant memories. Reflection is useful when the question requires a broader synthesis across those memories.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;/p&gt;

&lt;p&gt;reflection = client.reflect(&lt;br&gt;
    bank_id=BANK_ID,&lt;br&gt;
    query="What patterns and failed fixes should I consider?"&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;For incident response, that means the system can move from:&lt;/p&gt;

&lt;p&gt;“Here are some similar incidents.”&lt;/p&gt;

&lt;p&gt;toward:&lt;/p&gt;

&lt;p&gt;“Here is the recurring pattern that appears across those incidents.”&lt;/p&gt;

&lt;p&gt;That distinction is why I wanted Hindsight to be central to the design rather than adding a generic vector search layer and calling it memory.&lt;/p&gt;

&lt;p&gt;The second architecture: where the data comes from&lt;/p&gt;

&lt;p&gt;The data foundation is deliberately broader than one incident table.&lt;/p&gt;

&lt;p&gt;DATA FOUNDATION&lt;/p&gt;

&lt;p&gt;Rootly Logs&lt;br&gt;
       +&lt;br&gt;
PagerDuty Incident Response knowledge&lt;br&gt;
       +&lt;br&gt;
PagerDuty Postmortem knowledge&lt;br&gt;
       +&lt;br&gt;
100–150 realistic synthetic incident records&lt;br&gt;
       +&lt;br&gt;
50–100 realistic runbooks&lt;br&gt;
       +&lt;br&gt;
50–100 realistic postmortems&lt;br&gt;
       ↓&lt;br&gt;
HINDSIGHT CLOUD&lt;br&gt;
       ↓&lt;br&gt;
INCIDENT RESPONSE AGENT&lt;/p&gt;

&lt;p&gt;The purpose is to give the memory layer enough operational context to make historical recall meaningful.&lt;/p&gt;

&lt;p&gt;A memory system is only as useful as the information that enters it.&lt;/p&gt;

&lt;p&gt;If every retained record says only “incident resolved successfully,” there is very little for a future investigation to learn.&lt;/p&gt;

&lt;p&gt;The useful record is closer to:&lt;/p&gt;

&lt;p&gt;symptom → investigation → failed action → successful action → root cause → lesson → prevention&lt;/p&gt;

&lt;p&gt;That is the shape of knowledge an incident responder can reuse.&lt;/p&gt;

&lt;p&gt;The moment that changed the design&lt;/p&gt;

&lt;p&gt;The most interesting part of the interface is the resolved-incident screen.&lt;/p&gt;

&lt;p&gt;After the incident is resolved, the engineer is asked to capture:&lt;/p&gt;

&lt;p&gt;Root cause&lt;/p&gt;

&lt;p&gt;What worked&lt;/p&gt;

&lt;p&gt;What failed&lt;/p&gt;

&lt;p&gt;Lesson learned&lt;/p&gt;

&lt;p&gt;Prevention&lt;/p&gt;

&lt;p&gt;In the example shown in the application, the root cause is a connection-pool configuration problem introduced in deployment v2.8.0. Rolling back to v2.7.9 restored the connection-pool limits. Restarting the API gateway failed because it caused a traffic spike that made the database problem worse.&lt;/p&gt;

&lt;p&gt;The lesson is explicit:&lt;/p&gt;

&lt;p&gt;Check connection limits before restarting the gateway during a database-exhaustion event.&lt;/p&gt;

&lt;p&gt;The prevention step is also explicit:&lt;/p&gt;

&lt;p&gt;Add automated tests that verify connection limits before deployment.&lt;/p&gt;

&lt;p&gt;Then there is a Teach Organizational Memory action.&lt;/p&gt;

&lt;p&gt;That button represents the part of the design I care about most.&lt;/p&gt;

&lt;p&gt;The incident is not finished from the system's perspective when the service recovers. The resolution becomes future context.&lt;/p&gt;

&lt;p&gt;Why failed actions belong in memory&lt;/p&gt;

&lt;p&gt;One design decision I would keep even if I rebuilt the system from scratch is retaining failed actions.&lt;/p&gt;

&lt;p&gt;Incident documentation often emphasizes the final fix.&lt;/p&gt;

&lt;p&gt;That is understandable, but it loses valuable information.&lt;/p&gt;

&lt;p&gt;Imagine a future engineer sees:&lt;/p&gt;

&lt;p&gt;“Rollback deployment v2.8.0.”&lt;/p&gt;

&lt;p&gt;That is useful.&lt;/p&gt;

&lt;p&gt;But this is more useful:&lt;/p&gt;

&lt;p&gt;“Rollback deployment v2.8.0 restored connection-pool limits. Restarting the gateway first caused a traffic spike and made the database exhaustion worse.”&lt;/p&gt;

&lt;p&gt;The second record prevents a future engineer from repeating a known mistake.&lt;/p&gt;

&lt;p&gt;For incident response, what did not work can be as valuable as what did.&lt;/p&gt;

&lt;p&gt;This is also where persistent memory is different from a static runbook. A runbook describes a known procedure. Incident memory can preserve the experience around that procedure.&lt;/p&gt;

&lt;p&gt;Making memory visible&lt;/p&gt;

&lt;p&gt;I also wanted the system to make memory operations visible to the engineer.&lt;/p&gt;

&lt;p&gt;The overview screen exposes memory activity as:&lt;/p&gt;

&lt;p&gt;RECALL&lt;br&gt;
8 memories retrieved&lt;/p&gt;

&lt;p&gt;REFLECT&lt;br&gt;
Historical pattern synthesized&lt;/p&gt;

&lt;p&gt;RETAIN&lt;br&gt;
New learning stored&lt;/p&gt;

&lt;p&gt;The dashboard also exposes the relationship between active incidents, historical memory, memory records, and the Hindsight connection.&lt;/p&gt;

&lt;p&gt;That visibility matters because an incident recommendation should not feel like an unexplained answer from an LLM.&lt;/p&gt;

&lt;p&gt;An engineer should be able to ask:&lt;/p&gt;

&lt;p&gt;What historical information influenced this?&lt;/p&gt;

&lt;p&gt;Was the memory retrieved successfully?&lt;/p&gt;

&lt;p&gt;Did the system synthesize a broader pattern?&lt;/p&gt;

&lt;p&gt;What will be remembered after this incident?&lt;/p&gt;

&lt;p&gt;The UI is therefore part of the trust model, not just presentation.&lt;/p&gt;

&lt;p&gt;Keeping the engineer in control&lt;/p&gt;

&lt;p&gt;I deliberately kept a human review step in the incident workflow.&lt;/p&gt;

&lt;p&gt;The system can retrieve historical incidents and produce recommended investigation steps, but it does not get to silently decide that a production change should happen.&lt;/p&gt;

&lt;p&gt;The workflow is:&lt;/p&gt;

&lt;p&gt;Current incident&lt;br&gt;
      ↓&lt;br&gt;
Historical memory&lt;br&gt;
      ↓&lt;br&gt;
AI investigation&lt;br&gt;
      ↓&lt;br&gt;
Evidence-backed recommendation&lt;br&gt;
      ↓&lt;br&gt;
Human review&lt;br&gt;
      ↓&lt;br&gt;
Resolution&lt;br&gt;
      ↓&lt;br&gt;
Postmortem&lt;br&gt;
      ↓&lt;br&gt;
Organizational memory&lt;/p&gt;

&lt;p&gt;That separation is important.&lt;/p&gt;

&lt;p&gt;Historical memory can be wrong, incomplete, or no longer applicable. Infrastructure changes. Services are upgraded. Architecture changes. A recommendation based on a six-month-old incident should therefore be treated as evidence, not an instruction that bypasses engineering judgment.&lt;/p&gt;

&lt;p&gt;What I learned&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Storage is not memory&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Putting incident documents into a database does not automatically give an agent useful memory.&lt;/p&gt;

&lt;p&gt;Memory becomes useful when information can be retrieved in the context of a later decision.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Recall quality depends on what gets retained&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If the retained information is vague, future recall will also be vague.&lt;/p&gt;

&lt;p&gt;The most useful incident memories contain concrete symptoms, services, actions, outcomes, root causes, and lessons.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Failed fixes deserve first-class treatment&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The final successful remediation is only part of the incident story.&lt;/p&gt;

&lt;p&gt;Knowing what made the situation worse can prevent repeated mistakes.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Historical context should inform, not override&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A similar incident is evidence.&lt;/p&gt;

&lt;p&gt;It is not proof that the current incident has the same root cause.&lt;/p&gt;

&lt;p&gt;That is why the current incident, historical memories, and human review remain separate parts of the workflow.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The learning loop is the real product&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The interesting behavior is not a single successful recall.&lt;/p&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;p&gt;Incident → Memory → Investigation → Resolution → New Memory&lt;/p&gt;

&lt;p&gt;Every resolved incident can make the next relevant incident easier to investigate.&lt;/p&gt;

&lt;p&gt;Where Hindsight fits&lt;/p&gt;

&lt;p&gt;I used Hindsight because its memory model maps naturally onto this workflow.&lt;/p&gt;

&lt;p&gt;The official Hindsight GitHub repository describes memory around retain, recall, and reflect, while the Hindsight documentation explains the underlying concepts and APIs. The broader Vectorize explanation of agent memory is also useful for understanding why persistent memory is different from simply keeping more text in an LLM context window.&lt;/p&gt;

&lt;p&gt;For this system, the division is straightforward:&lt;/p&gt;

&lt;p&gt;Retain the validated lesson.&lt;/p&gt;

&lt;p&gt;Recall relevant experience during the next investigation.&lt;/p&gt;

&lt;p&gt;Reflect when several memories need to be synthesized into a broader pattern.&lt;/p&gt;

&lt;p&gt;That gives the incident agent something a normal one-shot prompt does not have: organizational experience that can persist across incidents.&lt;/p&gt;

&lt;p&gt;The part I would measure next&lt;/p&gt;

&lt;p&gt;The next engineering question is not whether the dashboard looks convincing.&lt;/p&gt;

&lt;p&gt;It is whether memory changes incident investigation in measurable ways.&lt;/p&gt;

&lt;p&gt;I would evaluate questions such as:&lt;/p&gt;

&lt;p&gt;How often does recall surface a genuinely relevant prior incident?&lt;/p&gt;

&lt;p&gt;How often does the historical recommendation match the eventual root cause?&lt;/p&gt;

&lt;p&gt;Which failed actions are successfully avoided?&lt;/p&gt;

&lt;p&gt;How does recall quality change as the memory bank grows?&lt;/p&gt;

&lt;p&gt;Which retained fields contribute most to useful future investigations?&lt;/p&gt;

&lt;p&gt;When does reflection provide information that recall alone does not?&lt;/p&gt;

&lt;p&gt;Those measurements would tell me whether the memory layer is actually improving the workflow rather than simply adding another component.&lt;/p&gt;

&lt;p&gt;Closing the loop&lt;/p&gt;

&lt;p&gt;The architecture ultimately comes down to a simple idea:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;          REMEMBER
              ↑
              │
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;INCIDENT → INVESTIGATE → RESOLVE&lt;br&gt;
                  │&lt;br&gt;
                  ↓&lt;br&gt;
              LEARN&lt;br&gt;
                  │&lt;br&gt;
                  └────────→ REMEMBER&lt;/p&gt;

&lt;p&gt;I did not want to build another assistant that gives an answer and forgets it.&lt;/p&gt;

&lt;p&gt;I wanted an incident-response system where a resolved outage becomes useful evidence for the next one.&lt;/p&gt;

&lt;p&gt;That is what Incident-Memory-Copilot is designed around: turning incident history into organizational memory, and organizational memory into better-informed incident response.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>devops</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
