<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: kadigevarshha2006-hub</title>
    <description>The latest articles on DEV Community by kadigevarshha2006-hub (@kadigevarshha2006hub).</description>
    <link>https://dev.to/kadigevarshha2006hub</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150455%2F75f5840a-c43f-4dc4-86a2-96dd91a3b05e.png</url>
      <title>DEV Community: kadigevarshha2006-hub</title>
      <link>https://dev.to/kadigevarshha2006hub</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kadigevarshha2006hub"/>
    <language>en</language>
    <item>
      <title>How I Built an Incident Response Agent That Remembers</title>
      <dc:creator>kadigevarshha2006-hub</dc:creator>
      <pubDate>Tue, 29 Sep 2026 17:10:00 +0000</pubDate>
      <link>https://dev.to/kadigevarshha2006hub/how-i-built-an-incident-response-agent-that-remembers-36ln</link>
      <guid>https://dev.to/kadigevarshha2006hub/how-i-built-an-incident-response-agent-that-remembers-36ln</guid>
      <description>&lt;p&gt;How I Built an Incident Response Agent That Remembers&lt;br&gt;
At 3 AM, an incident response system should not behave as if it has never seen an outage before.&lt;/p&gt;

&lt;p&gt;That was the problem I wanted to explore: how do we make an incident-response agent use the organization's previous operational experience instead of generating another generic list of troubleshooting steps?&lt;/p&gt;

&lt;p&gt;I built an incident-response system around three pieces: a machine-learning model for SLA breach risk, an incident agent for reasoning about the current failure, and Hindsight for retaining and recalling operational memory.&lt;/p&gt;

&lt;p&gt;The interesting part is not simply adding an LLM to incident management. It is giving the system a way to remember what happened before.&lt;/p&gt;

&lt;p&gt;The problem with starting from zero&lt;br&gt;
A new incident usually contains a mixture of structured information and messy operational context: the affected service, symptoms, priority, category, logs, and urgency.&lt;/p&gt;

&lt;p&gt;Incident report input&lt;/p&gt;

&lt;p&gt;A stateless assistant can process that information, but it does not automatically have access to the organization's previous incidents and the decisions engineers made during them.&lt;/p&gt;

&lt;p&gt;That creates a recurring problem.&lt;/p&gt;

&lt;p&gt;An engineer may have already solved a very similar failure months ago, but the next incident starts from scratch unless that knowledge is explicitly retrieved.&lt;/p&gt;

&lt;p&gt;I wanted the system to follow a different path:&lt;/p&gt;

&lt;p&gt;Current Incident&lt;br&gt;
      ↓&lt;br&gt;
Understand the current symptoms&lt;br&gt;
      ↓&lt;br&gt;
Predict SLA breach risk&lt;br&gt;
      ↓&lt;br&gt;
Recall similar historical incidents&lt;br&gt;
      ↓&lt;br&gt;
Use previous resolutions as evidence&lt;br&gt;
      ↓&lt;br&gt;
Produce a response plan&lt;br&gt;
      ↓&lt;br&gt;
Store the confirmed resolution&lt;br&gt;
The last step is particularly important. A successful incident should become useful context for a future incident rather than disappearing when the ticket is closed.&lt;/p&gt;

&lt;p&gt;The architecture&lt;br&gt;
The application uses a FastAPI backend with a browser-based operations dashboard.&lt;/p&gt;

&lt;p&gt;The backend separates the major responsibilities:&lt;/p&gt;

&lt;p&gt;model_service.py handles the machine-learning pipeline.&lt;br&gt;
hindsight_service.py provides the Hindsight integration.&lt;br&gt;
agent.py coordinates incident reasoning.&lt;br&gt;
app.py exposes the API endpoints and serves the dashboard.&lt;br&gt;
The repository also contains the historical incident data used to seed the Hindsight memory bank and automated tests for the incident flow.&lt;/p&gt;

&lt;p&gt;At a high level, the request moves through two different kinds of information.&lt;/p&gt;

&lt;p&gt;The first is quantitative:&lt;/p&gt;

&lt;p&gt;Incident features&lt;br&gt;
      ↓&lt;br&gt;
RandomForest&lt;br&gt;
      ↓&lt;br&gt;
SLA breach probability&lt;br&gt;
The second is experiential:&lt;/p&gt;

&lt;p&gt;Current incident&lt;br&gt;
      ↓&lt;br&gt;
Hindsight recall&lt;br&gt;
      ↓&lt;br&gt;
Relevant historical incidents&lt;br&gt;
      ↓&lt;br&gt;
Previous causes + resolutions&lt;br&gt;
The agent can then use both forms of information when constructing its response.&lt;/p&gt;

&lt;p&gt;Why Hindsight is the interesting part&lt;br&gt;
The Hindsight integration gives the application an explicit memory layer.&lt;/p&gt;

&lt;p&gt;Historical incident information can be retained and later recalled semantically. The project uses Hindsight's recall functionality to retrieve relevant incident memories rather than relying only on exact keyword matches.&lt;/p&gt;

&lt;p&gt;The basic idea is straightforward:&lt;/p&gt;

&lt;p&gt;memories = await hindsight.arecall(&lt;br&gt;
    bank_id=BANK_ID,&lt;br&gt;
    query=incident_context&lt;br&gt;
)&lt;br&gt;
Institutional memory retrieved by Hindsight&lt;/p&gt;

&lt;p&gt;The returned memories become context for the incident analysis.&lt;/p&gt;

&lt;p&gt;This changes the behavior of the system.&lt;/p&gt;

&lt;p&gt;Instead of asking only:&lt;/p&gt;

&lt;p&gt;"What could cause these symptoms?"&lt;/p&gt;

&lt;p&gt;the agent can effectively ask:&lt;/p&gt;

&lt;p&gt;"Have we seen something similar before, and what happened when we did?"&lt;/p&gt;

&lt;p&gt;That distinction matters in operational systems because previous incidents often contain information that isn't obvious from the current symptoms alone.&lt;/p&gt;

&lt;p&gt;For example, the repository includes historical scenarios involving database connection pools, authentication certificate problems, and Kubernetes JVM memory failures. Those memories can be recalled when a new incident resembles them.&lt;/p&gt;

&lt;p&gt;Memory is useful only if the system can learn&lt;br&gt;
Retrieval alone isn't enough.&lt;/p&gt;

&lt;p&gt;If the system can read old incidents but cannot remember newly confirmed solutions, the memory remains static.&lt;/p&gt;

&lt;p&gt;That is why the project includes a closed-loop flow.&lt;/p&gt;

&lt;p&gt;After an incident has been resolved, the confirmed root cause and remediation can be retained:&lt;/p&gt;

&lt;p&gt;await hindsight.aretain(&lt;br&gt;
    bank_id=BANK_ID,&lt;br&gt;
    content=resolution&lt;br&gt;
)&lt;br&gt;
A later incident can then retrieve that experience.&lt;/p&gt;

&lt;p&gt;The intended cycle is:&lt;/p&gt;

&lt;p&gt;Incident A&lt;br&gt;
   ↓&lt;br&gt;
Diagnose&lt;br&gt;
   ↓&lt;br&gt;
Resolve&lt;br&gt;
   ↓&lt;br&gt;
Retain resolution&lt;br&gt;
   ↓&lt;br&gt;
Incident B&lt;br&gt;
   ↓&lt;br&gt;
Recall Incident A&lt;br&gt;
   ↓&lt;br&gt;
Use its resolution as historical context&lt;br&gt;
This is a much more useful model of organizational memory than simply storing a conversation transcript.&lt;/p&gt;

&lt;p&gt;Why I kept the ML model separate from memory&lt;br&gt;
Another design decision was to avoid treating historical memory as a replacement for quantitative prediction.&lt;/p&gt;

&lt;p&gt;The project uses a Scikit-learn RandomForestClassifier pipeline to estimate SLA breach risk from structured incident attributes.&lt;/p&gt;

&lt;p&gt;The repository documents the model as being trained on ServiceNow incident data, with categorical dimensions including categories, subcategories, and assignment groups.&lt;/p&gt;

&lt;p&gt;That produces a different kind of signal from Hindsight.&lt;/p&gt;

&lt;p&gt;The ML model answers:&lt;/p&gt;

&lt;p&gt;"Based on the structured characteristics of this incident, how much SLA risk is associated with it?"&lt;/p&gt;

&lt;p&gt;Hindsight answers:&lt;/p&gt;

&lt;p&gt;"What relevant operational experiences do we already have?"&lt;/p&gt;

&lt;p&gt;Those are complementary questions.&lt;/p&gt;

&lt;p&gt;A probability score by itself doesn't explain how engineers should respond. A historical memory by itself doesn't necessarily quantify the urgency of the current ticket.&lt;/p&gt;

&lt;p&gt;Combining them gives the incident agent more context.&lt;/p&gt;

&lt;p&gt;SLA breach risk prediction&lt;/p&gt;

&lt;p&gt;The incident flow&lt;br&gt;
A user begins with the incident intake interface.&lt;/p&gt;

&lt;p&gt;The current service, symptoms, category, priority, and other incident information are submitted to the backend through the analysis endpoint.&lt;/p&gt;

&lt;p&gt;The backend then performs the relevant processing.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;/p&gt;

&lt;p&gt;POST /api/analyze&lt;br&gt;
       │&lt;br&gt;
       ├── Extract incident features&lt;br&gt;
       │&lt;br&gt;
       ├── Calculate SLA risk&lt;br&gt;
       │&lt;br&gt;
       ├── Recall historical memories&lt;br&gt;
       │&lt;br&gt;
       └── Generate incident analysis&lt;br&gt;
Incident response dashboard&lt;/p&gt;

&lt;p&gt;The dashboard can then present the resulting risk information, recalled historical context, and response plan.&lt;/p&gt;

&lt;p&gt;Safe verification workflow&lt;/p&gt;

&lt;p&gt;The repository also includes a resolution endpoint for the closed learning loop:&lt;/p&gt;

&lt;p&gt;POST /api/resolve&lt;br&gt;
       ↓&lt;br&gt;
Confirmed resolution&lt;br&gt;
       ↓&lt;br&gt;
Hindsight retain()&lt;br&gt;
Incident analysis and performance&lt;/p&gt;

&lt;p&gt;This separation makes the workflow easier to reason about: analysis and learning are related, but they are not the same operation.&lt;/p&gt;

&lt;p&gt;What surprised me while building it&lt;br&gt;
The hardest conceptual part wasn't getting an API endpoint to return a response.&lt;/p&gt;

&lt;p&gt;It was deciding what the agent should actually remember.&lt;/p&gt;

&lt;p&gt;Incident memory becomes useful when it contains operationally meaningful information: what failed, what evidence pointed toward the cause, what remediation was applied, and whether that remediation was confirmed.&lt;/p&gt;

&lt;p&gt;Simply dumping every interaction into memory would create a large collection of text without necessarily creating useful operational knowledge.&lt;/p&gt;

&lt;p&gt;That led to a principle I kept coming back to:&lt;/p&gt;

&lt;p&gt;Memory should capture experience, not just conversation.&lt;/p&gt;

&lt;p&gt;The distinction is important for future incident retrieval.&lt;/p&gt;

&lt;p&gt;If an engineer previously solved a JVM memory problem by changing a container-aware configuration, that remediation is potentially useful to another incident with similar symptoms.&lt;/p&gt;

&lt;p&gt;The useful memory isn't the fact that someone had a conversation about the incident. It is the operational knowledge contained in the resolution.&lt;/p&gt;

&lt;p&gt;What I learned&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Memory and reasoning solve different problems
An agent can reason about a current incident, but reasoning does not automatically give it organizational experience.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Memory provides the missing historical context.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Retrieval needs meaningful information
Good recall depends on what is retained.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Incident IDs alone aren't useful memories. Root causes, symptoms, services, remediation steps, and verification results provide much more useful context.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Quantitative and qualitative signals complement each other
The ML model provides a quantitative SLA-risk signal.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Hindsight provides historical operational context.&lt;/p&gt;

&lt;p&gt;Neither needs to replace the other.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Closed-loop learning changes the system over time
A resolved incident should not simply disappear.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Once a confirmed resolution is retained, it can become evidence for future incidents.&lt;/p&gt;

&lt;p&gt;That is the foundation of an incident-response system that can accumulate operational knowledge.&lt;/p&gt;

&lt;p&gt;Where this goes next&lt;br&gt;
The current project is an engineering implementation rather than the end of the problem.&lt;/p&gt;

&lt;p&gt;The next level is making the memory lifecycle increasingly reliable: deciding what should be retained, validating resolutions before they become reusable knowledge, improving retrieval quality, and measuring whether recalled incidents actually improve response outcomes.&lt;/p&gt;

&lt;p&gt;That is the part I find most interesting.&lt;/p&gt;

&lt;p&gt;The goal isn't to build an assistant that produces more text when an incident occurs.&lt;/p&gt;

&lt;p&gt;The goal is to build one that can say:&lt;/p&gt;

&lt;p&gt;"I've seen something like this before. Here's what happened, here's what we learned, and here's the evidence that makes that experience relevant now."&lt;/p&gt;

&lt;p&gt;That is what makes memory useful in an incident-response agent.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjmlrel6gzb0vkwlr0b37.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjmlrel6gzb0vkwlr0b37.png" alt=" " width="799" height="714"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2m78u211o7fq94pezb47.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2m78u211o7fq94pezb47.png" alt=" " width="800" height="538"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F76rtsyaso4cq4ay0euar.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F76rtsyaso4cq4ay0euar.png" alt=" " width="800" height="694"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9jm7g16m5944novund3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9jm7g16m5944novund3.png" alt=" " width="548" height="1600"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffb7m3m2858piuzyuoqad.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffb7m3m2858piuzyuoqad.png" alt=" " width="800" height="494"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg5wzhm66h9ezyy3qyifq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg5wzhm66h9ezyy3qyifq.png" alt=" " width="800" height="1082"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>llm</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
