<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Huda</title>
    <description>The latest articles on DEV Community by Huda (@hudakhan).</description>
    <link>https://dev.to/hudakhan</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149708%2Faa6782f9-a653-4556-ab1b-124d485113b4.png</url>
      <title>DEV Community: Huda</title>
      <link>https://dev.to/hudakhan</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/hudakhan"/>
    <language>en</language>
    <item>
      <title>Incident Response Agent — Intelligent Incident Analysis and Response Assistance</title>
      <dc:creator>Huda</dc:creator>
      <pubDate>Tue, 29 Sep 2026 12:48:41 +0000</pubDate>
      <link>https://dev.to/hudakhan/incident-response-agent-intelligent-incident-analysis-and-response-assistance-5a94</link>
      <guid>https://dev.to/hudakhan/incident-response-agent-intelligent-incident-analysis-and-response-assistance-5a94</guid>
      <description>&lt;p&gt;When an incident starts at 3 AM, the first question is rarely “Can we search our database?”&lt;/p&gt;

&lt;p&gt;It is usually something closer to:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“Has this happened before, and can I trust what the previous responder did?”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That distinction shaped how I think about our Incident Response Agent.&lt;/p&gt;

&lt;p&gt;The project is designed around a simple problem: incident history is valuable, but historical information by itself is not enough. A responder needs useful precedent, some explanation for why that precedent is relevant, and an indication of whether the associated runbook has actually been successful before.&lt;/p&gt;

&lt;p&gt;Our implementation approaches that problem without an external LLM or network dependency. The application runs locally using FastAPI, SQLAlchemy, SQLite, scikit-learn's TF-IDF implementation, and a vanilla JavaScript frontend.&lt;/p&gt;

&lt;p&gt;The result is less like asking an AI to invent a solution and more like giving an on-call engineer a structured memory of previous incidents.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with ordinary incident history
&lt;/h2&gt;

&lt;p&gt;Imagine that &lt;code&gt;payments-api&lt;/code&gt; starts returning upstream 502 errors.&lt;/p&gt;

&lt;p&gt;An engineer could search through old incident tickets and perhaps find something similar. But “similar” can mean several different things.&lt;/p&gt;

&lt;p&gt;Maybe the old incident affected the same service.&lt;/p&gt;

&lt;p&gt;Maybe it had the same error signature.&lt;/p&gt;

&lt;p&gt;Maybe it happened at the same severity.&lt;/p&gt;

&lt;p&gt;Maybe the description contains similar words.&lt;/p&gt;

&lt;p&gt;Or maybe the incident is old enough that the corresponding runbook should no longer be trusted without validation.&lt;/p&gt;

&lt;p&gt;A simple keyword search doesn't communicate all of that.&lt;/p&gt;

&lt;p&gt;This is why our system doesn't return only a single opaque similarity percentage. It creates an incident fingerprint with five dimensions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;text similarity&lt;/li&gt;
&lt;li&gt;service match&lt;/li&gt;
&lt;li&gt;error-signature match&lt;/li&gt;
&lt;li&gt;severity match&lt;/li&gt;
&lt;li&gt;recency&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That makes the result more useful to a human responder because the system exposes some of the reasoning behind its match.&lt;/p&gt;

&lt;h2&gt;
  
  
  Following the responder's workflow
&lt;/h2&gt;

&lt;p&gt;The application's workflow mirrors what an engineer actually does during an incident.&lt;/p&gt;

&lt;p&gt;First, a new incident is created with information such as its description, service, severity, and error signature.&lt;/p&gt;

&lt;p&gt;The search endpoint then compares the incident against resolved incidents:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;matches&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;similarity&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find_similar_incidents&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;query_text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;query_service&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;query_severity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;severity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;query_error_sig&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;error_signature&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;top_k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;top_k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important detail here is that the system isn't searching every record indiscriminately. The search router retrieves incidents whose status is &lt;code&gt;RESOLVED&lt;/code&gt;, then sends those candidates into the similarity engine.&lt;/p&gt;

&lt;p&gt;For the responder, this creates a much more focused question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which resolved incidents look like the one I'm dealing with now?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A concrete example
&lt;/h2&gt;

&lt;p&gt;The repository's tests use a payments incident involving:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;502-upstream-timeout&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The test query describes &lt;code&gt;payments-api&lt;/code&gt; returning upstream 502 errors while the connection pool appears saturated.&lt;/p&gt;

&lt;p&gt;The historical candidate contains the same service and error signature.&lt;/p&gt;

&lt;p&gt;The similarity test explicitly verifies that this incident ranks first and that both the service and error-signature matches are detected.&lt;/p&gt;

&lt;p&gt;That matters because it demonstrates an actual behavior of the implementation rather than a hypothetical example.&lt;/p&gt;

&lt;p&gt;The response also carries the fingerprint.&lt;/p&gt;

&lt;p&gt;Instead of seeing something like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Match: 87%&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;the responder can see evidence such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;same service&lt;/li&gt;
&lt;li&gt;identical error signature&lt;/li&gt;
&lt;li&gt;same severity class&lt;/li&gt;
&lt;li&gt;historical incident title&lt;/li&gt;
&lt;li&gt;textual similarity&lt;/li&gt;
&lt;li&gt;recency&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The search router even constructs a human-readable explanation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;why&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Matched on &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;why_bits&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That small design choice is important.&lt;/p&gt;

&lt;p&gt;During an incident, a recommendation that cannot explain itself creates another question for the engineer: &lt;em&gt;Why should I trust this?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;An explainable match at least gives the responder something concrete to inspect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory is more than storing incidents
&lt;/h2&gt;

&lt;p&gt;One of the interesting parts of this project is that “memory” isn't implemented as a collection of old incident descriptions alone.&lt;/p&gt;

&lt;p&gt;The database model connects an incident with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;root causes&lt;/li&gt;
&lt;li&gt;resolution steps&lt;/li&gt;
&lt;li&gt;runbook usage&lt;/li&gt;
&lt;li&gt;postmortems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A resolution step records the minute offset, actor, type of action, description, and optionally the runbook that was used.&lt;/p&gt;

&lt;p&gt;That means the system can preserve a sequence rather than just an outcome.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;T+0:&lt;/strong&gt; detection&lt;br&gt;
&lt;strong&gt;T+3:&lt;/strong&gt; diagnosis&lt;br&gt;
&lt;strong&gt;T+8:&lt;/strong&gt; mitigation&lt;br&gt;
&lt;strong&gt;T+15:&lt;/strong&gt; resolution&lt;/p&gt;

&lt;p&gt;The frontend describes this as a black-box-style incident timeline.&lt;/p&gt;

&lt;p&gt;From an engineer's perspective, that is useful because incident response is inherently temporal. Knowing that a particular fix worked is helpful. Knowing what responders tried before that fix can be even more useful when the next incident is slightly different.&lt;/p&gt;

&lt;h2&gt;
  
  
  The system also remembers whether a runbook worked
&lt;/h2&gt;

&lt;p&gt;This is probably the most interesting part of the project's approach.&lt;/p&gt;

&lt;p&gt;A runbook isn't treated as permanently trustworthy just because someone used it successfully once.&lt;/p&gt;

&lt;p&gt;Each runbook starts with an Elo-style rating of 1200.&lt;/p&gt;

&lt;p&gt;When a runbook is used, the system records the incident, outcome, previous rating, new rating, and timestamp.&lt;/p&gt;

&lt;p&gt;The rating update considers the severity of the incident.&lt;/p&gt;

&lt;p&gt;The implementation models SEV1 as a stronger “opponent” than SEV4. Consequently, successfully resolving a high-severity incident provides stronger evidence that the runbook can perform under pressure.&lt;/p&gt;

&lt;p&gt;A failure also has context.&lt;/p&gt;

&lt;p&gt;The system therefore avoids reducing runbook quality to a simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;worked 8 / used 10&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead, it tries to represent the confidence gained from different kinds of incident outcomes.&lt;/p&gt;

&lt;p&gt;There is another useful detail: ratings decay toward the 1200 baseline when a runbook hasn't been used recently.&lt;/p&gt;

&lt;p&gt;That reflects a practical reality of infrastructure: systems change.&lt;/p&gt;

&lt;p&gt;A runbook that successfully handled an incident eight months ago might still be useful, but its historical rating shouldn't necessarily be treated as equally strong forever.&lt;/p&gt;

&lt;h2&gt;
  
  
  From responder memory to organizational memory
&lt;/h2&gt;

&lt;p&gt;This creates an interesting loop.&lt;/p&gt;

&lt;p&gt;A responder receives suggestions from historical incidents.&lt;/p&gt;

&lt;p&gt;They use a runbook.&lt;/p&gt;

&lt;p&gt;They record what happened.&lt;/p&gt;

&lt;p&gt;The incident is resolved.&lt;/p&gt;

&lt;p&gt;The system calculates MTTR.&lt;/p&gt;

&lt;p&gt;A postmortem can then be generated from the incident's structured timeline and root causes.&lt;/p&gt;

&lt;p&gt;Finally, feedback about the runbook changes its rating.&lt;/p&gt;

&lt;p&gt;The next responder therefore doesn't start with exactly the same information.&lt;/p&gt;

&lt;p&gt;The system has accumulated another piece of evidence.&lt;/p&gt;

&lt;p&gt;That is the real value of persistent incident history: not simply remembering that an incident existed, but gradually building a more useful record of what happened and what actually worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The postmortem is intentionally not an AI-generated answer
&lt;/h2&gt;

&lt;p&gt;There is an important engineering decision here that I found particularly useful.&lt;/p&gt;

&lt;p&gt;The project does &lt;strong&gt;not&lt;/strong&gt; call an external LLM to generate the postmortem.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;postmortem_gen.py&lt;/code&gt; constructs the draft from the incident's own data.&lt;/p&gt;

&lt;p&gt;It uses the recorded severity, duration, error signature, root causes, and response timeline to generate sections such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;summary&lt;/li&gt;
&lt;li&gt;contributing factors&lt;/li&gt;
&lt;li&gt;lessons learned&lt;/li&gt;
&lt;li&gt;action items&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The generated postmortem is explicitly treated as a draft.&lt;/p&gt;

&lt;p&gt;That is a sensible boundary for incident response. A postmortem contains operational context and requires human judgment, especially around customer impact, blast radius, and blameless framing.&lt;/p&gt;

&lt;p&gt;Automating the first draft can remove repetitive writing without pretending that the generated text is the final source of truth.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens without this context?
&lt;/h2&gt;

&lt;p&gt;The difference can be thought of as two workflows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Without historical context:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Incident → investigate → search manually → decide what to try → resolve → document later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With the project's incident memory:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Incident → retrieve resolved precedents → inspect the similarity fingerprint → review associated runbooks → respond → record timeline → resolve → update runbook evidence → generate postmortem.&lt;/p&gt;

&lt;p&gt;The second workflow doesn't remove the engineer from the loop.&lt;/p&gt;

&lt;p&gt;That is important.&lt;/p&gt;

&lt;p&gt;The system isn't deciding that a particular runbook must be executed. It is organizing evidence so that the engineer can make a faster, more informed decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  There are real limitations
&lt;/h2&gt;

&lt;p&gt;The implementation is deliberately conservative in several areas.&lt;/p&gt;

&lt;p&gt;The biggest limitation is that similarity is lexical rather than semantic. TF-IDF works well when important words overlap, but two incidents describing the same root cause using completely different terminology may not match strongly.&lt;/p&gt;

&lt;p&gt;The project itself identifies embeddings as a possible future extension behind the existing similarity interface.&lt;/p&gt;

&lt;p&gt;The database is also SQLite. That is practical for the current application and demo, but the README explicitly notes that a larger organizational deployment would require a more scalable database such as PostgreSQL.&lt;/p&gt;

&lt;p&gt;And the postmortem generator is template-based. It is deterministic and works offline, but it isn't a substitute for a human review.&lt;/p&gt;

&lt;p&gt;These limitations are actually useful because they define where the system's confidence should stop.&lt;/p&gt;

&lt;h2&gt;
  
  
  One important distinction about Hindsight
&lt;/h2&gt;

&lt;p&gt;The project description that accompanied this work refers to Hindsight, but the supplied repository does not contain evidence of a Hindsight SDK, dependency, configuration, or &lt;code&gt;retain&lt;/code&gt;/&lt;code&gt;recall&lt;/code&gt; calls.&lt;/p&gt;

&lt;p&gt;The persistent-memory behavior demonstrated by this ZIP is implemented using the application's own SQLite/SQLAlchemy data model and deterministic similarity engine.&lt;/p&gt;

&lt;p&gt;So I would not describe the current repository as a Hindsight integration without additional project files proving that connection.&lt;/p&gt;

&lt;p&gt;That distinction matters in technical writing. A good engineering article should describe what the code actually does, not what we wish the architecture looked like.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned from the responder's perspective
&lt;/h2&gt;

&lt;p&gt;The biggest lesson for me is that incident memory is only useful when it helps answer three questions:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Have we seen something like this before?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does the system think it is similar?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I trust the historical fix?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Our implementation tries to answer those questions with concrete data rather than an unexplained AI score.&lt;/p&gt;

&lt;p&gt;It stores the incident.&lt;/p&gt;

&lt;p&gt;It preserves the response timeline.&lt;/p&gt;

&lt;p&gt;It connects incidents to runbooks.&lt;/p&gt;

&lt;p&gt;It records outcomes.&lt;/p&gt;

&lt;p&gt;It adjusts runbook trust.&lt;/p&gt;

&lt;p&gt;It accounts for freshness.&lt;/p&gt;

&lt;p&gt;And it turns the accumulated information into something the next responder can actually inspect.&lt;/p&gt;

&lt;p&gt;That changes the meaning of “we've seen this before.”&lt;/p&gt;

&lt;p&gt;Instead of being a vague memory from an old incident ticket, it becomes structured evidence that another engineer can investigate before taking action.&lt;/p&gt;

&lt;p&gt;For incident response, that distinction is small in code but significant in practice.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>backend</category>
      <category>fastapi</category>
    </item>
  </channel>
</rss>
