<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Syed Rafi</title>
    <description>The latest articles on DEV Community by Syed Rafi (@rafi_syed_64376918f70802d).</description>
    <link>https://dev.to/rafi_syed_64376918f70802d</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150593%2F43cc4a80-af50-4499-abd5-1b3e0eb5cb24.png</url>
      <title>DEV Community: Syed Rafi</title>
      <link>https://dev.to/rafi_syed_64376918f70802d</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rafi_syed_64376918f70802d"/>
    <language>en</language>
    <item>
      <title>Incident Response Agent — Intelligent Incident Analysis and Response Assistance</title>
      <dc:creator>Syed Rafi</dc:creator>
      <pubDate>Tue, 29 Sep 2026 17:41:19 +0000</pubDate>
      <link>https://dev.to/rafi_syed_64376918f70802d/incident-response-agent-intelligent-incident-analysis-and-response-assistance-5n9</link>
      <guid>https://dev.to/rafi_syed_64376918f70802d/incident-response-agent-intelligent-incident-analysis-and-response-assistance-5n9</guid>
      <description>&lt;p&gt;At 3 AM, Similarity Isn't Enough: Building an Incident-Response Workflow That Explains Its Recommendations&lt;br&gt;
When an incident starts at 3 AM, the difficult part is often not recognizing that something is broken. The difficult part is deciding what to do next.&lt;br&gt;
A familiar error can have several possible causes. A runbook may have worked before, but under different conditions. An incident from six months ago may look relevant while being too stale to trust blindly.&lt;br&gt;
That was the problem we wanted our Incident Response Agent to address.&lt;br&gt;
Instead of treating incident history as a searchable archive, the system turns historical incidents into structured evidence. A new incident can be compared against resolved incidents, the similarity can be broken down into understandable dimensions, previously used runbooks can be surfaced, the current response can be recorded as a timeline, and the eventual result can influence how much trust the system places in that runbook later.&lt;br&gt;
The interesting part is that this does not require a large language model or an external API. The implementation runs locally using deterministic components.&lt;br&gt;
The problem: incident history is only useful if responders can act on it&lt;br&gt;
An incident database can contain years of valuable operational knowledge and still be difficult to use during an outage.&lt;br&gt;
Suppose &lt;code&gt;payments-api&lt;/code&gt; starts returning upstream 502 errors. A responder may want answers to several questions immediately:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Have we seen this before?&lt;/li&gt;
&lt;li&gt;Was the previous incident on the same service?&lt;/li&gt;
&lt;li&gt;Was the error signature identical?&lt;/li&gt;
&lt;li&gt;Was it the same severity?&lt;/li&gt;
&lt;li&gt;What runbook did the previous responder use?&lt;/li&gt;
&lt;li&gt;Did that runbook actually work?&lt;/li&gt;
&lt;li&gt;How recently was that runbook validated?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A simple full-text search answers only part of this.&lt;br&gt;
Our system instead models the matching process explicitly.&lt;br&gt;
The similarity engine compares a new incident against resolved incidents and produces an &lt;strong&gt;incident fingerprint&lt;/strong&gt; containing five dimensions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Text similarity&lt;/li&gt;
&lt;li&gt;Service match&lt;/li&gt;
&lt;li&gt;Error-signature match&lt;/li&gt;
&lt;li&gt;Severity match&lt;/li&gt;
&lt;li&gt;Recency&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is important because the final score is not presented as the entire explanation.&lt;br&gt;
A responder can see &lt;em&gt;why&lt;/em&gt; an incident was considered relevant.&lt;br&gt;
From a new incident to historical precedents&lt;br&gt;
The search flow begins in the &lt;code&gt;/api/search/suggest&lt;/code&gt; endpoint.&lt;br&gt;
The endpoint retrieves resolved incidents and passes the new incident's information into &lt;code&gt;find_similar_incidents()&lt;/code&gt;.&lt;br&gt;
The matching engine uses TF-IDF and cosine similarity over incident information including the title, description, service, error signature, and root-cause summaries.&lt;br&gt;
The core implementation looks like this:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
WEIGHTS = {&lt;br&gt;
    "text_similarity": 0.40,&lt;br&gt;
    "service_match": 0.20,&lt;br&gt;
    "error_signature_match": 0.20,&lt;br&gt;
    "severity_match": 0.10,&lt;br&gt;
    "recency": 0.10,&lt;br&gt;
}&lt;br&gt;
This gives text similarity the largest contribution, while still allowing operational attributes to influence the result.&lt;br&gt;
For example, an incident on the same service with the same error signature receives explicit matches for those dimensions instead of depending entirely on vocabulary overlap.&lt;br&gt;
The result is an overall score plus the underlying fingerprint.&lt;br&gt;
That distinction matters.&lt;br&gt;
A score such as &lt;code&gt;0.82&lt;/code&gt; tells me that the system thinks two incidents are similar. The fingerprint tells me whether that similarity comes from the same service, the same error signature, matching severity, recent history, or simply similar language.&lt;br&gt;
The system deliberately avoids a black-box similarity explanation&lt;br&gt;
One of the implementation decisions I found particularly useful is the decision to keep the similarity dimensions separate.&lt;br&gt;
The fingerprint generated by the matcher contains:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
fingerprint = {&lt;br&gt;
    "text_similarity": round(float(text_sim), 3),&lt;br&gt;
    "service_match": service_match,&lt;br&gt;
    "error_signature_match": error_match,&lt;br&gt;
    "severity_match": severity_match,&lt;br&gt;
    "recency": round(recency, 3),&lt;br&gt;
}&lt;br&gt;
The API then uses those values to construct a plain-English explanation.&lt;br&gt;
For example, a recommendation can explain that the historical incident was on the same service, had an identical error signature, had the same severity class, and was previously associated with a particular runbook.&lt;br&gt;
That gives the responder something much more useful than an unexplained ranking.&lt;br&gt;
Similar incidents are only useful if their fixes are useful&lt;br&gt;
Finding an old incident is only half the workflow.&lt;br&gt;
The next question is:&lt;br&gt;
What did the responder actually do?&lt;br&gt;
The project connects historical incidents to resolution steps and runbooks. When a matching incident contains a runbook reference, the search endpoint surfaces that runbook as a recommendation.&lt;br&gt;
The system also sorts recommended runbooks by their current trust rating.&lt;br&gt;
That introduces another interesting design problem.&lt;br&gt;
A runbook used successfully twice is not necessarily equivalent to one that has repeatedly succeeded against serious incidents.&lt;br&gt;
Treating runbook trust like an Elo rating&lt;br&gt;
The project uses an Elo-style rating system rather than a simple success percentage.&lt;br&gt;
Every runbook starts at:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
STARTING_RATING = 1200.0&lt;/p&gt;

&lt;p&gt;Incident severity is treated as the strength of the opponent:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
SEVERITY_DIFFICULTY = {&lt;br&gt;
    "SEV1": 1600,&lt;br&gt;
    "SEV2": 1400,&lt;br&gt;
    "SEV3": 1200,&lt;br&gt;
    "SEV4": 1000,&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;A successful runbook against a SEV1 therefore provides stronger evidence than a successful runbook against a SEV4.&lt;br&gt;
The feedback endpoint records whether the runbook:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;worked,&lt;/li&gt;
&lt;li&gt;partially worked, or&lt;/li&gt;
&lt;li&gt;failed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rating is then updated and the usage is persisted.&lt;br&gt;
This creates a feedback loop:&lt;br&gt;
incident → historical match → runbook recommendation → outcome → updated trust → future recommendation&lt;br&gt;
That is more interesting than simply storing whether someone clicked "worked."&lt;/p&gt;

&lt;p&gt;The incident itself becomes a timeline&lt;br&gt;
The workflow does not stop after a recommendation.&lt;br&gt;
Responders can record resolution steps through:&lt;/p&gt;

&lt;p&gt;text&lt;br&gt;
POST /api/incidents/{incident_id}/steps&lt;br&gt;
Each step can include information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;minute offset,&lt;/li&gt;
&lt;li&gt;actor,&lt;/li&gt;
&lt;li&gt;step type,&lt;/li&gt;
&lt;li&gt;action text,&lt;/li&gt;
&lt;li&gt;associated runbook.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This creates what the project describes as a "black-box replay" of the incident.&lt;br&gt;
That matters for two reasons.&lt;br&gt;
First, it gives the responder a structured record of what happened.&lt;br&gt;
Second, it creates information that can be used after the incident is resolved.&lt;/p&gt;

&lt;p&gt;The system calculates MTTR when the incident is resolved:&lt;br&gt;
python&lt;br&gt;
delta = incident.resolved_at - incident.started_at&lt;br&gt;
incident.mttr_minutes = max(&lt;br&gt;
    1,&lt;br&gt;
    int(delta.total_seconds() // 60)&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;The response therefore becomes more than a temporary conversation. It becomes structured incident history.&lt;br&gt;
Closing the loop with a postmortem&lt;br&gt;
Once an incident is resolved, the system can generate a postmortem draft.&lt;br&gt;
Importantly, this is **not an LLM-generated postmortem.&lt;br&gt;
The repository's implementation builds the draft deterministically from the incident's stored information, root causes, and timeline.&lt;br&gt;
For example, the postmortem generator can identify the first mitigation step and turn it into a lesson about whether that action could have happened earlier.&lt;br&gt;
It can also generate action items based on root-cause categories such as deployment, capacity, dependency, configuration, and human error.&lt;br&gt;
This is a deliberate engineering choice.&lt;br&gt;
For incident response, a deterministic first draft has an advantage: its content comes from recorded incident data rather than from an external model inventing plausible-sounding details.&lt;br&gt;
It is still a draft. The repository explicitly positions it as something that should support, rather than replace, human review.&lt;br&gt;
What this looks like during a recurring incident&lt;br&gt;
Consider a new &lt;code&gt;payments-api&lt;/code&gt; incident with:&lt;/p&gt;

&lt;p&gt;text&lt;br&gt;
Description:&lt;br&gt;
payments-api is returning 502 upstream timeouts,&lt;br&gt;
and the connection pool looks saturated.&lt;/p&gt;

&lt;p&gt;Service:&lt;br&gt;
payments-api&lt;/p&gt;

&lt;p&gt;Severity:&lt;br&gt;
SEV1&lt;/p&gt;

&lt;p&gt;Error signature:&lt;br&gt;
502-upstream-timeout&lt;/p&gt;

&lt;p&gt;The project's tests verify that the corresponding historical payments incident ranks first when the same service and error signature are supplied.&lt;br&gt;
The resulting fingerprint can expose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;matching service,&lt;/li&gt;
&lt;li&gt;matching error signature,&lt;/li&gt;
&lt;li&gt;matching severity,&lt;/li&gt;
&lt;li&gt;textual similarity,&lt;/li&gt;
&lt;li&gt;recency.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The associated runbook can then be surfaced with an explanation of why the historical incident matched.&lt;br&gt;
The responder can follow the runbook, record the actions taken, resolve the incident, generate the postmortem, and finally report whether the runbook worked.&lt;br&gt;
That last step changes the future.&lt;br&gt;
The next responder does not inherit exactly the same static recommendation. The runbook's rating has been updated using the new evidence.&lt;/p&gt;

&lt;p&gt;What I learned from this implementation:&lt;br&gt;
The biggest lesson is that incident memory is more useful when it is structured around decisions.&lt;br&gt;
Simply remembering that an incident happened is not enough.&lt;/p&gt;

&lt;p&gt;The system needs to preserve relationships:&lt;br&gt;
incident → root cause → response step → runbook → outcome**&lt;br&gt;
The second lesson is that explainability can be built into the data model instead of being added later as a presentation layer.&lt;br&gt;
The similarity engine does not only return a number. It returns the dimensions behind the number.&lt;br&gt;
The third lesson is that historical information has a shelf life.&lt;br&gt;
The project implements rating decay toward the 1200 baseline when a runbook has not been used recently. The intention is straightforward: infrastructure changes, dependencies change, and an old successful procedure should not automatically retain the same level of trust forever.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the project does not solve
&lt;/h2&gt;

&lt;p&gt;There are important limitations.&lt;br&gt;
The similarity system uses TF-IDF rather than semantic embeddings. That means two incidents describing the same underlying problem with very different vocabulary may not match strongly.&lt;br&gt;
The runbook rating is also not a scientific measure of correctness. It is an operational trust signal based on the outcomes recorded by the system.&lt;br&gt;
The postmortem generator is deterministic and template-driven. It cannot independently investigate an incident or discover facts that were never recorded.&lt;br&gt;
Finally, the application currently uses SQLite. The repository describes this as appropriate for a team or demonstration environment, with a production-scale deployment potentially moving to PostgreSQL.&lt;br&gt;
These limitations are useful because they define exactly where the current system ends.&lt;/p&gt;

&lt;h2&gt;
  
  
  One important architectural clarification
&lt;/h2&gt;

&lt;p&gt;The original project brief called for a Hindsight-based memory story. However, the uploaded repository does &lt;strong&gt;not&lt;/strong&gt; demonstrate an actual Hindsight SDK, retain/recall API, or Hindsight service integration.&lt;br&gt;
The implemented system instead provides its own persistent incident-history mechanism through SQLAlchemy/SQLite, similarity search, runbook usage records, and postmortem data.&lt;br&gt;
That distinction matters.&lt;/p&gt;

&lt;p&gt;It would be inaccurate to describe TF-IDF incident retrieval as Hindsight memory simply because both systems involve historical information.&lt;br&gt;
The current implementation is its own deterministic incident-memory workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The most useful part of an incident-response system may not be the ability to find an old incident.&lt;br&gt;
It is the ability to turn that old incident into evidence that can be inspected, acted on, and eventually updated.&lt;/p&gt;

&lt;p&gt;Our Incident Response Agent connects those stages:&lt;br&gt;
find a precedent → understand why it matches → inspect its runbook → record the current response → measure the result → update future trust**&lt;br&gt;
For me, that is the core engineering lesson from the project.&lt;br&gt;
During an outage, historical knowledge should not just be available.&lt;br&gt;
It should be structured enough to help someone decide what deserves attention next — while still leaving the final decision with the engineer handling the incident.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>automation</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
