<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Arooba Hasan</title>
    <description>The latest articles on DEV Community by Arooba Hasan (@arooba_hasan_ce04e5314f88).</description>
    <link>https://dev.to/arooba_hasan_ce04e5314f88</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150147%2F18809492-ba4b-45cb-bb3f-66421a7e6a51.png</url>
      <title>DEV Community: Arooba Hasan</title>
      <link>https://dev.to/arooba_hasan_ce04e5314f88</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/arooba_hasan_ce04e5314f88"/>
    <language>en</language>
    <item>
      <title>Incident Response Agent — Intelligent Incident Analysis and Response Assistance</title>
      <dc:creator>Arooba Hasan</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:32:23 +0000</pubDate>
      <link>https://dev.to/arooba_hasan_ce04e5314f88/incident-response-agent-intelligent-incident-analysis-and-response-assistance-fh4</link>
      <guid>https://dev.to/arooba_hasan_ce04e5314f88/incident-response-agent-intelligent-incident-analysis-and-response-assistance-fh4</guid>
      <description>&lt;p&gt;When an incident happens at 3 a.m., having historical knowledge is valuable—but only if the engineer can understand why that history is relevant and whether it is still trustworthy.&lt;/p&gt;

&lt;p&gt;That was one of the ideas that stood out to me while examining our Incident Response Agent. The system does more than store previous incidents and search through them. It tries to turn incident history into practical guidance while making the reasoning behind that guidance visible.&lt;/p&gt;

&lt;p&gt;The most important lesson for me was that an incident-response assistant should not behave as if historical data is automatically correct.&lt;/p&gt;

&lt;p&gt;A previous runbook may have worked under very different circumstances. An incident may look similar because two services share the same error signature, while the underlying cause is different. A runbook that worked six months ago may also be less trustworthy today because infrastructure and configuration change.&lt;/p&gt;

&lt;p&gt;Our implementation therefore treats historical memory as evidence—not as an unquestionable answer.&lt;/p&gt;

&lt;p&gt;The problem: historical incidents are useful, but history gets stale&lt;/p&gt;

&lt;p&gt;Incident response often depends on institutional knowledge.&lt;/p&gt;

&lt;p&gt;An engineer sees an error and asks questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Have we seen this before?&lt;/li&gt;
&lt;li&gt;What happened last time?&lt;/li&gt;
&lt;li&gt;Which runbook did we use?&lt;/li&gt;
&lt;li&gt;Did that runbook actually work?&lt;/li&gt;
&lt;li&gt;Was the previous incident really comparable to this one?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A basic incident database can answer the first few questions by providing search results. But a list of historical incidents does not necessarily tell the responder which precedent deserves attention.&lt;/p&gt;

&lt;p&gt;Our project approaches the problem by creating an incident fingerprint.&lt;/p&gt;

&lt;p&gt;The similarity engine looks at five dimensions:&lt;/p&gt;

&lt;p&gt;WEIGHTS = {&lt;br&gt;
    "text_similarity": 0.40,&lt;br&gt;
    "service_match": 0.20,&lt;br&gt;
    "error_signature_match": 0.20,&lt;br&gt;
    "severity_match": 0.10,&lt;br&gt;
    "recency": 0.10,&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;Instead of producing only one unexplained number, the system keeps these dimensions separate.&lt;/p&gt;

&lt;p&gt;That distinction matters during an incident. A responder can see that a previous incident matched because it involved the same service and error signature, rather than being told that two incidents are simply “87% similar.”&lt;/p&gt;

&lt;p&gt;Making similarity explainable&lt;/p&gt;

&lt;p&gt;The text component uses TF-IDF and cosine similarity over information such as the incident title, description, error signature, service, and root-cause summaries.&lt;/p&gt;

&lt;p&gt;It is deliberately local and deterministic:&lt;/p&gt;

&lt;p&gt;vectorizer = TfidfVectorizer(&lt;br&gt;
    stop_words="english",&lt;br&gt;
    max_features=2000&lt;br&gt;
)&lt;br&gt;
tfidf = vectorizer.fit_transform(corpus)&lt;br&gt;
text_sims = cosine_similarity(&lt;br&gt;
    query_vec,&lt;br&gt;
    doc_vecs&lt;br&gt;
).flatten()&lt;/p&gt;

&lt;p&gt;The resulting fingerprint contains:&lt;/p&gt;

&lt;p&gt;{&lt;br&gt;
    "text_similarity": ...,&lt;br&gt;
    "service_match": ...,&lt;br&gt;
    "error_signature_match": ...,&lt;br&gt;
    "severity_match": ...,&lt;br&gt;
    "recency": ...&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;The system then combines those values using explicit weights.&lt;/p&gt;

&lt;p&gt;This design has an important reliability advantage: the recommendation can be inspected.&lt;/p&gt;

&lt;p&gt;For example, a result can be explained through conditions such as:&lt;/p&gt;

&lt;p&gt;same service, identical error signature, same severity class&lt;/p&gt;

&lt;p&gt;That is much easier for an engineer to challenge than an opaque similarity score.&lt;/p&gt;

&lt;p&gt;If the recommendation looks wrong, the responder has some information about why it appeared.&lt;/p&gt;

&lt;p&gt;But similarity is not enough&lt;/p&gt;

&lt;p&gt;Finding a similar incident is only half of the problem.&lt;/p&gt;

&lt;p&gt;Suppose a runbook was used during that historical incident. Should we automatically recommend it?&lt;/p&gt;

&lt;p&gt;Our project deliberately avoids treating a simple success percentage as the complete answer.&lt;/p&gt;

&lt;p&gt;Instead, each runbook starts with an Elo-style rating of 1200.&lt;/p&gt;

&lt;p&gt;The unusual part is that the “opponent” is the incident severity.&lt;/p&gt;

&lt;p&gt;SEVERITY_DIFFICULTY = {&lt;br&gt;
    "SEV1": 1600,&lt;br&gt;
    "SEV2": 1400,&lt;br&gt;
    "SEV3": 1200,&lt;br&gt;
    "SEV4": 1000,&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;A successful runbook on a SEV1 incident therefore provides stronger evidence than a successful runbook on a SEV4 incident.&lt;/p&gt;

&lt;p&gt;Likewise, failing against a routine SEV4 incident costs more rating than failing against a difficult SEV1 incident.&lt;/p&gt;

&lt;p&gt;The rating update is based on the standard Elo expectation formula:&lt;/p&gt;

&lt;p&gt;expected = expected_score(runbook_rating, opponent)&lt;br&gt;
new_rating = runbook_rating + K_FACTOR * (actual - expected)&lt;/p&gt;

&lt;p&gt;This is not intended to prove that a runbook is objectively correct. It is a way of representing accumulated evidence about how the runbook has behaved in incidents of different severity.&lt;/p&gt;

&lt;p&gt;Historical success also has a shelf life&lt;/p&gt;

&lt;p&gt;One of the most useful reliability ideas in the project is rating decay.&lt;/p&gt;

&lt;p&gt;A runbook can have a strong historical rating and still be stale.&lt;/p&gt;

&lt;p&gt;Infrastructure changes. Dependencies change. Configuration changes. Teams change operational procedures.&lt;/p&gt;

&lt;p&gt;The implementation therefore pulls an unused rating back toward the 1200 baseline:&lt;/p&gt;

&lt;p&gt;decay_factor = 0.5 ** (&lt;br&gt;
    days_since_last_use / DECAY_HALF_LIFE_DAYS&lt;br&gt;
)&lt;br&gt;
return round(&lt;br&gt;
    STARTING_RATING&lt;br&gt;
    + (rating - STARTING_RATING) * decay_factor,&lt;br&gt;
    1&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;The half-life is 120 days.&lt;/p&gt;

&lt;p&gt;That means the system does not interpret an old successful result as permanent proof.&lt;/p&gt;

&lt;p&gt;This is exactly the kind of behavior I think an incident-response system needs. Historical memory should become useful knowledge, but knowledge should remain subject to validation.&lt;/p&gt;

&lt;p&gt;Before and after: what changes for the responder?&lt;/p&gt;

&lt;p&gt;Consider a new payments-api incident reporting 502 upstream timeouts.&lt;/p&gt;

&lt;p&gt;Without the matching and ranking workflow, an engineer might search incident history manually, find several payment incidents, open them individually, and decide which runbook seems relevant.&lt;/p&gt;

&lt;p&gt;With the implemented workflow, the system searches resolved incidents and ranks the closest precedents.&lt;/p&gt;

&lt;p&gt;The project includes a test specifically checking that an incident with the same service and error signature is ranked appropriately:&lt;/p&gt;

&lt;p&gt;assert top_incident.service == "payments-api"&lt;br&gt;
assert fingerprint["service_match"] == 1.0&lt;br&gt;
assert fingerprint["error_signature_match"] == 1.0&lt;/p&gt;

&lt;p&gt;The responder therefore receives more than a historical incident title. They receive the similarity fingerprint and the runbooks associated with that historical response.&lt;/p&gt;

&lt;p&gt;But the responder still makes the final decision.&lt;/p&gt;

&lt;p&gt;That distinction is important.&lt;/p&gt;

&lt;p&gt;The system learns from feedback—but does not silently rewrite history&lt;/p&gt;

&lt;p&gt;When a runbook is used, the system accepts three outcomes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;worked&lt;/li&gt;
&lt;li&gt;partial&lt;/li&gt;
&lt;li&gt;failed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The feedback endpoint records the usage and updates the runbook’s Elo rating.&lt;/p&gt;

&lt;p&gt;new_rating, delta = elo.update_elo(&lt;br&gt;
    rb.elo_rating,&lt;br&gt;
    severity,&lt;br&gt;
    payload.outcome&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;The previous and new ratings are also stored with the runbook usage.&lt;/p&gt;

&lt;p&gt;This creates a feedback loop:&lt;/p&gt;

&lt;p&gt;incident → historical match → runbook recommendation → human action → outcome → updated trust&lt;/p&gt;

&lt;p&gt;The next recommendation can therefore benefit from previous operational experience.&lt;/p&gt;

&lt;p&gt;However, this feedback mechanism is still only as good as the evidence entered into the system. A wrong classification or an inaccurately recorded outcome can influence future recommendations.&lt;/p&gt;

&lt;p&gt;That is another reason not to treat the rating as an absolute truth.&lt;/p&gt;

&lt;p&gt;Postmortems are deliberately drafts&lt;/p&gt;

&lt;p&gt;The same principle appears in postmortem generation.&lt;/p&gt;

&lt;p&gt;The project does not call an external LLM to invent a postmortem. Instead, postmortem_gen.py constructs a deterministic draft from the incident’s structured data.&lt;/p&gt;

&lt;p&gt;It uses recorded root causes, the timeline, severity, error signature, and MTTR.&lt;/p&gt;

&lt;p&gt;For example, it can identify the first detection step and the first mitigating action from the timeline.&lt;/p&gt;

&lt;p&gt;detection_steps = [&lt;br&gt;
    s for s in steps&lt;br&gt;
    if s.step_type == "detection"&lt;br&gt;
]&lt;br&gt;
fix_steps = [&lt;br&gt;
    s for s in steps&lt;br&gt;
    if s.step_type in ("fix", "mitigation", "resolution")&lt;br&gt;
]&lt;/p&gt;

&lt;p&gt;The project explicitly treats the result as a draft.&lt;/p&gt;

&lt;p&gt;That is a good reliability boundary. A generated postmortem can organize information and save time, but it cannot know everything a human incident review may need to consider—such as customer impact, organizational context, or the appropriate blameless framing.&lt;/p&gt;

&lt;p&gt;The implementation itself makes this limitation clear.&lt;/p&gt;

&lt;p&gt;What the tests tell us&lt;/p&gt;

&lt;p&gt;The repository contains 17 automated tests, and I think they are an important part of the project’s story.&lt;/p&gt;

&lt;p&gt;The tests cover the similarity engine, Elo calculations, and API lifecycle.&lt;/p&gt;

&lt;p&gt;For example, the Elo tests verify that:&lt;/p&gt;

&lt;p&gt;assert delta_hard &amp;gt; delta_easy &amp;gt; 0&lt;/p&gt;

&lt;p&gt;for successful runbooks against different incident severities.&lt;/p&gt;

&lt;p&gt;They also verify that failure against an easier incident costs more:&lt;/p&gt;

&lt;p&gt;assert delta_easy &amp;lt; delta_hard &amp;lt; 0&lt;/p&gt;

&lt;p&gt;The similarity tests verify that the fingerprint contains all five dimensions and that results are ordered by score.&lt;/p&gt;

&lt;p&gt;The full API lifecycle is also tested.&lt;/p&gt;

&lt;p&gt;When I ran the repository’s test suite, all 17 tests passed. The test run did produce Python deprecation warnings around datetime.utcnow(), which is a useful maintenance issue to address rather than something to hide.&lt;/p&gt;

&lt;p&gt;Where the system should not be trusted&lt;/p&gt;

&lt;p&gt;The project documents several limitations, and they are important.&lt;/p&gt;

&lt;p&gt;First, TF-IDF similarity is lexical.&lt;/p&gt;

&lt;p&gt;If two engineers describe the same root cause using completely different terminology, the system may not recognize the relationship as strongly as a semantic embedding system might.&lt;/p&gt;

&lt;p&gt;Second, the project uses a single SQLite database.&lt;/p&gt;

&lt;p&gt;That is appropriate for the demonstrated application, but the README itself identifies PostgreSQL as a future direction before the system holds a much larger organization’s incident history.&lt;/p&gt;

&lt;p&gt;Third, postmortems are deterministic drafts rather than complete incident reviews.&lt;/p&gt;

&lt;p&gt;Finally, the runbook rating is evidence, not proof. A “Battle-tested” label does not mean a runbook is guaranteed to work during the next incident.&lt;/p&gt;

&lt;p&gt;One important implementation boundary&lt;/p&gt;

&lt;p&gt;There is also a distinction worth making clear about memory technology.&lt;/p&gt;

&lt;p&gt;The supplied repository does not contain an integration with Hindsight. There are no Hindsight dependencies or retain/recall calls in this version.&lt;/p&gt;

&lt;p&gt;Instead, the project’s institutional memory is implemented through its own SQLite data model, incident history, TF-IDF matching, runbook usage records, Elo ratings, and postmortem data.&lt;/p&gt;

&lt;p&gt;That does not make the memory mechanism uninteresting. In fact, it makes the reliability lesson easier to see: regardless of the underlying memory technology, historical information needs context, explainability, freshness, and human validation.&lt;/p&gt;

&lt;p&gt;Lessons learned&lt;/p&gt;

&lt;p&gt;The biggest lesson from this project is that an incident-response assistant should not try to eliminate human judgment.&lt;/p&gt;

&lt;p&gt;It should make human judgment faster and better informed.&lt;/p&gt;

&lt;p&gt;For me, the most valuable engineering choices were therefore not simply about finding similar incidents. They were the safeguards around that information:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;expose why an incident matched;&lt;/li&gt;
&lt;li&gt;distinguish different dimensions of similarity;&lt;/li&gt;
&lt;li&gt;track whether recommended runbooks actually worked;&lt;/li&gt;
&lt;li&gt;account for incident severity when updating trust;&lt;/li&gt;
&lt;li&gt;decay stale ratings;&lt;/li&gt;
&lt;li&gt;generate postmortems from recorded evidence;&lt;/li&gt;
&lt;li&gt;keep generated postmortems as drafts;&lt;/li&gt;
&lt;li&gt;test the ranking and feedback logic;&lt;/li&gt;
&lt;li&gt;and document the boundaries of the system.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That approach produces something more useful than an assistant that confidently gives an answer.&lt;/p&gt;

&lt;p&gt;It produces an assistant that can show its evidence—and gives the engineer enough context to decide whether that evidence still applies.&lt;/p&gt;

&lt;p&gt;Conclusion&lt;/p&gt;

&lt;p&gt;Incident history becomes valuable when it can influence the next response without becoming a source of blind confidence.&lt;/p&gt;

&lt;p&gt;Our Incident Response Agent takes a practical approach: store structured incident history, find comparable incidents, explain the match, rank runbooks using outcome history, reduce confidence in stale knowledge, and turn the response timeline into a reviewable postmortem draft.&lt;/p&gt;

&lt;p&gt;The system is not production-ready infrastructure, and it does not prove that its recommendations are correct. Its value is in creating a structured feedback loop around incident knowledge.&lt;/p&gt;

&lt;p&gt;That is the reliability lesson I would carry into a larger system: memory should inform the responder, not replace the responder.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>devops</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
