<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Murali Karthik</title>
    <description>The latest articles on DEV Community by Murali Karthik (@murali_karthik_412486a7f5).</description>
    <link>https://dev.to/murali_karthik_412486a7f5</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150804%2F26105662-aaad-49fd-99bf-51ccdac9fd57.jpg</url>
      <title>DEV Community: Murali Karthik</title>
      <link>https://dev.to/murali_karthik_412486a7f5</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/murali_karthik_412486a7f5"/>
    <language>en</language>
    <item>
      <title>AI Incident Response Agent</title>
      <dc:creator>Murali Karthik</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:26:04 +0000</pubDate>
      <link>https://dev.to/murali_karthik_412486a7f5/ai-incident-response-agent-2ljm</link>
      <guid>https://dev.to/murali_karthik_412486a7f5/ai-incident-response-agent-2ljm</guid>
      <description>&lt;p&gt;AI Incident Response Agent is an AI-powered, human-in-the-loop incident response system designed to investigate, diagnose, and recover from application incidents using real evidence instead of relying only on telemetry.&lt;/p&gt;

&lt;p&gt;The agent analyzes current incident metrics, understands the application's GitHub repository, identifies relevant source/configuration files, and uses historical incident knowledge through Hindsight to build an evidence-based diagnosis.&lt;/p&gt;

&lt;p&gt;Instead of blindly assuming a root cause from metrics, the agent cross-checks whether the suspected component actually exists in the application. When evidence conflicts or is insufficient, it explicitly marks the incident as requiring further investigation.&lt;/p&gt;

&lt;p&gt;The system supports a closed-loop workflow:&lt;/p&gt;

&lt;p&gt;Detect → Investigate → Diagnose → Human Approval → Recover → Verify → Learn&lt;/p&gt;

&lt;p&gt;Key capabilities include:&lt;/p&gt;

&lt;p&gt;🔍 Repository-aware incident investigation&lt;br&gt;
🤖 LLM-based root-cause analysis&lt;br&gt;
📊 Telemetry analysis&lt;br&gt;
🧠 Hindsight-powered historical incident recall&lt;br&gt;
👨‍💻 Human approval before state-changing recovery actions&lt;br&gt;
🔄 Investigation and re-investigation when evidence is insufficient&lt;br&gt;
✅ Post-recovery verification&lt;br&gt;
📚 Retaining confirmed incident learnings for future incidents&lt;br&gt;
🛡️ Evidence-based reasoning to avoid unsupported diagnoses&lt;/p&gt;

&lt;p&gt;The goal is to move incident response from “metrics say something is wrong” to “the agent investigates the actual application, explains why it believes something is wrong, and takes controlled action with human oversight.”&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>automation</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
