<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Laaqshya B</title>
    <description>The latest articles on DEV Community by Laaqshya B (@laaqshya_b_9fc73c5ad8f0b2).</description>
    <link>https://dev.to/laaqshya_b_9fc73c5ad8f0b2</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150715%2Fca8e7350-d2bb-4919-9924-61bbdde27b3a.png</url>
      <title>DEV Community: Laaqshya B</title>
      <link>https://dev.to/laaqshya_b_9fc73c5ad8f0b2</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/laaqshya_b_9fc73c5ad8f0b2"/>
    <language>en</language>
    <item>
      <title>Building a Cloud SRE Agent That Learns From Every Incident</title>
      <dc:creator>Laaqshya B</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:16:40 +0000</pubDate>
      <link>https://dev.to/laaqshya_b_9fc73c5ad8f0b2/building-a-cloud-sre-agent-that-learns-from-every-incident-1c37</link>
      <guid>https://dev.to/laaqshya_b_9fc73c5ad8f0b2/building-a-cloud-sre-agent-that-learns-from-every-incident-1c37</guid>
      <description>&lt;p&gt;Building a Cloud SRE Agent That Learns From Every Incident&lt;/p&gt;

&lt;p&gt;Production incidents are rarely completely new.&lt;/p&gt;

&lt;p&gt;A payment API timeout, a sudden error spike after deployment, or database connection-pool exhaustion may look like a new emergency. But in many cases, the organization has already seen the same warning signs, attempted a fix, and learned what worked—or what failed.&lt;/p&gt;

&lt;p&gt;The problem is that this knowledge is scattered across logs, incident tickets, chat messages, and postmortems. During an outage, engineers still have to search manually and reconstruct the same context under pressure.&lt;/p&gt;

&lt;p&gt;That is the problem we are solving with our memory-powered Cloud SRE Agent.&lt;/p&gt;

&lt;p&gt;he problem with ordinary AI incident assistants&lt;/p&gt;

&lt;p&gt;A normal AI assistant can summarize logs and suggest possible root causes. But if it has no persistent, reliable memory, every incident begins from zero.&lt;/p&gt;

&lt;p&gt;Even worse, storing every log snippet or LLM-generated conclusion as “memory” can be dangerous. A model may make an incorrect assumption, and if that assumption is saved without review, the same bad advice can influence future incidents.&lt;/p&gt;

&lt;p&gt;Memory is useful only when it is trustworthy.&lt;/p&gt;

&lt;p&gt;For an SRE agent, the goal is not simply to remember that an incident happened. The goal is to remember:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What service was affected&lt;/li&gt;
&lt;li&gt;What symptoms appeared&lt;/li&gt;
&lt;li&gt;What the validated root cause was&lt;/li&gt;
&lt;li&gt;Which remediation was attempted&lt;/li&gt;
&lt;li&gt;Whether the remediation actually resolved the incident&lt;/li&gt;
&lt;li&gt;Whether a human approved the action&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Our approach: verified incident memory&lt;/p&gt;

&lt;p&gt;Our Cloud SRE Agent monitors cloud logs and processes each potential incident in three stages:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Triage&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The agent determines whether an alert needs action, identifies the affected service, and measures severity.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Root-cause analysis&lt;br&gt;&lt;br&gt;
The agent examines the available signals and proposes the most likely cause.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Remediation&lt;br&gt;&lt;br&gt;
The agent recommends a concrete fix. With human approval, it can prepare a pull request instead of making an uncontrolled production change.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We extend this workflow with Hindsight, a persistent memory layer that lets the agent learn from resolved incidents over time.&lt;/p&gt;

&lt;p&gt;A safer five-stage memory pipeline&lt;/p&gt;

&lt;p&gt;Instead of storing raw logs or unverified AI conclusions directly, we use a controlled pipeline.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Incident capture&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The agent receives sanitized incident information such as service name, timestamp, error rate, deployment context, and relevant log patterns.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Payment API error rate increased to 18.4% shortly after a deployment. Database timeout errors are increasing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ol&gt;
&lt;li&gt;Memory recall&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Before generating a recommendation, the agent asks Hindsight for similar historical incidents.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Find previous Payment API incidents involving database timeouts, connection pools, or deployment-related error spikes.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This may retrieve a prior incident where increasing the connection-pool limit successfully resolved the same issue.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Analysis and remediation proposal&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The agent combines the current signals with the recalled evidence.&lt;/p&gt;

&lt;p&gt;Instead of producing a generic answer like “Check the database,” it can produce a grounded recommendation:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A similar incident was resolved by increasing the database connection-pool limit from 20 to 50. Current error patterns indicate possible connection-pool exhaustion. Recommend verifying the pool metrics and preparing the same change for human review.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ol&gt;
&lt;li&gt;Human approval&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The agent does not automatically treat every suggestion as correct.&lt;/p&gt;

&lt;p&gt;An engineer can approve the proposed remediation, request changes, reject it, or escalate the incident. This is especially important because incident response can affect production systems.&lt;/p&gt;

&lt;p&gt;The AI proposes. Humans retain control.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Outcome retention&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Only after the outcome is known does the system store a durable memory in Hindsight.&lt;/p&gt;

&lt;p&gt;A retained memory might look like this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Service: Payment API&lt;br&gt;&lt;br&gt;
Incident: Database timeout after deployment&lt;br&gt;&lt;br&gt;
Root cause: Database connection-pool exhaustion&lt;br&gt;&lt;br&gt;
Remediation: Increased connection-pool limit from 20 to 50&lt;br&gt;&lt;br&gt;
Outcome: Resolved after approval; error rate returned below 1%&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This turns an isolated incident into reusable operational knowledge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Hindsight matters
&lt;/h2&gt;

&lt;p&gt;Hindsight helps the Cloud SRE Agent go beyond a one-time chatbot response.&lt;/p&gt;

&lt;p&gt;It allows the system to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Recall similar incidents from weeks or months ago&lt;/li&gt;
&lt;li&gt;Connect current alerts with successful runbooks&lt;/li&gt;
&lt;li&gt;Identify recurring failure patterns&lt;/li&gt;
&lt;li&gt;Learn which remediation actions were approved&lt;/li&gt;
&lt;li&gt;Avoid repeating suggestions that previously failed&lt;/li&gt;
&lt;li&gt;Improve the specificity of recommendations over time&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The value is visible in a simple comparison.&lt;/p&gt;

&lt;h3&gt;
  
  
  Without persistent memory
&lt;/h3&gt;

&lt;p&gt;The agent sees a database timeout and gives a broad suggestion:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Check the database, deployment changes, and connection settings.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;With verified persistent memory&lt;/p&gt;

&lt;p&gt;The agent recalls two similar incidents and gives a focused recommendation:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Two resolved incidents match the current Payment API failure. Both involved connection-pool exhaustion after increased traffic or deployment changes. Recommend increasing the pool limit from 20 to 50 and verifying recovery before closing the incident.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the difference between an assistant that can generate text and an agent that can learn from operational experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Our live demonstration
&lt;/h2&gt;

&lt;p&gt;Our demo shows the agent handling two similar incidents.&lt;/p&gt;

&lt;p&gt;For the first incident, the agent performs triage, investigates the cause, and proposes a remediation. A human reviews the recommendation and approves the outcome for retention.&lt;/p&gt;

&lt;p&gt;For the second incident, the agent recalls the earlier validated memory from Hindsight. It recognizes the pattern faster, explains why the previous resolution is relevant, and proposes a more precise fix.&lt;/p&gt;

&lt;p&gt;This demonstrates the core idea behind our project:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every resolved incident should make the next incident easier to handle.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;What comes next&lt;/p&gt;

&lt;p&gt;Our next goal is to expand the agent with a runbook library built from approved incident outcomes. Instead of storing static documentation alone, the system can recommend runbooks based on the current service, error pattern, deployment context, and past success rate.&lt;/p&gt;

&lt;p&gt;Reliable AI operations are not about giving an agent unlimited autonomy. They are about combining AI reasoning, persistent memory, deterministic safeguards, and human approval.&lt;/p&gt;

&lt;p&gt;With Hindsight, our Cloud SRE Agent does not just react to incidents. It learns from them.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>devops</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
