<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gudoor Vijaya</title>
    <description>The latest articles on DEV Community by Gudoor Vijaya (@gudoor_vijaya_1b9768335b8).</description>
    <link>https://dev.to/gudoor_vijaya_1b9768335b8</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147270%2F6f9a1aa7-de5c-40f0-ae8c-c533fd22c791.png</url>
      <title>DEV Community: Gudoor Vijaya</title>
      <link>https://dev.to/gudoor_vijaya_1b9768335b8</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gudoor_vijaya_1b9768335b8"/>
    <language>en</language>
    <item>
      <title>What Building an AI SRE Agent Taught Us</title>
      <dc:creator>Gudoor Vijaya</dc:creator>
      <pubDate>Mon, 28 Sep 2026 13:46:42 +0000</pubDate>
      <link>https://dev.to/gudoor_vijaya_1b9768335b8/what-building-an-ai-sre-agent-taught-us-3nlb</link>
      <guid>https://dev.to/gudoor_vijaya_1b9768335b8/what-building-an-ai-sre-agent-taught-us-3nlb</guid>
      <description>&lt;p&gt;What Building an AI SRE Agent Taught Us&lt;/p&gt;

&lt;p&gt;The hardest part of building an AI SRE agent was not generating a diagnosis.&lt;/p&gt;

&lt;p&gt;It was designing the system around that diagnosis.&lt;/p&gt;

&lt;p&gt;OpsMind started from a straightforward problem: incident responders repeatedly encounter similar failures, but the useful experience from one incident is not always available when another incident occurs.&lt;/p&gt;

&lt;p&gt;We built the system around persistent agent memory so that previous incident outcomes could become part of future investigations.&lt;/p&gt;

&lt;p&gt;The result was a workflow combining current telemetry, Hindsight memory, AI reasoning, human approval, simulated remediation, and retained learning.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Architecture
&lt;/h2&gt;

&lt;p&gt;OpsMind follows this lifecycle:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Incident → Evidence → Recall → Diagnosis → Approval → Remediation → Retain&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Each stage has a different responsibility.&lt;/p&gt;

&lt;p&gt;Current logs and metrics provide evidence.&lt;/p&gt;

&lt;p&gt;Hindsight provides historical operational context.&lt;/p&gt;

&lt;p&gt;The AI SRE agent reasons over the information.&lt;/p&gt;

&lt;p&gt;A human approves remediation.&lt;/p&gt;

&lt;p&gt;The successful outcome becomes memory.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmqvmegh75c0ynjttqipn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmqvmegh75c0ynjttqipn.png" alt=" " width="469" height="213"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Figure 1 — OpsMind turns incident resolution into a persistent learning loop.&lt;/p&gt;

&lt;p&gt;The architecture looks simple when written as a sequence.&lt;/p&gt;

&lt;p&gt;Implementing it exposed several engineering decisions that were less obvious.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 1: Memory Is Not the Same as Knowledge
&lt;/h2&gt;

&lt;p&gt;Adding a vector or memory database does not automatically make an agent knowledgeable.&lt;/p&gt;

&lt;p&gt;The system needs to decide:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What should be remembered?&lt;/li&gt;
&lt;li&gt;When should it be remembered?&lt;/li&gt;
&lt;li&gt;What context should be stored?&lt;/li&gt;
&lt;li&gt;When should it be retrieved?&lt;/li&gt;
&lt;li&gt;How should retrieved information influence reasoning?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;OpsMind retains incident outcomes rather than indiscriminately storing every interaction.&lt;/p&gt;

&lt;p&gt;A successful incident creates a learning record containing the incident, diagnosis, actions, outcome, and resolution status.&lt;/p&gt;

&lt;p&gt;That makes the memory operationally meaningful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 2: Current Evidence Must Remain Primary
&lt;/h2&gt;

&lt;p&gt;Persistent memory can be powerful, but it creates a new failure mode.&lt;/p&gt;

&lt;p&gt;An agent may find an old incident that looks similar and assume that the same cause applies again.&lt;/p&gt;

&lt;p&gt;OpsMind addresses this by giving current logs and metrics priority.&lt;/p&gt;

&lt;p&gt;The AI receives current incident information and evidence before historical context is incorporated into the reasoning process.&lt;/p&gt;

&lt;p&gt;For INC-008, current telemetry showed 6.1-second latency, 26% HTTP 500 errors, and 97% database connection utilization.&lt;/p&gt;

&lt;p&gt;Hindsight then added historical incidents involving related database behavior.&lt;/p&gt;

&lt;p&gt;The model was instructed not to blindly copy previous remediation.&lt;/p&gt;

&lt;p&gt;That separation is one of the most important design choices in the system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 3: Retrieval Needs to Be Testable
&lt;/h2&gt;

&lt;p&gt;It is easy to say that an agent has memory.&lt;/p&gt;

&lt;p&gt;It is harder to demonstrate that the memory actually changes future behavior.&lt;/p&gt;

&lt;p&gt;We used INC-008 as a learning event.&lt;/p&gt;

&lt;p&gt;The incident was analyzed and resolved.&lt;/p&gt;

&lt;p&gt;Its successful outcome was retained.&lt;/p&gt;

&lt;p&gt;Then INC-007 was analyzed afterward.&lt;/p&gt;

&lt;p&gt;Hindsight returned INC-008 as historical context.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fusv15whaomhh2dxfcjhv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fusv15whaomhh2dxfcjhv.png" alt=" " width="800" height="404"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Figure 2 — The later INC-007 investigation recalls the newly retained INC-008 experience.&lt;/p&gt;

&lt;p&gt;This gave us a concrete test:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Can a future investigation retrieve an experience created by a previous investigation?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The answer in this workflow was yes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 4: Human Approval Changes the Architecture
&lt;/h2&gt;

&lt;p&gt;It would have been possible to connect the AI's recommended actions directly to an execution layer.&lt;/p&gt;

&lt;p&gt;We deliberately did not.&lt;/p&gt;

&lt;p&gt;OpsMind inserts a human approval gate between diagnosis and remediation.&lt;/p&gt;

&lt;p&gt;The agent can produce:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Root cause&lt;/li&gt;
&lt;li&gt;Evidence&lt;/li&gt;
&lt;li&gt;Reasoning&lt;/li&gt;
&lt;li&gt;Recommended actions&lt;/li&gt;
&lt;li&gt;Confidence&lt;/li&gt;
&lt;li&gt;Runbook&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The operator then reviews the proposed action.&lt;/p&gt;

&lt;p&gt;Only after approval does the resolution workflow proceed.&lt;/p&gt;

&lt;p&gt;The current remediation is simulated rather than connected to production infrastructure.&lt;/p&gt;

&lt;p&gt;This keeps the prototype's operational boundary explicit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 5: Agent Memory Creates a Feedback Loop
&lt;/h2&gt;

&lt;p&gt;Traditional incident handling often ends with:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Resolved → Closed&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The resolution may remain inside a ticketing system or someone's memory.&lt;/p&gt;

&lt;p&gt;OpsMind changes the lifecycle to:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Resolved → Retained → Available to Future Investigations&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;After successful remediation, the incident outcome is retained in Hindsight.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0rjxj8o24abnkjpjaued.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0rjxj8o24abnkjpjaued.png" alt=" " width="800" height="425"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Figure 3 — A successful incident outcome becomes persistent organizational memory.&lt;/p&gt;

&lt;p&gt;That creates a feedback loop.&lt;/p&gt;

&lt;p&gt;The next incident can use the previous experience.&lt;/p&gt;

&lt;p&gt;The next successful resolution can then become another memory.&lt;/p&gt;

&lt;p&gt;Over time, the system can accumulate operational experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 6: Integration Problems Are Part of the Work
&lt;/h2&gt;

&lt;p&gt;The Hindsight integration also produced a technical problem that was not visible in the initial architecture.&lt;/p&gt;

&lt;p&gt;The first recall implementation encountered:&lt;/p&gt;

&lt;p&gt;text&lt;br&gt;
Timeout context manager should be used inside a task&lt;/p&gt;

&lt;p&gt;The issue appeared because the synchronous memory client was being used within the asynchronous FastAPI environment.&lt;/p&gt;

&lt;p&gt;We changed the implementation to use Hindsight's asynchronous recall operation:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
response = await hindsight_client.arecall(&lt;br&gt;
    bank_id=HINDSIGHT_BANK_ID,&lt;br&gt;
    query=query,&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;and explicitly closed the client afterward:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
await hindsight_client.aclose()&lt;/p&gt;

&lt;p&gt;This solved the event-loop problem.&lt;/p&gt;

&lt;p&gt;The lesson was broader than this particular error.&lt;/p&gt;

&lt;p&gt;When integrating external services into an agent system, application execution models matter just as much as API design.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 7: A Good Demo Should Prove a Behavior
&lt;/h2&gt;

&lt;p&gt;For an AI agent, showing a generated answer is not enough.&lt;/p&gt;

&lt;p&gt;We wanted the workflow to demonstrate a state change.&lt;/p&gt;

&lt;p&gt;The sequence is:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;INC-008&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;↓&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Diagnosis&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;↓&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Human approval&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;↓&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Resolution&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;↓&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Memory retained&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;↓&lt;/p&gt;

&lt;p&gt;&lt;em&gt;INC-007&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;↓&lt;/p&gt;

&lt;p&gt;&lt;em&gt;INC-008 recalled as historical context&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F05ehpha5327ap2r7vgqn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F05ehpha5327ap2r7vgqn.png" alt=" " width="800" height="425"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Figure 4 — The memory relationship becomes observable when the later incident retrieves the earlier outcome.&lt;/p&gt;

&lt;p&gt;This is more meaningful than simply showing a chatbot response because it demonstrates persistent state across investigations.&lt;/p&gt;

&lt;h2&gt;
  
  
  What We Would Improve Next
&lt;/h2&gt;

&lt;p&gt;There are several areas that would need additional engineering for a production SRE environment.&lt;/p&gt;

&lt;p&gt;First, real observability integrations would be needed instead of controlled incident datasets.&lt;/p&gt;

&lt;p&gt;Second, remediation would need stronger safeguards, authorization, audit logging, rollback support, and failure handling.&lt;/p&gt;

&lt;p&gt;Third, memory retrieval would need continuous evaluation to determine whether retrieved incidents are genuinely relevant.&lt;/p&gt;

&lt;p&gt;Fourth, the system would need mechanisms for handling outdated or contradictory historical knowledge.&lt;/p&gt;

&lt;p&gt;Persistent memory introduces a new operational responsibility:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;not only deciding what the agent should remember, but also deciding when remembered information should no longer influence decisions.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Main Takeaway
&lt;/h2&gt;

&lt;p&gt;The most useful lesson from OpsMind was that agent memory is not an isolated feature.&lt;/p&gt;

&lt;p&gt;It changes the lifecycle of the application.&lt;/p&gt;

&lt;p&gt;Without memory:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Incident → Diagnosis → Resolution&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;With persistent memory:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Incident → Evidence → Historical Context → Diagnosis → Resolution → Learning → Future Incident&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Hindsight provides the persistent layer that makes the second workflow possible.&lt;/p&gt;

&lt;p&gt;But the rest of the architecture still matters.&lt;/p&gt;

&lt;p&gt;Evidence grounds the diagnosis.&lt;/p&gt;

&lt;p&gt;The AI interprets the evidence.&lt;/p&gt;

&lt;p&gt;Human approval controls remediation.&lt;/p&gt;

&lt;p&gt;Successful outcomes create new memory.&lt;/p&gt;

&lt;p&gt;Future investigations can then retrieve those experiences.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Building OpsMind changed our understanding of what an AI SRE agent should do.&lt;/p&gt;

&lt;p&gt;The objective is not simply to create an AI that can explain why an incident happened.&lt;/p&gt;

&lt;p&gt;The more interesting objective is to build a system that can &lt;em&gt;learn from successful incident resolution and make that experience available when the next problem appears.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;INC-008 demonstrated that loop.&lt;/p&gt;

&lt;p&gt;It was investigated using current telemetry and historical context, resolved through a controlled workflow, and retained as organizational memory.&lt;/p&gt;

&lt;p&gt;Later, INC-007 could retrieve that experience.&lt;/p&gt;

&lt;p&gt;That is the behavior we wanted from persistent agent memory:&lt;/p&gt;

&lt;p&gt;not a replacement for engineering judgment, but a mechanism for making previous operational experience reusable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/vectorize-io/hindsight?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hindsight on GitHub&lt;/a&gt;&lt;br&gt;
&lt;a href="https://hindsight.vectorize.io/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Hindsight Documentation&lt;/a&gt;&lt;br&gt;
&lt;a href="https://vectorize.io/what-is-agent-memory?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;What Is Agent Memory? — Vectorize&lt;/a&gt;&lt;br&gt;
&lt;a href="https://github.com/srivaniyadav174/OpsMind-AI-SRE-Incident-Response-Memory-Agent.git?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;OpsMind — GitHub Repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
