<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Saipriya41</title>
    <description>The latest articles on DEV Community by Saipriya41 (@saipriya41).</description>
    <link>https://dev.to/saipriya41</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4148733%2F70e7edba-79db-4a1e-9d07-7a73fb7b1a17.png</url>
      <title>DEV Community: Saipriya41</title>
      <link>https://dev.to/saipriya41</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/saipriya41"/>
    <language>en</language>
    <item>
      <title>I Built a DevOps Agent That Remembers How Incidents Were Fixed</title>
      <dc:creator>Saipriya41</dc:creator>
      <pubDate>Tue, 29 Sep 2026 07:46:11 +0000</pubDate>
      <link>https://dev.to/saipriya41/i-built-a-devops-agent-that-remembers-how-incidents-were-fixed-11on</link>
      <guid>https://dev.to/saipriya41/i-built-a-devops-agent-that-remembers-how-incidents-were-fixed-11on</guid>
      <description>&lt;p&gt;Most incident-response tools can tell you that something is broken. The harder problem is remembering that you have already seen something like it before.&lt;/p&gt;

&lt;p&gt;I built a DevOps incident-response agent around that idea: instead of treating every production error as a completely new question, the agent retrieves relevant incidents from persistent memory, gives that context to an LLM, and stores the resulting diagnosis and resolution for the next time the same class of failure appears.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5p4p9b3boqnyu5zd5ufj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5p4p9b3boqnyu5zd5ufj.png" alt="The DevOps Output webapplication" width="799" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The interesting part wasn't simply connecting an LLM to a log stream. It was giving the agent a memory that could survive beyond a single conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem With Starting From Zero
&lt;/h2&gt;

&lt;p&gt;Imagine a production system suddenly starts reporting:&lt;br&gt;
PostgreSQL connection pool exhausted&lt;br&gt;
sqlalchemy.exc.TimeoutError&lt;br&gt;
.&lt;br&gt;
.&lt;br&gt;
QueuePool limit exceeded&lt;br&gt;
An LLM can explain what a connection pool is. It can suggest increasing the pool size. It can recommend checking for connection leaks.&lt;/p&gt;

&lt;p&gt;But that isn't necessarily what I want from an incident-response system.&lt;/p&gt;

&lt;h3&gt;
  
  
  what I really want to know is:
&lt;/h3&gt;

&lt;p&gt;Have we seen this before, and what happened when we fixed it?&lt;/p&gt;

&lt;p&gt;A previous incident might reveal that the real problem wasn't the pool size at all. Perhaps an unindexed query caused connections to remain open under load, and the eventual fix was to roll back a deployment and correct the query.&lt;/p&gt;

&lt;p&gt;That historical context is much more valuable than a generic explanation.&lt;/p&gt;

&lt;p&gt;This is where I use Hindsight.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Architecture Is Deliberately Simple
&lt;/h3&gt;

&lt;p&gt;The core system has three stages:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                     │Incident
                     ▼
             ┌───────────────┐
             │    RECALL     │
             │   Hindsight   │
             └───────┬───────┘
                     │
              Past incidents
                     │
                     ▼
             ┌───────────────┐
             │    Groq LLM   │
             │   Reasoning   │
             └───────┬───────┘
                     │
              Diagnosis +
              remediation
                     │
                     ▼
             ┌───────────────┐
             │    RETAIN     │
             │   Hindsight   │
             └───────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The application is written in Python and exposed through Streamlit. Groq provides the LLM interface, while Hindsight provides persistent memory.&lt;/p&gt;

&lt;p&gt;The important architectural decision is that memory isn't treated as a static document that gets pasted into every prompt.&lt;/p&gt;

&lt;p&gt;The agent explicitly recalls relevant information for the current incident.&lt;/p&gt;

&lt;p&gt;Then, after producing a resolution, it retains the new experience.&lt;/p&gt;

&lt;p&gt;That gives the system a feedback loop:&lt;/p&gt;

&lt;p&gt;Incident&lt;br&gt;
   ↓&lt;br&gt;
Recall&lt;br&gt;
   ↓&lt;br&gt;
Reason&lt;br&gt;
   ↓&lt;br&gt;
Resolve&lt;br&gt;
   ↓&lt;br&gt;
Retain&lt;br&gt;
   ↓&lt;br&gt;
Future Incident&lt;br&gt;
   ↓&lt;br&gt;
Recall the previous experience&lt;/p&gt;

&lt;p&gt;The more incidents the system processes, the more useful its historical context can become.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftckho5myqspso70eawpf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftckho5myqspso70eawpf.png" alt="Giving the Errors occured Question or Query" width="799" height="399"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For background on the approach, I used the Hindsight documentation and the broader agent memory explanation from Vectorize.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Recall Step
&lt;/h2&gt;

&lt;p&gt;The first interesting piece of the implementation is the memory lookup.&lt;br&gt;
My Streamlit application sends the incident to Hindsight:&lt;br&gt;
def recall_memory(query):&lt;br&gt;
    headers = {&lt;br&gt;
        "Authorization": f"Bearer {HINDSIGHT_API_KEY}"&lt;br&gt;
    }&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;res = requests.post(
    f"{HINDSIGHT_BASE_URL}/banks/{HINDSIGHT_BANK_ID}/recall",
    json={"query": query},
    headers=headers,
    timeout=5
)

if res.status_code == 200:
    memories = res.json().get("memories", [])
    return "\n".join(
        [m.get("text", "") for m in memories]
    )

return ""
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;There is an important detail here.&lt;br&gt;
I'm not retrieving the entire incident history and dumping it into the prompt.&lt;br&gt;
The current incident becomes the query:&lt;/p&gt;

&lt;p&gt;PostgreSQL connection pool exhausted&lt;/p&gt;

&lt;p&gt;Hindsight returns relevant memories, which are then passed to the reasoning stage.&lt;/p&gt;

&lt;p&gt;This changes the LLM's job.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;p&gt;"What could possibly be causing this?"&lt;/p&gt;

&lt;p&gt;the agent can reason more like:&lt;/p&gt;

&lt;p&gt;"This looks similar to incidents we've already encountered.&lt;br&gt;
What did we learn from those incidents?"&lt;/p&gt;

&lt;p&gt;That distinction is the central idea behind the project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory Becomes Part of the Prompt
&lt;/h2&gt;

&lt;p&gt;Once Hindsight returns relevant memories, I construct the system instruction:&lt;/p&gt;

&lt;p&gt;system_instruction = (&lt;br&gt;
    "You are an expert DevOps Incident Response Agent "&lt;br&gt;
    "powered by Hindsight persistent memory.\n"&lt;br&gt;
    "Analyze system errors and recommend fixes based on "&lt;br&gt;
    "past incident post-mortems.\n"&lt;br&gt;
    f"Relevant Past Memory Context:\n{context}\n"&lt;br&gt;
)&lt;br&gt;
The LLM isn't replacing the memory system.&lt;br&gt;
It's sitting on top of it.&lt;br&gt;
The responsibilities are separated:&lt;/p&gt;

&lt;p&gt;Hindsight&lt;br&gt;
    ↓&lt;br&gt;
"What have we seen before?"&lt;/p&gt;

&lt;p&gt;LLM&lt;br&gt;
    ↓&lt;br&gt;
"Given that history, what should we do now?"&lt;br&gt;
That separation made the architecture easier to reason about.&lt;/p&gt;

&lt;p&gt;The memory layer doesn't need to be responsible for generating an answer, and the LLM doesn't need to pretend that it remembers everything itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Agent Then Retains What It Learned
&lt;/h2&gt;

&lt;p&gt;After the LLM generates its response, the incident and solution are stored:&lt;/p&gt;

&lt;p&gt;retain_memory(&lt;br&gt;
    f"User Incident: {user_input} | "&lt;br&gt;
    f"Agent Solution: {reply}"&lt;br&gt;
)&lt;br&gt;
The retention function sends that information back to Hindsight:&lt;br&gt;
def retain_memory(text):&lt;br&gt;
    requests.post(&lt;br&gt;
        f"{HINDSIGHT_BASE_URL}/banks/"&lt;br&gt;
        f"{HINDSIGHT_BANK_ID}/retain",&lt;br&gt;
        json={"content": text},&lt;br&gt;
        headers=headers,&lt;br&gt;
        timeout=5&lt;br&gt;
    )&lt;br&gt;
This is the part that turns the application from a conventional chatbot into a system with persistent incident memory.&lt;/p&gt;

&lt;p&gt;Suppose the first incident produces:&lt;br&gt;
Incident:&lt;br&gt;
PostgreSQL connection pool exhausted.&lt;/p&gt;

&lt;p&gt;Resolution:&lt;br&gt;
Rollback v2.4.1 and investigate the query causing&lt;br&gt;
connections to remain open.&lt;br&gt;
That experience becomes available to the next relevant incident.&lt;br&gt;
The second time something similar happens, the agent isn't starting from an empty context.&lt;br&gt;
It can retrieve the previous resolution and use it as evidence during its reasoning.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Concrete Incident Flow
&lt;/h2&gt;

&lt;p&gt;Consider a production checkout service.&lt;br&gt;
The live logs show:&lt;br&gt;
[19:35:01 STDOUT] Gateway routing healthcheck OK.&lt;br&gt;
[19:35:05 STDOUT] /api/v1/orders 200 OK - 42ms&lt;/p&gt;

&lt;p&gt;[19:35:12 FATAL ERROR 500]&lt;br&gt;
PostgreSQL connection pool exhausted at /checkout endpoint.&lt;/p&gt;

&lt;p&gt;sqlalchemy.exc.TimeoutError:&lt;br&gt;
QueuePool limit exceeded&lt;br&gt;
The incident is passed into the agent.&lt;/p&gt;

&lt;p&gt;Hindsight retrieves a previous incident involving the database connection pool.&lt;/p&gt;

&lt;p&gt;The historical memory contains something like:&lt;br&gt;
 Incident #204:&lt;br&gt;
Connection Leak in DB Proxy&lt;/p&gt;

&lt;p&gt;Root Cause:&lt;br&gt;
An unindexed JOIN query in release v2.4.1&lt;br&gt;
held open DB connections under spike load.&lt;/p&gt;

&lt;p&gt;Previous Resolution:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Scale DB connections dynamically&lt;/li&gt;
&lt;li&gt;Roll back deployment to v2.4.0
The LLM now has both pieces of information:
CURRENT INCIDENT
   +
HISTORICAL EXPERIENCE
   ↓
LLM REASONING
   ↓
RECOMMENDED MITIGATION&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is much closer to how an experienced engineer approaches an incident.&lt;/p&gt;

&lt;p&gt;The engineer doesn't just look at the current stack trace.&lt;/p&gt;

&lt;p&gt;They ask:&lt;/p&gt;

&lt;p&gt;"Didn't we have this problem two weeks ago?"&lt;/p&gt;

&lt;p&gt;The agent is trying to make that question automatic.&lt;/p&gt;

&lt;p&gt;From Chatbot to Incident-Response Interface&lt;br&gt;
The repository also contains a dashboard concept called OpsMind AI.&lt;/p&gt;

&lt;p&gt;The interface is organized around three areas:&lt;br&gt;
┌──────────────────┬──────────────────┬──────────────────┐&lt;br&gt;
│                  │                  │                  │&lt;br&gt;
│  LIVE ERROR LOG  │ HINDSIGHT MEMORY │    RUNBOOK       │&lt;br&gt;
│                  │                  │                  │&lt;br&gt;
│ Current failure  │ Previous cases   │ Recommended      │&lt;br&gt;
│ Stack traces     │ Root cause       │ mitigation       │&lt;br&gt;
│ Service state    │ Resolution       │ execution        │&lt;br&gt;
│                  │                  │                  │&lt;br&gt;
└──────────────────┴──────────────────┴──────────────────┘&lt;br&gt;
The first panel represents the operational reality: errors arriving from a system.&lt;br&gt;
The second represents the agent's memory.&lt;br&gt;
The third represents action.&lt;br&gt;
That separation matters because incident response isn't finished when an LLM writes a paragraph explaining the error.&lt;br&gt;
The useful output is something an engineer can act on.&lt;/p&gt;

&lt;h3&gt;
  
  
  For example:
&lt;/h3&gt;

&lt;p&gt;Suggested Mitigation&lt;/p&gt;

&lt;p&gt;☑ Identify faulty query parameters&lt;br&gt;
☐ Increase connection pool limit&lt;br&gt;
☐ Roll back to stable release v2.4.0&lt;br&gt;
The eventual production implementation can connect those recommendations to controlled runbooks rather than merely displaying them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Persistent Memory Changes the Design
&lt;/h2&gt;

&lt;p&gt;Without persistent memory, an agent generally has two choices.&lt;br&gt;
It can rely on the current conversation context, or the application can build increasingly large prompts containing historical information.&lt;br&gt;
Neither approach is particularly attractive for an incident system with a growing history.&lt;br&gt;
A dedicated memory layer gives the application a different model:&lt;br&gt;
┌─────────────────┐&lt;br&gt;
                │ Current Context │&lt;br&gt;
                └────────┬────────┘&lt;br&gt;
                         │&lt;br&gt;
                         ▼&lt;br&gt;
                ┌─────────────────┐&lt;br&gt;
                │ Hindsight       │&lt;br&gt;
                │ Persistent      │&lt;br&gt;
                │ Memory          │&lt;br&gt;
                └────────┬────────┘&lt;br&gt;
                         │&lt;br&gt;
                         ▼&lt;br&gt;
                ┌─────────────────┐&lt;br&gt;
                │ LLM Reasoning   │&lt;br&gt;
                └─────────────────┘&lt;br&gt;
The agent can selectively retrieve what matters instead of treating every historical incident as equally relevant.&lt;br&gt;
That is the reason Hindsight is central to the architecture rather than just another dependency.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Learned Building It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Memory is more interesting when it changes future behavior
&lt;/h3&gt;

&lt;p&gt;Saving conversations isn't enough.&lt;/p&gt;

&lt;p&gt;The useful question is:&lt;br&gt;
Does something the agent learned today improve what it can do tomorrow?&lt;br&gt;
That became the design criterion for the memory layer.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Retrieval should happen before reasoning
&lt;/h3&gt;

&lt;p&gt;I found it more useful to make historical context available before asking the LLM to analyze the incident.&lt;br&gt;
The sequence matters:&lt;br&gt;
Recall → Reason → Retain&lt;br&gt;
rather than:&lt;br&gt;
Reason → Hope the model remembers&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Incident resolution needs an action layer
&lt;/h3&gt;

&lt;p&gt;An explanation is useful, but production engineering ultimately requires action.&lt;br&gt;
That's why the dashboard concept includes a runbook layer.&lt;br&gt;
The long-term goal is not:&lt;br&gt;
AI:&lt;br&gt;
"Here is what you should probably do."&lt;br&gt;
It is:&lt;br&gt;
AI:&lt;br&gt;
"Here is the evidence.&lt;br&gt;
Here is the historical incident.&lt;br&gt;
Here is the proposed runbook.&lt;br&gt;
Here is what will happen if you execute it."&lt;br&gt;
with appropriate human approval and safeguards around execution.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Persistent memory introduces its own engineering problems
&lt;/h3&gt;

&lt;p&gt;Once an agent can remember previous incidents, memory quality becomes part of system quality.&lt;br&gt;
Bad historical information can lead to bad future recommendations.&lt;br&gt;
That means productionizing this architecture requires thinking about retention quality, stale resolutions, conflicting incidents, confidence, observability, and human review.&lt;br&gt;
Memory is powerful precisely because it has consequences.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. The most useful agent isn't necessarily the one with the largest prompt
&lt;/h3&gt;

&lt;p&gt;The project reinforced a simple idea for me:&lt;br&gt;
More context isn't automatically better context.&lt;br&gt;
The useful context is the context that is relevant to the incident being investigated.&lt;br&gt;
That is why the recall layer is such an important part of the system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Goes Next
&lt;/h2&gt;

&lt;p&gt;The current architecture gives me a foundation for a more complete incident-response system.&lt;br&gt;
The next stage is connecting the incident stream, memory retrieval, reasoning, and runbook execution into one controlled workflow:&lt;/p&gt;

&lt;p&gt;Production Logs&lt;br&gt;
      │&lt;br&gt;
      ▼&lt;br&gt;
Incident Detection&lt;br&gt;
      │&lt;br&gt;
      ▼&lt;br&gt;
Hindsight Recall&lt;br&gt;
      │&lt;br&gt;
      ▼&lt;br&gt;
LLM Analysis&lt;br&gt;
      │&lt;br&gt;
      ├───────────────┐&lt;br&gt;
      ▼               ▼&lt;br&gt;
Root Cause       Historical Evidence&lt;br&gt;
      │&lt;br&gt;
      ▼&lt;br&gt;
Runbook Proposal&lt;br&gt;
      │&lt;br&gt;
      ▼&lt;br&gt;
Human Approval&lt;br&gt;
      │&lt;br&gt;
      ▼&lt;br&gt;
Controlled Execution&lt;br&gt;
      │&lt;br&gt;
      ▼&lt;br&gt;
Incident Result&lt;br&gt;
      │&lt;br&gt;
      ▼&lt;br&gt;
Hindsight Retain&lt;/p&gt;

&lt;p&gt;The final step is the one I care about most.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fchrwsltubdoi5k97a8iu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fchrwsltubdoi5k97a8iu.png" alt="Working of DevOps" width="799" height="410"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every resolved incident becomes potential knowledge for the next one.&lt;/p&gt;

&lt;p&gt;That creates a system where incident response doesn't have to start from zero every time.&lt;/p&gt;

&lt;p&gt;The goal isn't to build an AI that magically knows how to operate production systems.&lt;/p&gt;

&lt;p&gt;It's to build an agent that can remember what happened before, reason from that history, and turn previous operational&lt;br&gt;
experience into useful context when the next failure arrives.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/Saipriya41/hindsight-agent.git" rel="noopener noreferrer"&gt;GITHUB-hindsight&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://hindsight.vectorize.io" rel="noopener noreferrer"&gt;Hindsight documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/vectorize-io/hindsight#quick-start" rel="noopener noreferrer"&gt;Hindsight API quickstart&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://hindsight.vectorize.io" rel="noopener noreferrer"&gt;Agent memory overview&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>automation</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
