<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Harshini Sai Makkapati</title>
    <description>The latest articles on DEV Community by Harshini Sai Makkapati (@harshinisai).</description>
    <link>https://dev.to/harshinisai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149973%2Faac0aca0-9867-492e-873c-a7977a787db1.png</url>
      <title>DEV Community: Harshini Sai Makkapati</title>
      <link>https://dev.to/harshinisai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/harshinisai"/>
    <language>en</language>
    <item>
      <title>How I Built an Incident Response Agent That Remembers What Killed Your Service Last Time</title>
      <dc:creator>Harshini Sai Makkapati</dc:creator>
      <pubDate>Tue, 29 Sep 2026 14:49:39 +0000</pubDate>
      <link>https://dev.to/harshinisai/how-i-built-an-incident-response-agent-that-remembers-what-killed-your-service-last-time-313j</link>
      <guid>https://dev.to/harshinisai/how-i-built-an-incident-response-agent-that-remembers-what-killed-your-service-last-time-313j</guid>
      <description>&lt;p&gt;Every on-call engineer has lived this: it's 2am, payment-api is down, and you're staring at a dashboard trying to remember — did we see this before? What did we try? What actually worked?&lt;/p&gt;

&lt;p&gt;The answer is almost always yes, you've seen it before. But that knowledge lives in a Slack thread from three months ago, a runbook nobody updated, or the head of the engineer who's now on vacation. So you start from scratch. You try restarting Redis. It doesn't work. You waste 20 minutes. You finally increase the database connection pool. Latency drops immediately. You write a note to yourself and forget about it.&lt;/p&gt;

&lt;p&gt;I built IncidentRecall to break that loop. It's an incident response agent that uses &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight agent memory&lt;/a&gt; to recall what happened in similar past incidents, tell you what worked, and — critically — warn you what failed so you don't waste time on it again.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the System Does
&lt;/h2&gt;

&lt;p&gt;The core flow is simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident → Recall similar memories → LLM recommendation → Engineer feedback → New memory → Future incidents learn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An engineer submits an incident — service name, description, symptoms, severity. The system queries &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; for semantically similar past incidents. Those memories become the evidence base for a structured LLM recommendation. The recommendation includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Matched historical incidents&lt;/strong&gt; — which past incidents are similar and why&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What to do&lt;/strong&gt; — grounded in what actually worked before, with a confidence count&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What NOT to try first&lt;/strong&gt; — actions that failed in similar situations, with evidence IDs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When the engineer resolves the incident, they submit the outcome — worked or failed. That outcome gets written back to Hindsight. The next similar incident retrieves it.&lt;/p&gt;

&lt;p&gt;The loop closes. The agent gets smarter with every incident.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2m8k6y35qnzlebfm72ya.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2m8k6y35qnzlebfm72ya.jpeg" alt="Incident form" width="800" height="429"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The incident form. Symptom tags are clickable — select all that apply. The agent uses these to build the Hindsight recall query.&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  The Memory Layer
&lt;/h2&gt;

&lt;p&gt;The most important architectural decision was keeping the memory layer thin and explicit. The rest of the application — the API, the recommendation engine — only ever calls two functions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# memory.py
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Retrieve relevant memories for a query.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;_hindsight&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;budget&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;low&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;incident resolution&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Store a memory.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;_hindsight&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the entire interface. &lt;code&gt;recall()&lt;/code&gt; and &lt;code&gt;retain()&lt;/code&gt;. Nothing else leaks through.&lt;/p&gt;

&lt;p&gt;This matters because &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;agent memory&lt;/a&gt; is not a database. You don't query it with SQL. You query it with natural language, and what comes back is semantically relevant text. The quality of what you get out depends entirely on the quality of what you put in.&lt;/p&gt;

&lt;p&gt;The recall query is built to be natural, not keyword-stuffed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Service: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Problem: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Symptoms: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;symptoms&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Find similar incidents, what fixes worked, what failed, root causes.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the memory written back after feedback is structured like a real incident report — not a JSON blob, not a summary, but a narrative with explicit &lt;code&gt;[WORKED]&lt;/code&gt; and &lt;code&gt;[FAILED]&lt;/code&gt; labels that the LLM can parse unambiguously:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;memory_content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Incident ID: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;incident_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Service: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;service&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Symptoms: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;symptoms&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Root Cause: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;root_cause&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Not specified&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Actions Attempted:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  - [&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;outcome_label&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;] &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;    Notes: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;notes&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;None&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Successful Fix&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; + req.action if req.result == &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="n"&gt;worked&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; else &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="n"&gt;Failed&lt;/span&gt; &lt;span class="nc"&gt;Actions &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;do&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="n"&gt;first&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; + req.action&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;[WORKED]&lt;/code&gt; / &lt;code&gt;[FAILED]&lt;/code&gt; labels are load-bearing. The recommendation engine parses them to build the &lt;code&gt;do_not_try_first&lt;/code&gt; list. If you write vague memories, you get vague recommendations.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fco4lk28bm90c13yw5lba.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fco4lk28bm90c13yw5lba.jpeg" alt="Memory Inspector" width="800" height="429"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The Memory Inspector shows exactly which memories Hindsight retrieved and their full content. INC-031D36 is a memory written during this session — the learning loop working in real time.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flyopexjbmssf6oa1pui8.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flyopexjbmssf6oa1pui8.jpeg" alt="Memory Inspector full" width="800" height="429"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The raw memory text written to Hindsight after engineer feedback. The structured narrative format — with explicit &lt;code&gt;[WORKED]&lt;/code&gt; labels — is what makes future recall useful.&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  The Recommendation Engine
&lt;/h2&gt;

&lt;p&gt;The LLM's job is not to be creative. It's to be a structured parser of evidence.&lt;/p&gt;

&lt;p&gt;The prompt is explicit about this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;HISTORICAL MEMORIES FROM HINDSIGHT (your ONLY evidence source):
These are real past incidents retrieved from the memory bank. Use them as your evidence.
Do NOT invent incidents. Do NOT use general knowledge to claim something worked or failed.
If the memories do not contain relevant evidence, say so.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The LLM returns a structured JSON object with matches, a recommendation, a &lt;code&gt;do_not_try_first&lt;/code&gt; list, and a confidence count. The confidence count is not a probability — it's a raw count: "this fix worked in 3 out of 3 similar incidents." That's more honest and more useful than a percentage.&lt;/p&gt;

&lt;p&gt;The recommendation engine also runs a &lt;strong&gt;baseline&lt;/strong&gt; — the same incident, but with no memories passed in. This gives you a side-by-side comparison: what would a generic SRE recommendation look like versus what the memory-grounded agent recommends? The difference is the value of memory made visible.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flsk5kk9q5n74ybrnvjdq.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flsk5kk9q5n74ybrnvjdq.jpeg" alt="Recommendation panel" width="800" height="429"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The recommendation panel. Green = what to do, grounded in historical evidence. Red = what NOT to try first, with the exact incident IDs where it failed.&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  What the Learning Loop Actually Looks Like
&lt;/h2&gt;

&lt;p&gt;Here's a concrete example. payment-api goes down with high latency and database connection exhaustion. Without memory, the agent says: check recent deployments, review logs, follow standard runbook. Useful, but generic.&lt;/p&gt;

&lt;p&gt;With memory, it says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Based on 4 similar historical incidents (INC-1002, INC-1005, INC-1008): Increase database connection pool from 50 to 150. This fix worked in 3 of 4 similar incidents retrieved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Avoid First:&lt;/strong&gt; Restart Redis cache — failed in INC-1002, INC-1005, INC-1008. Restart failed in 3/3 similar incidents.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's the difference. The agent knows that restarting Redis is the instinctive wrong move for this service, because it's been tried and failed three times. It saves you 20 minutes at 2am.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F00wqmf0lxhjtwqv4vcgx.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F00wqmf0lxhjtwqv4vcgx.jpeg" alt="Historical matches" width="800" height="429"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;4 historical matches retrieved from Hindsight. Each card shows root cause, what worked (green), what failed (red), and the deployment change that triggered it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;After the engineer resolves the incident, they submit the outcome via the feedback form:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcw2dnw2rlff7wsei0do4.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcw2dnw2rlff7wsei0do4.jpeg" alt="Feedback form" width="800" height="429"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The feedback form. The action field pre-fills from the recommendation. Both "Fix Worked" and "Fix Failed" outcomes are stored — failed outcomes are what power the Avoid First warnings.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# POST /feedback
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;incident_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;INC-A1B2C3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Increased database connection pool from 50 to 200&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;worked&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Latency returned to normal within 2 minutes.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;root_cause&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Connection pool undersized for current traffic&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The next similar incident retrieves this memory. The confidence count goes up. The pattern strengthens.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxjdfksaugjnj8627oqhx.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxjdfksaugjnj8627oqhx.jpeg" alt="Memory updated" width="800" height="429"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The Memory Updated confirmation. This outcome is now in Hindsight. The next similar incident will retrieve it.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Hindsight
&lt;/h2&gt;

&lt;p&gt;I looked at a few options for the memory layer. The reason I chose &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; was the &lt;code&gt;retain&lt;/code&gt; / &lt;code&gt;recall&lt;/code&gt; API design. It's opinionated in the right way — you write natural language in, you get semantically relevant natural language back. There's no schema to maintain, no embedding pipeline to manage, no vector database to operate.&lt;/p&gt;

&lt;p&gt;For an incident response use case, that matters. Incidents are described in natural language. The fixes are described in natural language. The recall query is natural language. Forcing all of that through a structured schema would lose information. Hindsight keeps the full fidelity of the text and handles the semantic retrieval.&lt;/p&gt;

&lt;p&gt;The other thing that mattered: both &lt;code&gt;retain()&lt;/code&gt; and &lt;code&gt;recall()&lt;/code&gt; are single function calls. The integration is genuinely thin. The memory layer in this project is about 50 lines of real code. That's the right size for infrastructure that should be invisible.&lt;/p&gt;




&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Write memories like incident reports, not like database records.&lt;/strong&gt;&lt;br&gt;
The quality of recall depends on the quality of what you retain. Structured narrative with explicit outcome labels (&lt;code&gt;[WORKED]&lt;/code&gt;, &lt;code&gt;[FAILED]&lt;/code&gt;) gives the LLM unambiguous signal. Vague summaries produce vague recommendations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Store failed outcomes, not just successful ones.&lt;/strong&gt;&lt;br&gt;
The &lt;code&gt;do_not_try_first&lt;/code&gt; list is only possible because failed actions are stored. Most systems only record what worked. That's half the information. Knowing what failed — and in which specific incidents — is often more valuable than knowing what worked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The baseline comparison is the best demo you can give.&lt;/strong&gt;&lt;br&gt;
Showing the memory-grounded recommendation alongside the no-memory baseline makes the value of agent memory immediately obvious. Without it, you're asking people to imagine the difference. With it, they can see it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Keep the memory interface minimal.&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;recall()&lt;/code&gt; and &lt;code&gt;retain()&lt;/code&gt;. That's it. The rest of the application doesn't need to know how memory works. This made it trivial to swap between a live Hindsight backend and a deterministic mock for development — same interface, different implementation behind a single environment variable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Confidence counts beat confidence percentages.&lt;/strong&gt;&lt;br&gt;
"Worked in 3/3 similar incidents" is more credible than "87% confidence." Engineers are skeptical. Show them the evidence count, not a number that sounds like it came from nowhere.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;The current system stores memories per-session in the demo layer and per-bank in live mode. The natural next step is per-service memory banks — so cart-service incidents don't pollute payment-api recall results, and each service builds its own institutional knowledge over time.&lt;/p&gt;

&lt;p&gt;The other obvious extension is automatic pattern detection: if the same action fails three times for the same service, surface a proactive alert before the next incident even happens.&lt;/p&gt;

&lt;p&gt;The foundation for both is already there. The memory layer is the hard part. Once you have reliable recall and retain, the rest is just building on top of it.&lt;/p&gt;




&lt;p&gt;If you're building agents that need to learn from past interactions — not just retrieve static documents, but actually accumulate operational knowledge over time — &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; is worth looking at. The &lt;code&gt;retain&lt;/code&gt; / &lt;code&gt;recall&lt;/code&gt; model is the right abstraction for this class of problem.&lt;/p&gt;

&lt;p&gt;The full project is on GitHub. The backend is FastAPI + Groq + Hindsight. The frontend is React + Vite. Demo mode requires no API keys — you can run the full learning loop locally in under five minutes.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>python</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
