<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: poojitha chougani</title>
    <description>The latest articles on DEV Community by poojitha chougani (@poojitha_chougani_81831b3).</description>
    <link>https://dev.to/poojitha_chougani_81831b3</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147306%2F274e2a58-4769-4a97-8734-6a6edc87c2b6.png</url>
      <title>DEV Community: poojitha chougani</title>
      <link>https://dev.to/poojitha_chougani_81831b3</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/poojitha_chougani_81831b3"/>
    <language>en</language>
    <item>
      <title>What if your deployment agent could remember every failure it ever fixed</title>
      <dc:creator>poojitha chougani</dc:creator>
      <pubDate>Tue, 29 Sep 2026 03:13:54 +0000</pubDate>
      <link>https://dev.to/poojitha_chougani_81831b3/what-if-your-deployment-agent-could-remember-every-failure-it-ever-fixed-51lc</link>
      <guid>https://dev.to/poojitha_chougani_81831b3/what-if-your-deployment-agent-could-remember-every-failure-it-ever-fixed-51lc</guid>
      <description>&lt;p&gt;&lt;a href="https://pipelinesage.streamlit.app/" rel="noopener noreferrer"&gt;🚀 Try PipelineSage Live&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Deployment #1057 of our payment service failed with a database migration timeout after 30 seconds. Twenty days earlier, deployment #1017 had failed the same way, someone had fixed it, and the fix was written down.&lt;/p&gt;

&lt;p&gt;Nobody remembered it.&lt;/p&gt;

&lt;p&gt;I built &lt;strong&gt;PipelineSage&lt;/strong&gt; to close that gap. It reads a failed deployment, retrieves relevant past incidents using &lt;strong&gt;Hindsight&lt;/strong&gt;, and recommends a fix grounded in that history. When a human confirms that the fix worked, the outcome is written back to Hindsight so the next failure can learn from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the system does
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0ae8z9uncd6weazvf91y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0ae8z9uncd6weazvf91y.png" alt=" " width="799" height="354"&gt;&lt;/a&gt;&lt;br&gt;
The flow is simple:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A CI/CD failure arrives with its service, environment, commit, and error.&lt;/li&gt;
&lt;li&gt;PipelineSage queries &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; for relevant past incidents.&lt;/li&gt;
&lt;li&gt;Retrieved memories are filtered, deduplicated, and ranked.&lt;/li&gt;
&lt;li&gt;The current failure and relevant memories are sent to an LLM (&lt;code&gt;openai/gpt-oss-120b&lt;/code&gt; on Groq).&lt;/li&gt;
&lt;li&gt;The agent produces historical evidence, diagnosis, and a recommended fix.&lt;/li&gt;
&lt;li&gt;A human reviews the recommendation.&lt;/li&gt;
&lt;li&gt;If the fix works, the confirmed outcome is written back to Hindsight.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The project is split into &lt;code&gt;agent/&lt;/code&gt;, &lt;code&gt;memory/&lt;/code&gt;, &lt;code&gt;services/&lt;/code&gt;, and a Streamlit interface.&lt;/p&gt;

&lt;p&gt;I deliberately didn't build my own memory layer. The &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight documentation&lt;/a&gt; already covers the difficult parts of storing and retrieving agent memories. My integration is a small wrapper around it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9hczfb7tt52effn7ckxe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9hczfb7tt52effn7ckxe.png" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Memory is a source of truth
&lt;/h2&gt;

&lt;p&gt;Memory introduces a problem that stateless agents don't have: the agent can retrieve the wrong precedent and confidently present it as history.&lt;/p&gt;

&lt;p&gt;So I designed PipelineSage around one rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Whatever the agent says about history must be traceable to something Hindsight actually returned.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  Writing memories for a cold reader
&lt;/h3&gt;

&lt;p&gt;Every incident is retained with fixed fields:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;retain_incident&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
DevOps pipeline incident.

Deployment: #&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;deployment_id&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Service: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;service&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Environment: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;environment&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Status: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Failure:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;error&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Root cause:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;root_cause&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Not yet confirmed.&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Resolution:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;resolution&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Not yet resolved.&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Outcome:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;outcome&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;No outcome recorded.&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Related historical incident:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;related_historical_incident&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Not specified.&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The defaults are intentional. An unresolved incident is stored as &lt;code&gt;Not yet resolved.&lt;/code&gt; rather than left blank.&lt;/p&gt;

&lt;p&gt;I also store a &lt;code&gt;Related historical incident&lt;/code&gt;. This creates a simple trail showing which earlier incident influenced a later resolution.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recall returns candidates, not answers
&lt;/h2&gt;

&lt;p&gt;Retrieval uses Hindsight's &lt;code&gt;recall&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4096&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;budget&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I run multiple queries using the incident's service and error text, then merge and clean the results.&lt;/p&gt;

&lt;p&gt;One important safeguard is excluding the current deployment from its own history:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Do not use the current deployment as history.
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;deployment_id&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deployment #&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;deployment_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;
    &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deployment: #&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;deployment_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;continue&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I also had to deal with near-duplicate memories. Five retrieved memories do not necessarily mean five independent pieces of evidence. Several can be different versions of the same incident.&lt;/p&gt;

&lt;p&gt;My first ranking approach also had keyword bonuses, including a bonus for &lt;code&gt;"500"&lt;/code&gt; because I already knew the answer I wanted. That was effectively an answer key, not a retrieval test.&lt;/p&gt;

&lt;p&gt;I replaced it with general signals such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Same service&lt;/li&gt;
&lt;li&gt;Successful documented outcome&lt;/li&gt;
&lt;li&gt;Similar failure text&lt;/li&gt;
&lt;li&gt;Penalties for unrelated services or failure classes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The retrieval should work for failures the system has never seen before.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prompt rules are not guarantees
&lt;/h2&gt;

&lt;p&gt;The system prompt tells the model that Hindsight memories are the source of truth:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Never invent historical deployments, fixes, outcomes, numbers,
batch sizes, timeout values, configuration values, or
infrastructure changes.

If a historical successful resolution contains an exact value
such as "500 records", preserve that value exactly.

If historical evidence is insufficient, clearly state that.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Temperature is set to &lt;code&gt;0.1&lt;/code&gt;, and the output is separated into historical evidence, diagnosis, and recommended fix.&lt;/p&gt;

&lt;p&gt;But prompts are not enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened on #1057
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F21oz8bvd2gsm556n60cl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F21oz8bvd2gsm556n60cl.png" alt=" " width="800" height="1793"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For deployment #1057, the agent retrieved memories pointing to #1017.&lt;/p&gt;

&lt;p&gt;The historical incident described a payment-service migration that exceeded the 30-second limit and succeeded after being split into batches of &lt;strong&gt;500 records&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;PipelineSage recommended:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Process the rows in batches of 500&lt;/li&gt;
&lt;li&gt;Test in staging&lt;/li&gt;
&lt;li&gt;Re-run the migration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important part was that the &lt;strong&gt;500-record value came from the retrieved memory&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;But the model also added that the batch size had been &lt;em&gt;“proven to keep each migration step within the 30-second timeout.”&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That claim was not present in the memory.&lt;/p&gt;

&lt;p&gt;This exposed an important limitation: a prompt can tell the model not to invent facts, but it cannot guarantee compliance.&lt;/p&gt;

&lt;p&gt;The next improvement I want is a mechanical verification step that flags numeric or timing claims in the recommendation when they cannot be found in the retrieved memories.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing the memory loop
&lt;/h2&gt;

&lt;p&gt;When the recommended fix works, a human confirms it.&lt;/p&gt;

&lt;p&gt;PipelineSage then writes the outcome back to Hindsight with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Status: &lt;code&gt;SUCCESS&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Resolution: &lt;code&gt;Process historical transaction rows in batches of 500 records&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Outcome: Human-confirmed success&lt;/li&gt;
&lt;li&gt;Related historical incident: &lt;code&gt;#1017&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This human gate is important.&lt;/p&gt;

&lt;p&gt;If the agent automatically wrote its own recommendation into memory, an unverified guess could become a future "documented precedent." Human confirmation keeps the memory based on what actually happened rather than what the model predicted.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Write memories for a cold reader.&lt;/strong&gt; Fixed fields and explicit unresolved states make historical incidents easier to interpret later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Treat recall as candidate generation.&lt;/strong&gt; Exclude the current event, expect duplicates, and don't mistake multiple memories for independent evidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't let tests contain the answer.&lt;/strong&gt; If your query already names the fix you're trying to retrieve, you're not really testing retrieval.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompts need mechanical backstops.&lt;/strong&gt; The model can still make unsupported claims even when the prompt explicitly prohibits them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gate memory writes on human-verified outcomes.&lt;/strong&gt; Reading from memory is one thing; writing new facts into memory is where mistakes can compound.&lt;/p&gt;

&lt;p&gt;I started with memory because the same failures keep costing the same hour.&lt;/p&gt;

&lt;p&gt;What surprised me is how little of the work was about storage, and how much was about deciding &lt;strong&gt;what the agent is allowed to believe.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Learn More
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight on GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;What is Agent Memory? — Vectorize&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://pipelinesage.streamlit.app/" rel="noopener noreferrer"&gt;Try PipelineSage Live&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>devops</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
