<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: keerthi</title>
    <description>The latest articles on DEV Community by keerthi (@24311a12f4).</description>
    <link>https://dev.to/24311a12f4</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150107%2F97e72049-61c0-417d-82c4-25f51efab6c3.png</url>
      <title>DEV Community: keerthi</title>
      <link>https://dev.to/24311a12f4</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/24311a12f4"/>
    <language>en</language>
    <item>
      <title>I built an incident-memory agent that remembers why previous fixes failed</title>
      <dc:creator>keerthi</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:41:58 +0000</pubDate>
      <link>https://dev.to/24311a12f4/i-built-an-incident-memory-agent-that-remembers-why-previous-fixes-failed-15lf</link>
      <guid>https://dev.to/24311a12f4/i-built-an-incident-memory-agent-that-remembers-why-previous-fixes-failed-15lf</guid>
      <description>&lt;h1&gt;
  
  
  Why Hindsight Needs the Reason a Fix Failed
&lt;/h1&gt;

&lt;p&gt;The most useful thing an incident agent can remember is not simply that someone restarted a service. It is that the restart reduced errors for 25 minutes, the failures returned, and the real problem was still unresolved.&lt;/p&gt;

&lt;p&gt;That idea became the foundation of an incident memory system I built using &lt;strong&gt;Hindsight&lt;/strong&gt; as the long-term memory layer. Instead of storing only successful resolutions, the system also remembers failed troubleshooting attempts, engineer corrections, and the reasoning behind each decision. The goal is simple: help future investigators learn from previous incidents without confusing historical evidence with current facts.&lt;/p&gt;

&lt;h2&gt;
  
  
  From Incident Report to Long-Term Memory
&lt;/h2&gt;

&lt;p&gt;An investigation usually begins with limited information:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Incident title&lt;/li&gt;
&lt;li&gt;Service name&lt;/li&gt;
&lt;li&gt;Severity&lt;/li&gt;
&lt;li&gt;Error message&lt;/li&gt;
&lt;li&gt;Symptoms&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This information is enough to start investigating, but it is not enough to create valuable long-term knowledge.&lt;/p&gt;

&lt;p&gt;Only after an engineer identifies the root cause and completes the investigation does the incident become useful for future incidents. At that point the memory stores:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Root cause&lt;/li&gt;
&lt;li&gt;Final resolution&lt;/li&gt;
&lt;li&gt;Failed attempts&lt;/li&gt;
&lt;li&gt;Why those attempts failed&lt;/li&gt;
&lt;li&gt;Lessons learned&lt;/li&gt;
&lt;li&gt;Engineer verification&lt;/li&gt;
&lt;li&gt;Corrections made later&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Rather than letting the investigation flow interact directly with the storage backend, I created a small abstraction layer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;MemoryBank&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;backend&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;hindsight&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;local-fallback&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;record&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;MemoryRecord&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nx"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;NewIncident&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;
  &lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;RecalledMemory&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;MemoryRecord&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nf"&gt;clear&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This keeps the investigation workflow independent from the storage implementation. The application decides &lt;strong&gt;what should be remembered&lt;/strong&gt;, while Hindsight decides &lt;strong&gt;how memories are retained and retrieved&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Failed Attempts Matter
&lt;/h2&gt;

&lt;p&gt;One design decision turned out to be surprisingly important.&lt;/p&gt;

&lt;p&gt;Instead of storing only the successful solution, every failed troubleshooting attempt is stored together with the reason it failed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;FailedAttempt&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;why_it_failed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This small structure changes how future incidents are investigated.&lt;/p&gt;

&lt;p&gt;Imagine a Payments API incident where the database connection pool becomes exhausted during peak traffic.&lt;/p&gt;

&lt;p&gt;The engineering team tries several actions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Attempt 1&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Restart the application pods.&lt;/p&gt;

&lt;p&gt;Result:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Errors disappear.&lt;/li&gt;
&lt;li&gt;Twenty-five minutes later they return.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Attempt 2&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Increase the database connection limit.&lt;/p&gt;

&lt;p&gt;Result:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Failures occur later.&lt;/li&gt;
&lt;li&gt;Database memory usage increases.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Eventually the engineers discover that a retry helper never releases database connections.&lt;/p&gt;

&lt;p&gt;The permanent fix is to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;release connections in a &lt;code&gt;finally&lt;/code&gt; block,&lt;/li&gt;
&lt;li&gt;limit the pool size,&lt;/li&gt;
&lt;li&gt;monitor pool utilization.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If memory stores only the final resolution, future engineers lose valuable diagnostic information.&lt;/p&gt;

&lt;p&gt;If memory stores only "Restarted pods", future engineers may waste valuable time repeating the same temporary workaround.&lt;/p&gt;

&lt;p&gt;The useful knowledge is actually:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Restarting the service temporarily reduced errors because leaked connections were cleared, but the leak continued to exist.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That explanation becomes reusable engineering knowledge.&lt;/p&gt;




&lt;h2&gt;
  
  
  Preserving Trust in Memory
&lt;/h2&gt;

&lt;p&gt;Not every memory should carry the same level of confidence.&lt;/p&gt;

&lt;p&gt;Some incidents are imported as sample data.&lt;/p&gt;

&lt;p&gt;Some are verified by engineers.&lt;/p&gt;

&lt;p&gt;Others are corrected later when new information becomes available.&lt;/p&gt;

&lt;p&gt;To preserve that history, each memory keeps additional metadata.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;MemoryRecord&lt;/span&gt; &lt;span class="kd"&gt;extends&lt;/span&gt; &lt;span class="nx"&gt;Incident&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;source&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;seed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;engineer_saved&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;engineer_correction&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;verified_at&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;corrections&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;[];&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes it possible to distinguish between:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;demonstration data,&lt;/li&gt;
&lt;li&gt;engineer-confirmed incidents,&lt;/li&gt;
&lt;li&gt;corrected historical records.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead of overwriting previous knowledge, corrections become part of the memory's history.&lt;/p&gt;




&lt;h2&gt;
  
  
  Using Historical Memory Carefully
&lt;/h2&gt;

&lt;p&gt;When a new incident arrives, the application searches memory using information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;service&lt;/li&gt;
&lt;li&gt;title&lt;/li&gt;
&lt;li&gt;symptoms&lt;/li&gt;
&lt;li&gt;error message&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Hindsight retrieves the most relevant previous incidents.&lt;/p&gt;

&lt;p&gt;Those memories are then presented as historical context—not as conclusions.&lt;/p&gt;

&lt;p&gt;The investigation prompt follows a few simple rules.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Keep current incident evidence separate from historical memories.&lt;/li&gt;
&lt;li&gt;Never assume a previous root cause is the current root cause.&lt;/li&gt;
&lt;li&gt;Clearly distinguish successful and failed previous actions.&lt;/li&gt;
&lt;li&gt;If multiple historical incidents disagree, request more current evidence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This prevents historical knowledge from becoming false certainty.&lt;/p&gt;

&lt;p&gt;For example, suppose a checkout service begins timing out again during heavy traffic.&lt;/p&gt;

&lt;p&gt;The system recalls a previous Payments API incident involving connection leaks.&lt;/p&gt;

&lt;p&gt;Instead of saying:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This is definitely another connection leak.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;the investigation says something closer to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A previous incident involving this service experienced a connection leak. Restarting the service only reduced errors temporarily. Check current connection usage, retry logic, and pool utilization before concluding the same root cause exists.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Historical experience guides the investigation without replacing current evidence.&lt;/p&gt;




&lt;h2&gt;
  
  
  Handling Conflicting Memories
&lt;/h2&gt;

&lt;p&gt;Real systems evolve.&lt;/p&gt;

&lt;p&gt;The same service might fail for completely different reasons six months apart.&lt;/p&gt;

&lt;p&gt;One historical incident might point to a database leak.&lt;/p&gt;

&lt;p&gt;Another might identify an upstream dependency failure.&lt;/p&gt;

&lt;p&gt;Rather than combining these into a single confident answer, the application surfaces both memories and explicitly asks for additional evidence.&lt;/p&gt;

&lt;p&gt;Conflicting memories are treated as useful information rather than mistakes.&lt;/p&gt;

&lt;p&gt;This encourages investigation instead of assumption.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Learned
&lt;/h2&gt;

&lt;p&gt;Building this project changed how I think about long-term memory for AI agents.&lt;/p&gt;

&lt;p&gt;The most valuable memory is often not the successful fix.&lt;/p&gt;

&lt;p&gt;It is understanding &lt;strong&gt;why previous fixes failed&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Negative results prevent engineers from repeating ineffective troubleshooting steps.&lt;/p&gt;

&lt;p&gt;Engineer verification increases confidence.&lt;/p&gt;

&lt;p&gt;Corrections preserve knowledge instead of hiding mistakes.&lt;/p&gt;

&lt;p&gt;Most importantly, historical memory should narrow the investigation—not replace it.&lt;/p&gt;

&lt;p&gt;Hindsight provides a powerful foundation for retaining and retrieving engineering experience, but the surrounding application still needs clear rules about what memories mean and how they should influence decision making.&lt;/p&gt;

&lt;p&gt;For incident response, memory should answer one question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What should I investigate next?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;—not—&lt;/p&gt;

&lt;p&gt;**"What already happened?"&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6942w2m1x1j7vltbhefd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6942w2m1x1j7vltbhefd.png" alt=" " width="800" height="92"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2jkkz1u2q9j6uxzis3he.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2jkkz1u2q9j6uxzis3he.png" alt=" " width="799" height="469"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3xctpvu02fvr6orbnuie.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3xctpvu02fvr6orbnuie.png" alt=" " width="329" height="646"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1pppzjaywwiwzzyrs631.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1pppzjaywwiwzzyrs631.png" alt=" " width="317" height="524"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa4cgsq2raz820h3ktdfi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa4cgsq2raz820h3ktdfi.png" alt=" " width="800" height="543"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>typescript</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Why Hindsight Needs the Reason a Fix Failed</title>
      <dc:creator>keerthi</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:22:59 +0000</pubDate>
      <link>https://dev.to/24311a12f4/why-hindsight-needs-the-reason-a-fix-failed-531h</link>
      <guid>https://dev.to/24311a12f4/why-hindsight-needs-the-reason-a-fix-failed-531h</guid>
      <description>&lt;h1&gt;
  
  
  Why Hindsight Needs the Reason a Fix Failed
&lt;/h1&gt;

&lt;p&gt;The most useful thing an incident agent can remember is not that someone restarted a service. It is that the restart bought 25 minutes, the errors returned, and the connection leak was still there.&lt;/p&gt;

&lt;p&gt;That distinction is the center of the incident memory system I built. It collects what engineers learned during an outage, retrieves relevant experience when another incident arrives, and gives an investigator that context without letting history masquerade as current evidence. I use Hindsight for the long-term memory layer because incident knowledge is more than a pile of similar-looking error messages.&lt;/p&gt;

&lt;h2&gt;
  
  
  From incident form to durable memory
&lt;/h2&gt;

&lt;p&gt;An incident starts with the information an on-call engineer actually has: a title, service, severity, symptoms, and perhaps an error string. That is enough to begin an investigation, but not enough to make a useful memory. The record becomes valuable after someone adds a root cause, the steps that exposed it, the resolution, and the attempts that did not work.&lt;/p&gt;

&lt;p&gt;The project keeps that record behind a small interface. The investigation flow does not need to know how memory is stored; it asks the bank to retain an incident or recall a few records for a new one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;MemoryBank&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;readonly&lt;/span&gt; &lt;span class="nx"&gt;backend&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;hindsight&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;local-fallback&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;record&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;MemoryRecord&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;NewIncident&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;RecalledMemory&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;MemoryRecord&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nf"&gt;clear&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the production system, the Hindsight-backed adapter maps those operations to a dedicated incident memory bank. Keeping that boundary explicit matters: the incident workflow owns the meaning of a record, while Hindsight owns long-term retention and retrieval. The &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight GitHub repository&lt;/a&gt; and &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight documentation&lt;/a&gt; describe the retain, recall, and reflect operations that make this more than a transcript archive.&lt;/p&gt;

&lt;p&gt;The record is deliberately structured. A failed attempt is an action paired with its explanation, not a note buried in a postmortem paragraph:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;FailedAttempt&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;why_it_failed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;MemoryRecord&lt;/span&gt; &lt;span class="kd"&gt;extends&lt;/span&gt; &lt;span class="nx"&gt;Incident&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;source&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;seed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;engineer_saved&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;engineer_correction&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;verified_at&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;corrections&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;[];&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This structure captures provenance as well as content. A seed record, an engineer-confirmed resolution, and a later correction should not carry identical authority. When an engineer saves an incident, the workflow records the root cause, resolution, failed action, and lesson, and marks the record verified. When someone corrects a recalled incident, the correction remains attached to that incident rather than becoming an unrelated new document.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the failure belongs beside the fix
&lt;/h2&gt;

&lt;p&gt;The payment incident in the repository makes the design choice concrete. Under peak load, the Payments API exhausted its database connection pool. Restarting the pods cleared the pools and reduced errors for about 25 minutes. Then the leak filled them again. Increasing the database connection limit bought more time, but also increased memory pressure. Neither action fixed the code path that kept connections open inside a retry loop.&lt;/p&gt;

&lt;p&gt;The eventual repair was to release connections in a &lt;code&gt;finally&lt;/code&gt; block, cap pool size per pod, and alert on pool utilization. If memory retained only “restarted pods” and “raised max_connections,” a later investigation could repeat both interventions while missing the actual lesson. If it retained only the final resolution, it would omit the strongest diagnostic clue: a restart that briefly helps can point toward a resource leak.&lt;/p&gt;

&lt;p&gt;So I send failed attempts and their explanations along with the successful resolution. This is where Hindsight’s model of agent memory is useful. Retain stores an experience; recall finds relevant experiences later; reflect can reason across those experiences to form a broader view. The &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;Vectorize overview of agent memory&lt;/a&gt; draws a useful line between remembering text and learning from accumulated experience. In incident response, that line is the difference between finding an old ticket and understanding what the old ticket should change about today’s investigation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recall is context, not a verdict
&lt;/h2&gt;

&lt;p&gt;When a new incident comes in, the application queries memory using the title, service, symptoms, and error message. The production Hindsight integration can use more than literal string overlap: its retrieval combines semantic, keyword, entity, and temporal signals. That matters when the new incident says “pool checkout timed out” and the old postmortem says “connection slots exhausted.” The wording differs; the operational mechanism may not.&lt;/p&gt;

&lt;p&gt;The application passes the recalled records into the investigation as a separate section. It labels the current event and historical memories explicitly, then gives the model a strict rule:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- Keep CURRENT INCIDENT and PAST INCIDENT MEMORIES strictly separate.
- Never state that a past incident IS the current root cause.
- Clearly distinguish SUCCESSFUL PREVIOUS ACTIONS from FAILED PREVIOUS ACTIONS.
- If conflicting memories are provided, say additional current evidence is required.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That instruction is not decorative prompt hygiene. Similar symptoms do not prove the same cause. A service can time out because of a connection leak, a slow dependency, a network change, or a bad rollout. The old incident should help the engineer choose what to inspect, not allow the agent to skip inspection.&lt;/p&gt;

&lt;p&gt;For each recalled memory, the interface also shows why it was returned: matching terms, service, age, engineer verification, and any corrections. Confidence is a property of the evidence we have, not a number the language model gets to invent. Older or unverified memories can still be useful, but the engineer should be able to see their limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a repeat incident looks like
&lt;/h2&gt;

&lt;p&gt;Suppose checkout errors return with a database timeout during peak traffic. Hindsight recalls the earlier Payments API incident because it shares the service and connection symptoms, even if the error message has changed. The investigation can then say: a previous incident found connections leaking in a retry helper; restarting pods only hid the issue temporarily; increasing the connection ceiling delayed failure and raised database memory pressure.&lt;/p&gt;

&lt;p&gt;That is a better starting point than “try restarting the pods.” The agent can recommend checking active and idle database connections, comparing pool checkout latency with the deploy timeline, and tracing the retry path. It should still ask for current evidence before calling the root cause. If the current connection counts are normal and a dependency is timing out, the old incident is a useful contrast, not the answer.&lt;/p&gt;

&lt;p&gt;The workflow also handles disagreement deliberately. If two recalled records for the same service point to different causes, the system surfaces both and asks for more evidence instead of blending them into a confident-sounding compromise. Hindsight can help retrieve and synthesize the history; the application still has to define what uncertainty means during an incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Store the negative result, not just the action.&lt;/strong&gt; “Restarted pods” is an event. “Errors returned after 25 minutes because connections continued leaking” is experience an investigator can use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Preserve provenance.&lt;/strong&gt; Engineer verification and corrections change how much weight a memory deserves. Keep that history visible rather than flattening every record into the same text field.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use memory to choose tests, not to declare causes.&lt;/strong&gt; Historical evidence is valuable because it narrows what to inspect. Production evidence still decides what happened this time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Treat disagreement as data.&lt;/strong&gt; Conflicting postmortems can reflect different failure modes, changed architecture, or weak records. Hiding the conflict makes the answer cleaner and the investigation worse.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make retention part of incident closure.&lt;/strong&gt; If saving the resolution is optional, the memory bank will mostly contain incidents that were easy to document. Capture the failed actions and the final explanation while the people who know the details are still there.&lt;/p&gt;

&lt;p&gt;Incident response already produces the raw material for a useful memory system: hypotheses, failed mitigations, confirmed causes, and corrections. Hindsight gives that history a long-lived place to accumulate and a way to return when it matters. The engineering work is deciding what the agent is allowed to conclude from it. My rule is simple: remember exactly why the last fix failed, then make the next investigator prove whether the same failure is happening again.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffz19vwtf77kmida9t40i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffz19vwtf77kmida9t40i.png" alt=" " width="800" height="543"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzs1y1356phaat2jv3tuj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzs1y1356phaat2jv3tuj.png" alt=" " width="800" height="92"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdqxysq2lmzt57jndtred.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdqxysq2lmzt57jndtred.png" alt=" " width="329" height="646"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv30x4q694y4d9nn7qf3n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv30x4q694y4d9nn7qf3n.png" alt=" " width="317" height="524"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsp44awclvlwrqy4q24oi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsp44awclvlwrqy4q24oi.png" alt=" " width="799" height="223"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ig7cx8pao9jt8kpmee8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ig7cx8pao9jt8kpmee8.png" alt=" " width="800" height="543"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>react</category>
      <category>typescript</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
