<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: jayasridasari</title>
    <description>The latest articles on DEV Community by jayasridasari (@jayasridasari).</description>
    <link>https://dev.to/jayasridasari</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147786%2Fc3791463-1630-42d1-beaf-55e88ac06970.png</url>
      <title>DEV Community: jayasridasari</title>
      <link>https://dev.to/jayasridasari</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jayasridasari"/>
    <language>en</language>
    <item>
      <title>Why I Never Let Hindsight Memories Override Current Tool Evidence</title>
      <dc:creator>jayasridasari</dc:creator>
      <pubDate>Mon, 28 Sep 2026 18:08:24 +0000</pubDate>
      <link>https://dev.to/jayasridasari/why-i-never-let-hindsight-memories-override-current-tool-evidence-3l9l</link>
      <guid>https://dev.to/jayasridasari/why-i-never-let-hindsight-memories-override-current-tool-evidence-3l9l</guid>
      <description>&lt;p&gt;I've watched a demo where an AI agent confidently declared the root cause of an outage, and I've watched an on-call engineer nod along and apply the "fix" anyway, only to find the real problem was something else entirely. That experience is the reason the incident-response agent I built treats memory as a witness, never a judge.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the system does
&lt;/h2&gt;

&lt;p&gt;The agent investigates a production incident the way a competent on-call engineer would: pull current metrics, pull recent logs, check whether anything like this has happened before, then reason about what's actually going on. The twist is the "check whether this has happened before" step isn't a grep through old tickets — it's a call to &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt;, a persistent memory service, which returns structured facts about prior incidents: root causes, resolutions, outcomes, and lessons, tied to the incidents that produced them.&lt;/p&gt;

&lt;p&gt;Here's the shape of a single investigation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident → getMetrics() → Hindsight recall → Groq analysis
   → engineer confirms fix → Hindsight retain → next incident recall
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The service under investigation is a payment API that starts returning HTTP 503s. The metrics tool reports current numbers — error rate, p95 latency, Redis latency, connection pool utilization. Hindsight, if it has anything relevant, returns memories from past incidents with the same service and similar symptoms. A &lt;a href="https://groq.com/" rel="noopener noreferrer"&gt;Groq&lt;/a&gt;-hosted model gets both current evidence and historical memory and returns a hypothesis, reasoning, a recommended next action, a confidence level, and — critically — a statement of what it's still uncertain about.&lt;/p&gt;

&lt;p&gt;Once an engineer confirms a fix worked, the incident's root cause, resolution, and lesson get written back into Hindsight through a &lt;code&gt;retain&lt;/code&gt; call, so the next 503 spike on the same service benefits from what actually happened this time, not just from what the model guessed.&lt;/p&gt;

&lt;p&gt;Backend is Express and TypeScript, frontend is a Vite/React dashboard that shows the investigation trace step by step, and the whole thing is designed so you can watch each dependency call succeed or fail independently. The &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;Vectorize write-up on agent memory&lt;/a&gt; is a good primer on why this pattern — separating "what an agent recalls" from "what an agent currently observes" — matters more than people expect until they've been burned by an agent that treats speculation as fact.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    subgraph Frontend
        UI["React / Vite dashboard"]
    end
    subgraph Backend["Express API (backend/src)"]
        APP["app.ts investigation workflow"]
        TOOLS["tools/metrics.ts, tools/logs.ts\n(deterministic current evidence)"]
        MEM["memory/memory.service.ts"]
        LLM["llm/analysis.service.ts"]
    end
    HINDSIGHT[("Hindsight\nretain / recall")]
    GROQ[("Groq\nchat completion")]

    UI --&amp;gt;|"investigate incident"| APP
    APP --&amp;gt; TOOLS
    APP --&amp;gt; MEM
    MEM &amp;lt;--&amp;gt;|"retain / recall"| HINDSIGHT
    APP --&amp;gt; LLM
    LLM --&amp;gt;|"current evidence + historical memories"| GROQ
    GROQ --&amp;gt;|"hypothesis, confidence, uncertainty"| LLM
    LLM --&amp;gt; APP
    APP --&amp;gt;|"trace, analysis, runbooks"| UI&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;&lt;em&gt;Hindsight sits alongside the investigation workflow, not inside it — the API only reaches it through the &lt;code&gt;MemoryService&lt;/code&gt; boundary, and every recall result is labeled before it reaches the model.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fczcg2vkfsgb0zggb2twb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fczcg2vkfsgb0zggb2twb.png" alt=" " width="800" height="389"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The dashboard itself keeps that same discipline visible: connection status for Hindsight and the LLM provider sits at the top of every screen, so a disconnected dependency is never silently papered over.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fac0wismbu0pmyd7bu7m8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fac0wismbu0pmyd7bu7m8.png" alt=" " width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Each stage of an investigation — intake, recall, investigate, analyze — is tracked independently, and a failed Hindsight recall doesn't block the stages that don't depend on it. That's the same evidence/memory separation from the code, surfaced as UI state instead of a prompt string.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core problem: memory is dangerous if you don't label it
&lt;/h2&gt;

&lt;p&gt;Early on I made the mistake every memory-backed agent seems to make once: I let historical context and current evidence sit in the same bucket in the prompt. The model would blend them. It would say things like "the error rate is caused by Redis pool exhaustion" with the same tone whether that came from a live metrics reading or from something Hindsight recalled about an incident from three weeks ago. That's a bad habit to bake into an incident response tool, because the entire point of on-call reasoning is knowing what you've verified versus what you suspect.&lt;/p&gt;

&lt;p&gt;So the architecture decision I actually care about in this codebase isn't "use Hindsight" — it's "never let Hindsight's answer look like today's answer."&lt;/p&gt;

&lt;h2&gt;
  
  
  How the separation actually works in code
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;AnalysisInput&lt;/code&gt; the model receives keeps three categories of information distinct at the type level:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;AnalysisInput&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nl"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Incident&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nl"&gt;metrics&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ServiceMetrics&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nl"&gt;logs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ServiceLogEntry&lt;/span&gt;&lt;span class="p"&gt;[];&lt;/span&gt;
    &lt;span class="nl"&gt;memories&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;IncidentMemory&lt;/span&gt;&lt;span class="p"&gt;[];&lt;/span&gt;
    &lt;span class="nl"&gt;memoryMode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;enabled&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;disabled&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nl"&gt;priorSolutionAttempts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;SolutionAttempt&lt;/span&gt;&lt;span class="p"&gt;[];&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;metrics&lt;/code&gt; and &lt;code&gt;logs&lt;/code&gt; are current tool evidence — deterministic, non-negotiable facts about right now. &lt;code&gt;memories&lt;/code&gt; is whatever Hindsight recalled. &lt;code&gt;memoryMode&lt;/code&gt; exists purely so I can run the same incident through the pipeline with recall on or off and diff the outputs, which turned out to be the single most useful debugging tool I built for this project.&lt;/p&gt;

&lt;p&gt;The system prompt is where the labeling gets enforced explicitly, because I learned the hard way that a JSON schema alone doesn't stop a model from conflating sources:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;systemPrompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;You are an incident-response analyst. Produce a cautious, evidence-grounded hypothesis, never a confirmed root cause.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Current metrics and log entries are current tool evidence. Recalled incident memories are historical evidence and must be labeled as such.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Metadata-derived runbooks are previously validated historical procedures, not actions already performed on the current incident.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Review prior solution attempts: do not blindly repeat a failed solution, and treat partial outcomes as evidence that requires further investigation.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;memoryMode&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;disabled&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;This is a no-memory baseline: Hindsight recall was intentionally skipped. Do not imply that historical experience was checked.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;If the historical memory list is empty, say no relevant experience was returned by Hindsight.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Do not invent logs, metrics, tool results, incidents, or resolutions.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Return only a JSON object with string fields possibleRootCause, reasoning, recommendedNextAction, uncertainty, and confidence set to low, medium, or high.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice the model is never told "here's the root cause" from memory — it's told memory is historical and it has to reconcile that with current evidence on its own, out loud, in the &lt;code&gt;reasoning&lt;/code&gt; field. And if recall is disabled, it's explicitly forbidden from implying it checked something it didn't. I added that line after noticing a baseline run once produced a reasoning paragraph that casually referenced "past incidents" it had never actually seen — a small hallucination, but exactly the kind that erodes trust in a tool whose entire value proposition is "trustworthy augmented judgment."&lt;/p&gt;

&lt;p&gt;Retain calls are equally deliberate about what they store. When an engineer's fix works, I write structured, attributable content back to Hindsight rather than a vague summary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="s2"&gt;`Incident: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;description&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s2"&gt;`Service: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;service&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s2"&gt;`Symptoms: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;symptoms&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;evidence&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="s2"&gt;`Observed evidence: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nf"&gt;formatMetrics&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s2"&gt;`Root cause: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rootCause&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s2"&gt;`Resolution: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;resolution&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s2"&gt;`Outcome: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s2"&gt;`Solution verification: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;verificationStatus&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;VERIFIED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="s2"&gt;`Lesson learned: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;lesson&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;runbook&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="s2"&gt;`Validated runbook: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;runbook&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;title&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Boolean&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;bankId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;context&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;resolved production incident experience&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;documentId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`bugslayers-demo-&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;experience&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;...`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;incidentId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;service&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;verificationStatus&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;runbookStatus&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;verificationStatus&lt;/code&gt; field is what turns this from "an LLM's guess about what happened" into "an audited record." A resolution only becomes a &lt;code&gt;runbookStatus=validated&lt;/code&gt; runbook recommendation if an engineer confirmed it worked. If a fix failed or only partially worked, that gets retained too — with &lt;code&gt;verificationStatus: "FAILED"&lt;/code&gt; or &lt;code&gt;"PARTIAL"&lt;/code&gt; — specifically so the next investigation's prompt includes a line telling the model not to blindly repeat something that already didn't work. That's a detail I almost skipped, and it turned out to be one of the more important ones: an agent that recalls only successes will happily suggest the same failed fix twice.&lt;/p&gt;

&lt;p&gt;Recall queries are built from what's actually happening, not from a static incident ID, which is what makes the memory relevant instead of just present:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;recallIncidents&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Incident&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;metrics&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ServiceMetrics&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;IncidentMemory&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ensureBank&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="s2"&gt;`Service: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;service&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;`Incident: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;description&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;`Symptoms: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;symptoms&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;`Current evidence: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nf"&gt;formatMetrics&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;metrics&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="c1"&gt;// ...&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What this looks like in practice
&lt;/h2&gt;

&lt;p&gt;Take the payment API incident: error rate at 18.2%, p95 latency at 840ms, Redis latency at 420ms, connection pool at 100% utilization.&lt;/p&gt;

&lt;p&gt;With recall disabled, the model has to reason from current numbers alone. It can reasonably suspect Redis pressure from the pool utilization figure, but it has no way to know that a prior 503 spike on this exact service was traced to pool exhaustion and fixed by doubling the pool size from 50 to 100. Its &lt;code&gt;uncertainty&lt;/code&gt; field says as much — it hasn't seen this exact shape of failure resolved before.&lt;/p&gt;

&lt;p&gt;With recall enabled, Hindsight returns that prior incident's memory: same service, same symptom cluster, root cause "Redis connection pool exhaustion," resolution "increase pool size from 50 to 100," outcome "error rate returned to normal," and a &lt;code&gt;runbookStatus=validated&lt;/code&gt; runbook with concrete steps. The model's &lt;code&gt;possibleRootCause&lt;/code&gt; becomes more specific, &lt;code&gt;confidence&lt;/code&gt; goes up, and — this is the part I care about — the &lt;code&gt;reasoning&lt;/code&gt; field explicitly separates "current evidence shows X" from "a similar incident was previously resolved by Y," instead of merging them into one unqualified claim. The recommended action still has to be justified against current evidence; historical experience informs it but doesn't approve it.&lt;/p&gt;

&lt;p&gt;When the engineer applies the fix and reports success, the pool-saturation runbook gets retained with a fresh source incident ID attached. The next time this service has a 503 spike, Hindsight can now cite two incidents, not one, and the runbook recommendation groups them by runbook ID so the dashboard shows "this procedure has worked twice" rather than two disconnected notes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Separate evidence from experience at the type level, not just in the prompt.&lt;/strong&gt; Keeping &lt;code&gt;metrics&lt;/code&gt;/&lt;code&gt;logs&lt;/code&gt; and &lt;code&gt;memories&lt;/code&gt; as distinct fields in &lt;code&gt;AnalysisInput&lt;/code&gt; made it structurally awkward to accidentally blend them, which matters more than any amount of prompt wording once the codebase has multiple contributors.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A memory system needs a "no memory" mode you can actually invoke.&lt;/strong&gt; Building &lt;code&gt;memoryMode: "disabled"&lt;/code&gt; as a first-class option, not a debug hack, let me directly compare Hindsight-informed reasoning against a cold baseline on the same incident. That comparison is the fastest way to prove memory is pulling its weight instead of just adding noise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verification status is what makes retained memory trustworthy.&lt;/strong&gt; Storing every resolution attempt — including failed and partial ones — with an explicit &lt;code&gt;verificationStatus&lt;/code&gt; stopped the agent from treating "a model once suggested this" as equivalent to "an engineer confirmed this worked." Without that distinction, a memory system just accumulates confident-sounding guesses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Attribute recalled facts to their source incidents.&lt;/strong&gt; Grouping recalled runbook metadata by runbook ID and keeping the list of source incident IDs turned "the AI thinks this might work" into "this exact procedure resolved two prior incidents," which is a very different sentence to read at 3 a.m.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tell the model explicitly when it's being tested cold.&lt;/strong&gt; The line forbidding the model from implying it checked history during a disabled-recall run seems like a small thing, but it's the difference between a baseline you can trust and one that quietly cheats.&lt;/p&gt;

&lt;p&gt;If you're building anything where an LLM's job is to reconcile "what just happened" with "what we've learned before," the discipline that matters isn't the retrieval algorithm — the &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight GitHub repo&lt;/a&gt; handles that well. It's making sure your prompt, your types, and your storage layer all agree on which claims are provisional and which ones are earned.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>typescript</category>
      <category>agents</category>
      <category>memory</category>
    </item>
  </channel>
</rss>
