DEV Community

jayasridasari
jayasridasari

Posted on

Why I Never Let Hindsight Memories Override Current Tool Evidence

I've watched a demo where an AI agent confidently declared the root cause of an outage, and I've watched an on-call engineer nod along and apply the "fix" anyway, only to find the real problem was something else entirely. That experience is the reason the incident-response agent I built treats memory as a witness, never a judge.

What the system does

The agent investigates a production incident the way a competent on-call engineer would: pull current metrics, pull recent logs, check whether anything like this has happened before, then reason about what's actually going on. The twist is the "check whether this has happened before" step isn't a grep through old tickets — it's a call to Hindsight, a persistent memory service, which returns structured facts about prior incidents: root causes, resolutions, outcomes, and lessons, tied to the incidents that produced them.

Here's the shape of a single investigation:

Incident → getMetrics() → Hindsight recall → Groq analysis
   → engineer confirms fix → Hindsight retain → next incident recall
Enter fullscreen mode Exit fullscreen mode

The service under investigation is a payment API that starts returning HTTP 503s. The metrics tool reports current numbers — error rate, p95 latency, Redis latency, connection pool utilization. Hindsight, if it has anything relevant, returns memories from past incidents with the same service and similar symptoms. A Groq-hosted model gets both current evidence and historical memory and returns a hypothesis, reasoning, a recommended next action, a confidence level, and — critically — a statement of what it's still uncertain about.

Once an engineer confirms a fix worked, the incident's root cause, resolution, and lesson get written back into Hindsight through a retain call, so the next 503 spike on the same service benefits from what actually happened this time, not just from what the model guessed.

Backend is Express and TypeScript, frontend is a Vite/React dashboard that shows the investigation trace step by step, and the whole thing is designed so you can watch each dependency call succeed or fail independently. The Vectorize write-up on agent memory is a good primer on why this pattern — separating "what an agent recalls" from "what an agent currently observes" — matters more than people expect until they've been burned by an agent that treats speculation as fact.

flowchart LR
    subgraph Frontend
        UI["React / Vite dashboard"]
    end
    subgraph Backend["Express API (backend/src)"]
        APP["app.ts investigation workflow"]
        TOOLS["tools/metrics.ts, tools/logs.ts\n(deterministic current evidence)"]
        MEM["memory/memory.service.ts"]
        LLM["llm/analysis.service.ts"]
    end
    HINDSIGHT[("Hindsight\nretain / recall")]
    GROQ[("Groq\nchat completion")]

    UI -->|"investigate incident"| APP
    APP --> TOOLS
    APP --> MEM
    MEM <-->|"retain / recall"| HINDSIGHT
    APP --> LLM
    LLM -->|"current evidence + historical memories"| GROQ
    GROQ -->|"hypothesis, confidence, uncertainty"| LLM
    LLM --> APP
    APP -->|"trace, analysis, runbooks"| UI

Hindsight sits alongside the investigation workflow, not inside it — the API only reaches it through the MemoryService boundary, and every recall result is labeled before it reaches the model.

The dashboard itself keeps that same discipline visible: connection status for Hindsight and the LLM provider sits at the top of every screen, so a disconnected dependency is never silently papered over.

Each stage of an investigation — intake, recall, investigate, analyze — is tracked independently, and a failed Hindsight recall doesn't block the stages that don't depend on it. That's the same evidence/memory separation from the code, surfaced as UI state instead of a prompt string.

The core problem: memory is dangerous if you don't label it

Early on I made the mistake every memory-backed agent seems to make once: I let historical context and current evidence sit in the same bucket in the prompt. The model would blend them. It would say things like "the error rate is caused by Redis pool exhaustion" with the same tone whether that came from a live metrics reading or from something Hindsight recalled about an incident from three weeks ago. That's a bad habit to bake into an incident response tool, because the entire point of on-call reasoning is knowing what you've verified versus what you suspect.

So the architecture decision I actually care about in this codebase isn't "use Hindsight" — it's "never let Hindsight's answer look like today's answer."

How the separation actually works in code

The AnalysisInput the model receives keeps three categories of information distinct at the type level:

export interface AnalysisInput {
    incident: Incident;
    metrics: ServiceMetrics;
    logs: ServiceLogEntry[];
    memories: IncidentMemory[];
    memoryMode: "enabled" | "disabled";
    priorSolutionAttempts: SolutionAttempt[];
}
Enter fullscreen mode Exit fullscreen mode

metrics and logs are current tool evidence — deterministic, non-negotiable facts about right now. memories is whatever Hindsight recalled. memoryMode exists purely so I can run the same incident through the pipeline with recall on or off and diff the outputs, which turned out to be the single most useful debugging tool I built for this project.

The system prompt is where the labeling gets enforced explicitly, because I learned the hard way that a JSON schema alone doesn't stop a model from conflating sources:

const systemPrompt = [
  "You are an incident-response analyst. Produce a cautious, evidence-grounded hypothesis, never a confirmed root cause.",
  "Current metrics and log entries are current tool evidence. Recalled incident memories are historical evidence and must be labeled as such.",
  "Metadata-derived runbooks are previously validated historical procedures, not actions already performed on the current incident.",
  "Review prior solution attempts: do not blindly repeat a failed solution, and treat partial outcomes as evidence that requires further investigation.",
  input.memoryMode === "disabled"
    ? "This is a no-memory baseline: Hindsight recall was intentionally skipped. Do not imply that historical experience was checked."
    : "If the historical memory list is empty, say no relevant experience was returned by Hindsight.",
  "Do not invent logs, metrics, tool results, incidents, or resolutions.",
  "Return only a JSON object with string fields possibleRootCause, reasoning, recommendedNextAction, uncertainty, and confidence set to low, medium, or high.",
].join(" ");
Enter fullscreen mode Exit fullscreen mode

Notice the model is never told "here's the root cause" from memory — it's told memory is historical and it has to reconcile that with current evidence on its own, out loud, in the reasoning field. And if recall is disabled, it's explicitly forbidden from implying it checked something it didn't. I added that line after noticing a baseline run once produced a reasoning paragraph that casually referenced "past incidents" it had never actually seen — a small hallucination, but exactly the kind that erodes trust in a tool whose entire value proposition is "trustworthy augmented judgment."

Retain calls are equally deliberate about what they store. When an engineer's fix works, I write structured, attributable content back to Hindsight rather than a vague summary:

const content = [
  `Incident: ${experience.incident.description}`,
  `Service: ${experience.incident.service}`,
  `Symptoms: ${experience.incident.symptoms.join("; ")}`,
  experience.evidence ? `Observed evidence: ${formatMetrics(experience.evidence)}` : "",
  `Root cause: ${experience.rootCause}`,
  `Resolution: ${experience.resolution}`,
  `Outcome: ${experience.outcome}`,
  `Solution verification: ${experience.verificationStatus ?? "VERIFIED"}`,
  `Lesson learned: ${experience.lesson}`,
  experience.runbook ? `Validated runbook: ${experience.runbook.title}` : "",
].filter(Boolean).join("\n");

await this.client.retain(this.bankId, content, {
  context: "resolved production incident experience",
  documentId: `bugslayers-demo-${experience.incident.id}...`,
  metadata: { incidentId, service, verificationStatus, runbookStatus, ... },
});
Enter fullscreen mode Exit fullscreen mode

The verificationStatus field is what turns this from "an LLM's guess about what happened" into "an audited record." A resolution only becomes a runbookStatus=validated runbook recommendation if an engineer confirmed it worked. If a fix failed or only partially worked, that gets retained too — with verificationStatus: "FAILED" or "PARTIAL" — specifically so the next investigation's prompt includes a line telling the model not to blindly repeat something that already didn't work. That's a detail I almost skipped, and it turned out to be one of the more important ones: an agent that recalls only successes will happily suggest the same failed fix twice.

Recall queries are built from what's actually happening, not from a static incident ID, which is what makes the memory relevant instead of just present:

async recallIncidents(incident: Incident, metrics: ServiceMetrics): Promise<IncidentMemory[]> {
  await this.ensureBank();
  const query = [
    `Service: ${incident.service}`,
    `Incident: ${incident.description}`,
    `Symptoms: ${incident.symptoms.join("; ")}`,
    `Current evidence: ${formatMetrics(metrics)}`,
  ].join("\n");
  // ...
}
Enter fullscreen mode Exit fullscreen mode

What this looks like in practice

Take the payment API incident: error rate at 18.2%, p95 latency at 840ms, Redis latency at 420ms, connection pool at 100% utilization.

With recall disabled, the model has to reason from current numbers alone. It can reasonably suspect Redis pressure from the pool utilization figure, but it has no way to know that a prior 503 spike on this exact service was traced to pool exhaustion and fixed by doubling the pool size from 50 to 100. Its uncertainty field says as much — it hasn't seen this exact shape of failure resolved before.

With recall enabled, Hindsight returns that prior incident's memory: same service, same symptom cluster, root cause "Redis connection pool exhaustion," resolution "increase pool size from 50 to 100," outcome "error rate returned to normal," and a runbookStatus=validated runbook with concrete steps. The model's possibleRootCause becomes more specific, confidence goes up, and — this is the part I care about — the reasoning field explicitly separates "current evidence shows X" from "a similar incident was previously resolved by Y," instead of merging them into one unqualified claim. The recommended action still has to be justified against current evidence; historical experience informs it but doesn't approve it.

When the engineer applies the fix and reports success, the pool-saturation runbook gets retained with a fresh source incident ID attached. The next time this service has a 503 spike, Hindsight can now cite two incidents, not one, and the runbook recommendation groups them by runbook ID so the dashboard shows "this procedure has worked twice" rather than two disconnected notes.

Lessons learned

Separate evidence from experience at the type level, not just in the prompt. Keeping metrics/logs and memories as distinct fields in AnalysisInput made it structurally awkward to accidentally blend them, which matters more than any amount of prompt wording once the codebase has multiple contributors.

A memory system needs a "no memory" mode you can actually invoke. Building memoryMode: "disabled" as a first-class option, not a debug hack, let me directly compare Hindsight-informed reasoning against a cold baseline on the same incident. That comparison is the fastest way to prove memory is pulling its weight instead of just adding noise.

Verification status is what makes retained memory trustworthy. Storing every resolution attempt — including failed and partial ones — with an explicit verificationStatus stopped the agent from treating "a model once suggested this" as equivalent to "an engineer confirmed this worked." Without that distinction, a memory system just accumulates confident-sounding guesses.

Attribute recalled facts to their source incidents. Grouping recalled runbook metadata by runbook ID and keeping the list of source incident IDs turned "the AI thinks this might work" into "this exact procedure resolved two prior incidents," which is a very different sentence to read at 3 a.m.

Tell the model explicitly when it's being tested cold. The line forbidding the model from implying it checked history during a disabled-recall run seems like a small thing, but it's the difference between a baseline you can trust and one that quietly cheats.

If you're building anything where an LLM's job is to reconcile "what just happened" with "what we've learned before," the discipline that matters isn't the retrieval algorithm — the Hindsight GitHub repo handles that well. It's making sure your prompt, your types, and your storage layer all agree on which claims are provisional and which ones are earned.

Top comments (0)