DEV Community

Ramyalatha
Ramyalatha

Posted on

Why I Stopped Trusting LLM Root Causes Until Hindsight Backed Them Up

The first time I ran my incident-response agent against a fake Payment API outage, it gave me a confident, well-written, completely unearned answer. It said Redis connection pool exhaustion was the root cause. It was right — but only by coincidence. It had no memory of ever seeing this failure before, no evidence it had checked anything beyond the metrics I handed it, and no way to distinguish a lucky guess from a grounded diagnosis. That gap between "sounds right" and "is right" is the whole reason I built this thing, and it's the reason Hindsight ended up being the most important dependency in the codebase, not an add-on.

What the system does

The agent is an incident-response investigator for a single, well-understood failure mode: a payment service returning HTTP 503s under Redis connection pool saturation. An engineer opens an active incident, hits investigate, and the backend does four things in sequence: pull current metrics, recall relevant historical incidents from Hindsight, pull recent logs, and hand all of it to an LLM (Groq, in my case) to produce a hypothesis.

The architecture is intentionally boring:

  • frontend/ — a React/Vite dashboard that shows the incident, the trace of tool calls, and the model's hypothesis side by side with any recalled memory.
  • backend/src/tools/ — deterministic metrics and log generators. No randomness, no flakiness. If the model's story doesn't match these numbers, that's a bug in the model, not the tools.
  • backend/src/memory/ — the boundary to Hindsight. This is the file I've rewritten the most.
  • backend/src/llm/ — the Groq client and response validation.

Each investigation produces a structured trace — metrics, recall, logs, analysis — with timing and status for every step, because I got tired of not knowing which stage silently failed when something went wrong (see below).

Here's where Hindsight actually sits in the stack — it's a peer dependency next to the LLM, not a cache in front of it:

flowchart LR
    UI["React/Vite Dashboard"] -->|investigate incident| API["Express API (backend/src/app.ts)"]
    API --> Tools["Metrics + Logs Tools\n(deterministic, current evidence)"]
    API -->|recall query| Hindsight[("Hindsight\nincident memory bank")]
    Tools --> Prompt["Analysis Prompt Builder"]
    Hindsight -->|recalled memories + runbooks| Prompt
    Prompt --> LLM["Groq LLM\n(hypothesis + confidence)"]
    LLM --> API
    API -->|retain resolved outcome| Hindsight
    API --> UI

Current evidence (metrics, logs) and historical evidence (Hindsight recall) flow into the prompt builder as separate, labeled inputs — they never merge into one context blob before reaching the model. When an incident is resolved, the outcome flows back into Hindsight via retain, so the next investigation on that service starts one data point ahead.

Here's what that looks like in the actual dashboard — current evidence and recalled Hindsight memory rendered as separate panels, exactly as described above:

Incident dashboard showing current evidence and recalled Hindsight memory side by side

The core problem: memory that lies by omission

The dangerous failure mode for an LLM-based investigator isn't a wrong answer. It's a right-sounding answer that isn't actually grounded in anything. Early on, when I hardcoded the historical incident directly into the prompt, the model would cite "past incidents" even in test runs where I'd intentionally passed it an empty memory list. It was pattern-matching on training data, not on my data. That's the point where I realized I needed a real memory system with an explicit contract, not a JSON blob I trusted the model to only-sometimes reference.

That's what pulled in Hindsight. It's a separate memory service the agent talks to over its own API — not a vector database I query directly, but a system with its own retain and recall semantics and its own extraction mission. Reading through the Hindsight docs reframed how I thought about the problem: memory for an agent isn't "store text, search text later." It's closer to what's described in Vectorize's writeup on agent memory — memory needs its own retention policy, its own retrieval semantics, and a way to keep provenance attached to every fact, or it becomes just another hallucination vector with worse UX.

How retain actually works

When an incident is resolved, I don't just dump a summary into Hindsight. I construct a structured narrative and attach real metadata:

async retainIncident(experience: IncidentExperience): Promise<void> {
  await this.ensureBank();
  const content = [
    `Incident: ${experience.incident.description}`,
    `Service: ${experience.incident.service}`,
    `Symptoms: ${experience.incident.symptoms.join("; ")}`,
    experience.evidence ? `Observed evidence: ${formatMetrics(experience.evidence)}` : "",
    `Root cause: ${experience.rootCause}`,
    `Resolution: ${experience.resolution}`,
    `Outcome: ${experience.outcome}`,
    `Solution verification: ${experience.verificationStatus ?? "VERIFIED"}`,
    `Lesson learned: ${experience.lesson}`,
    experience.runbook ? `Validated runbook: ${experience.runbook.title}` : "",
  ].filter(Boolean).join("\n");

  await this.client.retain(this.bankId, content, {
    context: "resolved production incident experience",
    documentId: `bugslayers-demo-${experience.incident.id}${experience.attemptId ? `-attempt-${experience.attemptId}` : ""}`,
    metadata: {
      incidentId: experience.incident.id,
      service: experience.incident.service,
      verificationStatus: experience.verificationStatus ?? "VERIFIED",
      ...(experience.runbook ? {
        runbookId: experience.runbook.id,
        runbookSteps: JSON.stringify(experience.runbook.steps),
        runbookStatus: (experience.verificationStatus ?? "VERIFIED").toLowerCase(),
      } : {}),
    },
  });
}
Enter fullscreen mode Exit fullscreen mode

Two decisions here mattered more than I expected. First, I retain a verificationStatus — VERIFIED, FAILED, or PARTIAL — not just successful resolutions. If an engineer tries a fix and it doesn't work, that gets retained too, with a lesson explicitly telling future investigations not to repeat it blindly. Second, runbook metadata only gets a runbookStatus: "validated" tag when a human actually confirmed the fix worked. A hypothesis the model proposed but nobody verified never gets promoted to a validated runbook, no matter how confident the analysis sounded.

How recall keeps the model honest

On the query side, recall is built from the current incident and current tool evidence, not from the model's own reasoning:

async recallIncidents(incident: Incident, metrics: ServiceMetrics): Promise<IncidentMemory[]> {
  await this.ensureBank();
  const query = [
    `Service: ${incident.service}`,
    `Incident: ${incident.description}`,
    `Symptoms: ${incident.symptoms.join("; ")}`,
    `Current evidence: ${formatMetrics(metrics)}`,
  ].join("\n");
  const response = await this.client.recall(this.bankId, query, { budget: "mid", maxTokens: 1800 });
  return response.results.slice(0, 6).map(({ id, text, type, context, metadata }) =>
    ({ id, text, type: type ?? "unknown", context: context ?? null, metadata: metadata ?? null }));
}
Enter fullscreen mode Exit fullscreen mode

The results come back as a list of discrete memories, each tagged with its source incident ID and metadata, not a single flattened summary. That matters downstream: I group recalled facts by runbookId and collect every sourceIncidentId that contributed to a recommendation, so the UI can show "this runbook was validated across 3 prior incidents" instead of an unattributed claim.

The part I care about most is what happens in the prompt itself. Historical memory and current evidence are never merged into one undifferentiated context blob:

input.memoryMode === "disabled"
  ? "This is a no-memory baseline: Hindsight recall was intentionally skipped. Do not imply that historical experience was checked."
  : "If the historical memory list is empty, say no relevant experience was returned by Hindsight.",
"Metadata-derived runbooks are previously validated historical procedures, not actions already performed on the current incident.",
"Review prior solution attempts: do not blindly repeat a failed solution, and treat partial outcomes as evidence that requires further investigation.",
Enter fullscreen mode Exit fullscreen mode

The API even supports an explicit memoryMode: "disabled" flag on the investigate endpoint, purely so I can run a real side-by-side: same incident, same tools, memory on versus off. That comparison is the whole demo, and it's also the best regression test I have for whether the system is actually using its memory or just narrating around it.

What it looks like in practice

Cold start, no memory: the agent investigates the Payment API 503s using only current metrics and logs. It correctly flags Redis latency and pool utilization as suspicious but hedges everything — medium confidence, "further investigation into Redis pool sizing may be warranted." Reasonable, generic, exactly what you'd expect from evidence alone.

After I retain the resolved historical incident and re-run the same investigation with recall enabled, the response changes shape. The recalled memory shows up as a labeled, separate block — same Redis pool signature, previously confirmed root cause, a validated runbook with concrete steps and its prior outcome ("the 50-to-100 pool increase returned the error rate to normal"). The model's hypothesis references that recalled experience explicitly rather than presenting it as its own reasoning, and confidence moves from medium to high because there's now a matching, verified precedent — not because the model got more articulate.

Then I close the loop: apply the recommended fix, confirm success, and the outcome — including the before/after metrics — gets retained back into Hindsight as a new document. The next incident on this service starts with two prior data points instead of one.

Lessons learned

Separate "what happened" from "what we think caused it." The system prompt is explicit that current tool evidence and recalled historical memory are different categories, and the model is instructed never to treat memory as proof about the current incident. This single distinction eliminated most of the false-confidence problem I started with.

Retain failures, not just successes. A FAILED or PARTIAL solution attempt is retained with the same rigor as a verified fix. Without this, the agent would eventually recommend the same broken fix twice, because "we haven't succeeded yet" and "we've never tried this" look identical if you only remember successes.

Don't let the model self-certify a runbook. Runbook status only flips to validated when a human confirms the outcome, via the solution-feedback endpoint, and that confirmation is captured in the retained metadata. The model can hypothesize; only a human can validate.

Traceability isn't optional once memory is involved. Every investigation returns a full trace of which tool ran, whether it succeeded, and how long it took — including whether Hindsight recall was skipped, attempted, or failed. Once you're mixing live tool output with recalled memory in one prompt, silent failures in either path produce answers that look fine and are quietly wrong.

Memory needs its own mission, not just a table. Configuring Hindsight's bank with an explicit retainMission and reflectMission — extract confirmed causes, preserve uncertainty, distinguish hypotheses from verified findings — did more for answer quality than any prompt tweak on the analysis side. The retrieval system needs its own opinion about what's worth remembering; you can't bolt that onto a generic vector store after the fact.

The unglamorous truth is that the LLM call is the easy 20% of this system. The hard 80% is deciding what gets remembered, under what confidence, attributed to what source, and how you stop a plausible-sounding model from quietly overstating what it actually knows. That's the part Hindsight is built for, and it's the part I'd rebuild first if I were starting over.

Top comments (0)