I Watched an Agent Get Smarter After One Retained Incident
The first time I ran my incident-response agent against a simulated Payment API outage, it produced a reasonable-sounding hypothesis with no idea it had seen this exact failure mode before. The second time, after I told it to remember what happened, it opened its analysis by citing the prior incident's root cause and the runbook that fixed it. Same code, same LLM, same outage. The only thing that changed was whether the agent had memory.
What the system does
The project is an incident-response agent for a single simulated service: a payment-api that throws HTTP 503s because a Redis connection pool is saturated. When the incident fires, the backend runs a fixed pipeline: pull current metrics, pull recent logs, recall any related historical incidents, and hand all three to an LLM for analysis. The LLM doesn't get to invent evidence — it only reasons over what the tools and memory actually returned.
The stack is deliberately boring: Express on the backend, React and Vite for the dashboard, Groq for the reasoning step, and Hindsight for persistent memory. Boring is a feature here — the story I wanted to tell wasn't about a clever model, it was about what happens when you give a fairly plain LLM pipeline a place to put its experience.
Incident → getMetrics() → Hindsight recall → Groq analysis
→ engineer confirms fix → Hindsight retain → next incident recall
Here's where Hindsight actually sits in the stack, relative to the tools and the model:
POST /investigate
React Dashboard ───────────────▶ Express API
│
┌────────────────┼───────────────────┐
▼ ▼ ▼
getMetrics tool getRecentLogs tool Hindsight (retain/recall)
│ │ │
└───────┬────────┘ │
▼ │
Groq Analysis ◀── labeled historical memories
│
▼
hypothesis + confidence ──▶ React Dashboard ──▶ confirmed fix ──▶ Hindsight
Hindsight isn't a side-channel the model can ignore — it's a required input on the recall path, and a required write on the confirm path.

The dashboard's incident view: lifecycle stages across the top, current evidence on the left, and the Hindsight memory panel on the right — kept visually and structurally separate.
The core problem: current evidence isn't the same as institutional memory
Every incident-response demo I'd seen before this stitched together metrics and logs and threw them at a model. That's fine for the first time you see a failure. It's useless the tenth time, because nothing persists between investigations — the on-call engineer who fixed it last month might remember, but the agent starts from zero every single time.
I didn't want a chat log or a JSON file of "past incidents" bundled into the prompt. That's not memory, that's a fixture, and it doesn't scale past a demo. I wanted the agent to actually retain what it learned when an incident was resolved, and recall only the parts that were relevant to a new one. That's the specific problem Hindsight is built for — retain structured experience, then recall it by semantic relevance instead of grepping a table.
How retain and recall are wired in
The memory boundary lives in one file, memory.service.ts, and it does two things: turn a resolved incident into text worth retaining, and turn an active incident into a query worth recalling against.
Retaining an incident means writing out everything worth remembering — not just the root cause, but the evidence that led there, the investigation steps, and, if a fix was validated, the runbook that fixed it:
async retainIncident(experience: IncidentExperience): Promise<void> {
await this.ensureBank();
const content = [
`Incident: ${experience.incident.description}`,
`Symptoms: ${experience.incident.symptoms.join("; ")}`,
experience.evidence ? `Observed evidence: ${formatMetrics(experience.evidence)}` : "",
`Root cause: ${experience.rootCause}`,
`Resolution: ${experience.resolution}`,
`Outcome: ${experience.outcome}`,
`Lesson learned: ${experience.lesson}`,
experience.runbook ? `Validated runbook: ${experience.runbook.title}` : "",
].filter(Boolean).join("\n");
await this.client.retain(this.bankId, content, {
context: "resolved production incident experience",
documentId: `bugslayers-demo-${experience.incident.id}`,
metadata: { incidentId: experience.incident.id, runbookStatus: "verified" },
});
}
Recall is the mirror image: build a query out of the current incident's service, symptoms, and live evidence, and let Hindsight decide what's relevant.
async recallIncidents(incident: Incident, metrics: ServiceMetrics): Promise<IncidentMemory[]> {
await this.ensureBank();
const query = [
`Service: ${incident.service}`,
`Incident: ${incident.description}`,
`Symptoms: ${incident.symptoms.join("; ")}`,
`Current evidence: ${formatMetrics(metrics)}`,
].join("\n");
const response = await this.client.recall(this.bankId, query, { budget: "mid", maxTokens: 1800 });
return response.results.slice(0, 6).map(({ id, text, type, context, metadata }) =>
({ id, text, type: type ?? "unknown", context: context ?? null, metadata: metadata ?? null }));
}
The bank itself is configured with an explicit retain and reflect mission, which matters more than it looks:
this.bankReady ??= this.client.createBank(this.bankId, {
name: "BugSlayers Incident Memory",
reflectMission: "Remember verified production incident experience... Treat recalled experience as historical context, never as proof about a current incident.",
retainMission: "Extract incident symptoms, services, evidence, confirmed causes, investigation steps, resolutions, outcomes, and reusable lessons. Preserve uncertainty and distinguish hypotheses from confirmed findings.",
});
That's not boilerplate — it's the difference between Hindsight extracting a clean, structured "lesson learned" from a messy incident record versus just indexing raw text.
The part I almost got wrong: trusting memory too much
The obvious failure mode with agent memory is a model that treats a recalled incident as proof about the current one. Two incidents with the same symptoms can have different causes, and an agent that says "we've seen this before, it's definitely the Redis pool" is more dangerous than one with no memory at all, because it sounds confident for the wrong reasons.
So the system prompt for the analysis step explicitly separates current evidence from historical memory, and tells the model not to conflate them:
const systemPrompt = [
"You are an incident-response analyst. Produce a cautious, evidence-grounded hypothesis, never a confirmed root cause.",
"Current metrics and log entries are current tool evidence. Recalled incident memories are historical evidence and must be labeled as such.",
"Metadata-derived runbooks are previously validated historical procedures, not actions already performed on the current incident.",
input.memoryMode === "disabled"
? "This is a no-memory baseline: Hindsight recall was intentionally skipped."
: "If the historical memory list is empty, say no relevant experience was returned by Hindsight.",
"Do not invent logs, metrics, tool results, incidents, or resolutions.",
].join(" ");
I also didn't want every recalled memory treated as an actionable runbook. A runbook only gets surfaced as a recommendation if its metadata explicitly says it was validated, and if the same runbook shows up across multiple retained incidents, it gets grouped and its source incidents tracked instead of listed as duplicate suggestions:
export function getRunbookRecommendations(memories: IncidentMemory[]): RunbookRecommendation[] {
const recommendations = new Map<string, RunbookRecommendation>();
for (const memory of memories) {
const metadata = memory.metadata;
if (!metadata || !["validated", "verified"].includes(metadata.runbookStatus ?? "")) continue;
// ...group by runbookId, accumulate sourceIncidentIds
}
return [...recommendations.values()];
}
That gate — runbookStatus must be validated or verified — is a small line of code, but it's the thing that keeps the agent from recommending a fix just because it once appeared somewhere in memory.
What actually happens, end to end
The investigation endpoint accepts an explicit memoryMode: "disabled" flag, which is what lets me run the same incident twice and compare. With memory disabled, the trace shows a skipped recall step and the analysis reasons from metrics and logs alone — a plausible but unconfirmed guess about a connection or resource limit. With memory enabled against a bank that already has the historical Payment API incident retained, the trace shows Hindsight actually being queried, and the analysis comes back citing Redis connection pool exhaustion directly, with the validated runbook (verify Redis reachability, check pool utilization, increase pool size, monitor error rate) attached and its source incident ID listed.

The investigation trace: each step — metrics, logs, Hindsight recall, Groq analysis — is logged and labeled independently, so a failed or skipped recall is visible instead of silently folded into the hypothesis.
Resolving the incident and confirming the fix worked is a separate step from saving that experience — apply-solution and solution-feedback update the live incident state, but retaining the outcome into Hindsight is an explicit learn call. I kept those separate on purpose: I didn't want a fix that looked like it worked in the moment to get silently canonized into memory before an engineer confirmed it actually held.
Lessons learned
Separate "what the tools say now" from "what we've seen before" at the prompt level, not just in your head. It's tempting to merge current and historical evidence into one blob for the model. Keeping them labeled separately in the prompt is what stops the model from treating a memory as a fact.
Gate recommendations on a validation flag, not just presence in memory. Anything retained is searchable, but not everything retained should be actionable. A runbookStatus metadata field was cheap to add and did most of the work of keeping recall from causing overconfident agents.
Make "no memory" a real, testable state, not just an empty array. Building an explicit no-recall baseline mode into the investigation endpoint was what let me actually see the before/after, instead of assuming memory was helping.
Retaining and confirming a fix should be two different actions. Don't let "the metrics look better" automatically become "write this to permanent memory." A human confirmation step in between is cheap insurance against retaining a fix that only looked successful.
Scope your deletes. The reset path only deletes documents whose IDs are prefixed with a known namespace. Wiping "all memory" during testing is an easy way to nuke something you didn't mean to, and a prefix check costs nothing.
None of this required a complicated agent framework. It required treating memory as a first-class, separately-reasoned-about input — not a hidden convenience feature bolted onto a prompt. Once retain and recall were explicit, boring functions with explicit validation rules, the agent went from "guesses every time" to "gets better at the same job it's done before," which is really the only reason to give an agent memory in the first place. If you're building something similar, the Hindsight docs and the broader case for agent memory are worth reading before you design your own retain/recall boundary — it's easy to get the separation between current evidence and historical memory wrong on the first try.
Top comments (0)