When an API goes down at 2 AM, the last thing an on-call engineer needs is generic advice.
Yet almost every LLM-based troubleshooting tool treats every outage as day zero. When our checkout service started returning HTTP 401 Unauthorized right after a routine release, a standard LLM copilot dutifully suggested checking JWT signatures, inspecting expiration timestamps, verifying algorithm headers, and checking database connectivity.
Standard textbook advice. But it missed the real story: two weeks earlier, another deployment had failed with the exact same error because Kubernetes pods were mounting an outdated secret version (payments-jwt-v3) after the issuer rotated to payments-jwt-v4.
The organization had already solved this problem. The knowledge existed, but the AI copilot had amnesia.
To fix this, we built BugFix—an incident-response system powered by agent memory using Hindsight. Instead of treating AI hallucinations as organizational truth, BugFix compares a stateless baseline with a memory-grounded diagnosis, and only retains resolutions explicitly verified by human engineers.
The Architecture: Making Memory Durable and Isolated
Most AI agent prototypes try to solve memory by shoving past chat history into a giant context window. That falls apart in production: context costs balloon, hallucinations multiply, and noise from unrelated teams drowns out critical signals.
We designed BugFix around three core tenets:
-
Project-Scoped Banks: Incident memories are partitioned by team/project (
bugfix-<project-id>). What happens in payments never pollutes the auth cluster. - Evidence-Gated Retention: Speculative model diagnoses are never saved to memory automatically. Only human-verified fixes become durable facts.
- Structured Incident Contracts: Strict schemas ensure both LLM diagnoses and memory records remain deterministic and machine-readable.
flowchart LR
A[Incident Report: Symptoms & Logs] --> H[Hindsight Recall]
A --> S[Stateless Diagnosis Baseline]
H --> G[Memory-Grounded Diagnosis]
S --> UI[Side-by-Side Comparison]
G --> UI
UI --> C[Engineer Confirms Root Cause & Fix]
C --> R[Hindsight Retain]
R -.->|Compounds Knowledge| H
How It Works in Code
The system integrates Node.js and Express with the @vectorize-io/hindsight-client and Groq's high-speed inference engine.
1. Recalling Project-Scoped Context
When an incident is reported, BugFix queries the project bank using multi-vector search (semantic, keyword, and entity graph retrieval) documented in the Hindsight documentation:
async function recallMemories(projectId, incident) {
await ensureBank(projectId);
const query = [
incident.title,
incident.service,
incident.environment,
incident.symptoms,
incident.logs
].filter(Boolean).join("\n");
const result = await hindsight.recall(bankId(projectId), query, {
budget: "high",
maxTokens: 1400,
types: ["world", "experience", "observation"],
});
return (result.results || []).slice(0, 6).map((memory) => ({
id: memory.id,
text: memory.text,
type: memory.type,
context: memory.context,
occurredAt: memory.occurred_start || memory.mentioned_at,
entities: memory.entities || [],
tags: memory.tags || [],
}));
}
2. Side-by-Side Diagnosis Generation
We query the model twice in parallel: once with zero memory (the stateless baseline) and once with recalled operational facts.
We enforce a strict security boundary in the prompt: logs and recalled memory are treated purely as untrusted data, never as executable instructions.
const [baseline, memoryGuided] = await Promise.all([
incident.memoryMode === "compare" ? diagnose(incident, []) : Promise.resolve(null),
diagnose(incident, memories),
]);
3. Retaining Verified Resolutions
When the on-call engineer fixes the issue, they submit the confirmed root cause and remediation. Only then does BugFix call hindsight.retain():
async function storeResolution(projectId, incident, resolution) {
await ensureBank(projectId);
const content = [
`Confirmed incident: ${incident.title}`,
`Service: ${incident.service}`,
`Environment: ${incident.environment}`,
`Severity: ${incident.severity}`,
`Symptoms: ${incident.symptoms}`,
`Root cause: ${resolution.rootCause}`,
`Verified fix: ${resolution.fix}`,
`Outcome: ${resolution.outcome}`,
resolution.technologies.length
? `Technologies: ${resolution.technologies.join(", ")}`
: "",
].filter(Boolean).join("\n");
await hindsight.retain(bankId(projectId), content, {
context: `confirmed incident resolution for ${incident.service}`,
documentId: `incident-${incident.id}`,
tags: [
"confirmed-resolution",
`service:${incident.service.toLowerCase().replace(/[^a-z0-9]+/g, "-")}`
],
timestamp: new Date().toISOString(),
});
}
Real-World Results: Before vs. After Memory
Here is the exact difference observed in our investigation workflow:
| Metric / Step | Stateless AI Baseline | BugFix with Hindsight Memory |
|---|---|---|
| Initial Guess | "Check JWT signature algorithm, verify private key format, test token expiration." | "Production pods likely mounting old payments-jwt-v3 secret rather than rotated payments-jwt-v4." |
| Actionable Checks | 5 generic verification steps across codebase | Specific kubectl secret inspection command |
| Time to Resolution (MTTR) | 25–40 minutes of manual debugging | Under 2 minutes |
| Evidence Trace | Black-box output | 4 linked facts with exact timestamps & tags |
Key Takeaways for Building AI Agents
- Don't Let the Model's Guess Become Your Memory: The most dangerous failure mode in agent architectures is saving AI outputs back into the knowledge base automatically. Memory must be verified by humans before retention.
- Project Isolation is Non-Negotiable: A shared global vector store causes cross-talk between services. Partitioning memory by project bank keeps recall crisp and prevents noisy context.
- Treat Context as Evidence, Not Instructions: Always wrap recalled memories and logs in strict XML or data boundaries to prevent prompt injection and keep diagnostics grounded.
- Durable Memory Outperforms Long Context: Rather than feeding thousands of lines of raw logs into context windows, retaining structured resolutions turns ephemeral incidents into compounding operational intelligence.
Links & Resources
- Project Repository: GitHub - bugfix-memory-agent
- Memory Engine: Hindsight GitHub Repository
- Documentation: Hindsight Docs
- Deep Dive: What is Agent Memory?
Top comments (0)