DEV Community

Sushma Kumari
Sushma Kumari

Posted on

Why Your Incident Copilot Needs Memory, Not Just Prompts

When an API goes down at 2 AM, the last thing an on-call engineer needs is generic advice.

Yet almost every LLM-based troubleshooting tool treats every outage as day zero. When our checkout service started returning HTTP 401 Unauthorized right after a routine release, a standard LLM copilot dutifully suggested checking JWT signatures, inspecting expiration timestamps, verifying algorithm headers, and checking database connectivity.

Standard textbook advice. But it missed the real story: two weeks earlier, another deployment had failed with the exact same error because Kubernetes pods were mounting an outdated secret version (payments-jwt-v3) after the issuer rotated to payments-jwt-v4.

The organization had already solved this problem. The knowledge existed, but the AI copilot had amnesia.

To fix this, we built BugFix—an incident-response system powered by agent memory using Hindsight. Instead of treating AI hallucinations as organizational truth, BugFix compares a stateless baseline with a memory-grounded diagnosis, and only retains resolutions explicitly verified by human engineers.


The Architecture: Making Memory Durable and Isolated

Most AI agent prototypes try to solve memory by shoving past chat history into a giant context window. That falls apart in production: context costs balloon, hallucinations multiply, and noise from unrelated teams drowns out critical signals.

We designed BugFix around three core tenets:

  1. Project-Scoped Banks: Incident memories are partitioned by team/project (bugfix-<project-id>). What happens in payments never pollutes the auth cluster.
  2. Evidence-Gated Retention: Speculative model diagnoses are never saved to memory automatically. Only human-verified fixes become durable facts.
  3. Structured Incident Contracts: Strict schemas ensure both LLM diagnoses and memory records remain deterministic and machine-readable.
flowchart LR
    A[Incident Report: Symptoms & Logs] --> H[Hindsight Recall]
    A --> S[Stateless Diagnosis Baseline]
    H --> G[Memory-Grounded Diagnosis]
    S --> UI[Side-by-Side Comparison]
    G --> UI
    UI --> C[Engineer Confirms Root Cause & Fix]
    C --> R[Hindsight Retain]
    R -.->|Compounds Knowledge| H

How It Works in Code

The system integrates Node.js and Express with the @vectorize-io/hindsight-client and Groq's high-speed inference engine.

1. Recalling Project-Scoped Context

When an incident is reported, BugFix queries the project bank using multi-vector search (semantic, keyword, and entity graph retrieval) documented in the Hindsight documentation:

async function recallMemories(projectId, incident) {
    await ensureBank(projectId);

    const query = [
        incident.title,
        incident.service,
        incident.environment,
        incident.symptoms,
        incident.logs
    ].filter(Boolean).join("\n");

    const result = await hindsight.recall(bankId(projectId), query, {
        budget: "high",
        maxTokens: 1400,
        types: ["world", "experience", "observation"],
    });

    return (result.results || []).slice(0, 6).map((memory) => ({
        id: memory.id,
        text: memory.text,
        type: memory.type,
        context: memory.context,
        occurredAt: memory.occurred_start || memory.mentioned_at,
        entities: memory.entities || [],
        tags: memory.tags || [],
    }));
}
Enter fullscreen mode Exit fullscreen mode

2. Side-by-Side Diagnosis Generation

We query the model twice in parallel: once with zero memory (the stateless baseline) and once with recalled operational facts.

We enforce a strict security boundary in the prompt: logs and recalled memory are treated purely as untrusted data, never as executable instructions.

const [baseline, memoryGuided] = await Promise.all([
    incident.memoryMode === "compare" ? diagnose(incident, []) : Promise.resolve(null),
    diagnose(incident, memories),
]);
Enter fullscreen mode Exit fullscreen mode

3. Retaining Verified Resolutions

When the on-call engineer fixes the issue, they submit the confirmed root cause and remediation. Only then does BugFix call hindsight.retain():

async function storeResolution(projectId, incident, resolution) {
    await ensureBank(projectId);

    const content = [
        `Confirmed incident: ${incident.title}`,
        `Service: ${incident.service}`,
        `Environment: ${incident.environment}`,
        `Severity: ${incident.severity}`,
        `Symptoms: ${incident.symptoms}`,
        `Root cause: ${resolution.rootCause}`,
        `Verified fix: ${resolution.fix}`,
        `Outcome: ${resolution.outcome}`,
        resolution.technologies.length
            ? `Technologies: ${resolution.technologies.join(", ")}`
            : "",
    ].filter(Boolean).join("\n");

    await hindsight.retain(bankId(projectId), content, {
        context: `confirmed incident resolution for ${incident.service}`,
        documentId: `incident-${incident.id}`,
        tags: [
            "confirmed-resolution",
            `service:${incident.service.toLowerCase().replace(/[^a-z0-9]+/g, "-")}`
        ],
        timestamp: new Date().toISOString(),
    });
}
Enter fullscreen mode Exit fullscreen mode

Real-World Results: Before vs. After Memory

Here is the exact difference observed in our investigation workflow:

Metric / Step Stateless AI Baseline BugFix with Hindsight Memory
Initial Guess "Check JWT signature algorithm, verify private key format, test token expiration." "Production pods likely mounting old payments-jwt-v3 secret rather than rotated payments-jwt-v4."
Actionable Checks 5 generic verification steps across codebase Specific kubectl secret inspection command
Time to Resolution (MTTR) 25–40 minutes of manual debugging Under 2 minutes
Evidence Trace Black-box output 4 linked facts with exact timestamps & tags

Key Takeaways for Building AI Agents

  1. Don't Let the Model's Guess Become Your Memory: The most dangerous failure mode in agent architectures is saving AI outputs back into the knowledge base automatically. Memory must be verified by humans before retention.
  2. Project Isolation is Non-Negotiable: A shared global vector store causes cross-talk between services. Partitioning memory by project bank keeps recall crisp and prevents noisy context.
  3. Treat Context as Evidence, Not Instructions: Always wrap recalled memories and logs in strict XML or data boundaries to prevent prompt injection and keep diagnostics grounded.
  4. Durable Memory Outperforms Long Context: Rather than feeding thousands of lines of raw logs into context windows, retaining structured resolutions turns ephemeral incidents into compounding operational intelligence.

Links & Resources

Top comments (0)