DEV Community

Gunja Ramana
Gunja Ramana

Posted on

I Gave Incident Response a Long-Term Memory With Hindsight

Production incidents have an annoying habit of looking new even when a team has already solved something similar before.

I built IncidentMind around a simple idea:

Instead of asking an incident-response agent to reason from a blank context every time, give it persistent memory of what engineers previously observed, decided, fixed, and learned.

The interesting part was not simply adding an AI chat interface. It was making memory part of the incident-response control flow.

The Problem Was Not a Lack of Incident Data

Engineering teams already produce operational knowledge through incidents, post-mortems, runbooks, and resolution notes.

The problem is bringing the right experience into the next incident at the moment it matters.

For example, two payment incidents can both produce 503 errors while having completely different causes. One historical incident might involve database connection-pool exhaustion during a traffic spike. Another might involve an external payment provider becoming slow and causing socket exhaustion.

A useful agent cannot simply see:

payment + 503
Enter fullscreen mode Exit fullscreen mode

and repeat the first fix it finds.

I wanted IncidentMind to answer a more useful question:

What have we learned from previous incidents that is relevant to what is happening now?

That made Hindsight the central memory layer rather than an optional database sitting behind the interface.

The Architecture

IncidentMind uses a React + TypeScript frontend with a Node.js + Express backend.

The backend separates incident handling from memory operations through an IMemoryService abstraction. The Hindsight implementation is handled through HindsightMemoryService, connected to a Hindsight Cloud bank.

The analysis path is intentionally explicit:

Current Incident
      |
      v
Hindsight RECALL
      |
      v
Historical Experiences
      |
      v
Hindsight REFLECT
      |
      v
Groq LLM Reasoning
      |
      v
Human-confirmed Investigation
      |
      v
Resolution / Post-mortem
      |
      v
Hindsight RETAIN
      |
      v
Future Incidents
Enter fullscreen mode Exit fullscreen mode

The responsibilities are separated:

  • Hindsight → long-term organizational memory
  • Reflect → comparison and synthesis of historical experiences
  • Groq LLM → reasoning over the available evidence
  • Engineer → final investigation and production decision

I used the Hindsight Documentation while implementing this lifecycle.

Making Hindsight the Source of Truth

One of the most important changes I made was removing a shortcut from the early prototype.

Initially, local historical data could be used to make sure the interface always displayed an incident match. That was useful during development, but it weakened the actual product behavior.

If the interface displayed a historical incident, I could not always prove that Hindsight had actually recalled it.

So I changed the application path so recalled operational memory comes from Hindsight.

const query = `Service: ${service}. Title: ${title}. Symptoms: ${symptoms}.
Find past engineering incidents with similar service, symptoms,
root causes, decisions, and successful fixes.`;

const recallResult = await this.client!.recall(this.bankId, query, {
  budget: 'mid'
});
Enter fullscreen mode Exit fullscreen mode

If Recall returns no relevant memories, IncidentMind keeps that result empty.

It does not invent a historical match.

Known operational experiences can be ingested into the Hindsight bank through RETAIN. Once stored, those experiences become available to future Recall operations.

That distinction matters because the memory layer should be real, traceable, and useful, rather than decorative.

Recall Is Only the Beginning

Retrieving a similar incident does not automatically mean its resolution should be copied.

IncidentMind therefore uses Hindsight Reflect to compare the current incident with recalled experiences.

const query = `Current incident: "${incident.title}" on service "${incident.service}".
Symptoms: "${incident.symptoms}".
Recalled Historical Memories from Hindsight:
${recalledSummary}

Analyze and compare the current incident against the recalled historical memories:
1. Identify relevant historical patterns and key similarities.
2. Note important differences.
3. Highlight which past resolutions apply and which should not be blindly reused.
4. Recommend prioritized next investigation steps.`;
Enter fullscreen mode Exit fullscreen mode

This creates a useful separation:

Stage Purpose
Recall What have we experienced before?
Reflect How does that experience compare with this incident?
LLM Reasoning What should the engineer investigate next?

This separation is important because similar symptoms do not necessarily mean the same root cause.

The LLM Reasons Over Evidence, Not a Fixed Scenario

After Recall and Reflect, the backend sends the current incident information and recalled organizational memories to Groq.

The prompt explicitly tells the model not to assume a particular service or root cause:

Do NOT assume any specific service, incident ID, or root cause
unless it is explicitly present in the input telemetry or recalled memory.

If no past memories match, explicitly state that historical evidence
is limited and provide diagnostic steps based on available telemetry.
Enter fullscreen mode Exit fullscreen mode

This allows the reasoning layer to work with different incident categories instead of being designed around one fixed example.

For example, the same workflow can be applied to incidents involving:

  • authentication
  • orders
  • notifications
  • web services
  • databases
  • queues
  • previously unseen symptoms

The recommendation remains advisory.

IncidentMind does not silently change production systems. The engineer reviews the available evidence and confirms the appropriate response.

The Learning Loop

The most important part of the architecture happens after an incident is resolved.

When the engineer records information such as the actual root cause, resolution, successful actions, failed actions, runbook information, and outcome, IncidentMind can store that experience in Hindsight.

const retainRes = await this.client!.retain(
  this.bankId,
  structuredContent,
  {
    context: `Post-Mortem for ${incident.title}`,
    tags: [
      incident.service.toLowerCase().replace(/\s+/g, '-'),
      incident.severity.toLowerCase()
    ],
    metadata: {
      incidentId: incident.id,
      service: incident.service,
      severity: incident.severity,
      rootCause: incident.actualRootCause || 'Under investigation'
    }
  }
);
Enter fullscreen mode Exit fullscreen mode

The important behavior is not simply whether RETAIN succeeds.

The intended learning cycle is:

Resolve
   ↓
Retain
   ↓
Future Incident
   ↓
Recall
   ↓
Relevant Historical Experience
Enter fullscreen mode Exit fullscreen mode

This creates a continuous memory loop where previous engineering work can become useful evidence for future incidents.

Before and After

Before Persistent Memory

A new incident could follow a mostly generic troubleshooting path:

New Incident
      ↓
Generic Logs
      ↓
Generic Metrics
      ↓
Recent Deployment Checks
      ↓
Broad Troubleshooting
Enter fullscreen mode Exit fullscreen mode

After Hindsight Became Part of the Workflow

The process becomes:

New Incident
      ↓
Recall Relevant Organizational Experience
      ↓
Reflect on Similarities and Differences
      ↓
Prioritize Investigation
      ↓
Engineer Confirms the Response
      ↓
Resolve
      ↓
Retain the New Experience
Enter fullscreen mode Exit fullscreen mode

The difference is not that IncidentMind automatically knows the answer.

The difference is that the next responder can begin with evidence from previous engineering work.

What I Learned

1. Memory Needs a Clear Contract

I had to decide exactly what belonged in memory.

Storing only incident titles and symptoms would not be enough. Useful operational memory can include root causes, decisions, successful actions, failed actions, and outcomes.

2. Similarity Is Not Applicability

A similar incident is evidence, not proof.

That is why Recall and Reflect have different responsibilities. Historical experience should influence investigation without becoming an automatic prescription.

3. Empty Memory Is a Valid Result

A trustworthy memory system must be able to say:

“We do not have relevant historical evidence.”

That is better than creating a false historical match simply because the interface expects an answer.

4. Performance Has to Be Designed Around Real Memory Calls

Remote Recall, Reflect, and LLM reasoning introduce additional processing time.

Instead of removing these operations, I kept them because they create the intended memory-driven workflow. I then added staged progress feedback and an analysis cache so repeated requests do not unnecessarily redo the same analysis.

5. Human Control Belongs in the Architecture

IncidentMind recommends an investigation path.

It does not silently change production systems.

Historical memory informs the engineer; it does not replace the engineer.

Building Incident Response That Remembers

The main lesson from IncidentMind is that adding memory is not simply about storing previous conversations.

Memory becomes useful when it is connected to the complete operational lifecycle:

Experience
    ↓
RETAIN
    ↓
RECALL
    ↓
REFLECT
    ↓
REASON
    ↓
Human Decision
    ↓
New Experience
Enter fullscreen mode Exit fullscreen mode

That creates a system where every resolved incident has the potential to become useful context for a future one.

Project and References

The complete IncidentMind project is available in the IncidentMind GitHub Repository.

For implementation details and background:

Architecture

IncidentMind Architecture

UI

IncidentMind Screenshot

recall

Hindsight Recall Screenshot

retain

Hindsight Retain Screenshot

Top comments (0)