DEV Community

Battula Aditya
Battula Aditya

Posted on

IncidentMind: Giving AI Agents Persistent Memory for Production Incidents

Production incidents rarely happen in isolation.

A service may fail because of a deployment, a database connection problem, a configuration change, or a combination of several factors. Teams often have encountered similar problems before, but the useful knowledge from those incidents can be difficult to bring into the next investigation.

That led us to a simple question:

What if an incident-response AI agent could remember what the organization had already learned?

We built IncidentMind, an AI-powered incident-response prototype that uses Hindsight as persistent organizational memory. The goal is not simply to generate another incident report, but to connect a current investigation with knowledge retained from previous incidents.

The Problem: Incident Response Starts With Context

When an incident begins, an engineer typically has to assemble context from several places:

What service is affected?

How severe is the incident?

What symptoms are appearing?

Was there a recent deployment?

Has something similar happened before?

What was tried previously?

What actually fixed it?

The last few questions are where organizational memory becomes important.

Without persistent memory, an AI assistant can reason about the information provided in the current conversation, but it does not automatically have access to the lessons learned from previous incidents.

The workflow becomes:

Current incident → Current context → Investigation

With persistent memory, the workflow can become:

Current incident → Relevant historical memory → Investigation

That additional context is what we wanted to explore with IncidentMind.

What We Built

IncidentMind provides an incident investigation workflow around a FastAPI backend.

An engineer provides information such as the incident ID, service, severity, symptoms, and deployment. IncidentMind can then retrieve relevant historical evidence from Hindsight and provide that context to the investigation process.

The application uses:

  • Python
  • FastAPI
  • Hindsight
  • LLM

The investigation service combines the current incident with historical evidence before asking the AI model to produce an investigation report.

The report is structured around:

Incident summary

Root-cause hypotheses

Evidence supporting those hypotheses

Recommended investigation steps

Recommended remediation actions

Important uncertainty or missing evidence

The system is deliberately instructed not to invent facts and to distinguish evidence from hypotheses.


Figure 1. IncidentMind prototype showing the incident context used for an investigation.

Why Persistent Memory Matters

The interesting part of IncidentMind is not simply generating text from an incident description.

The important part is the memory loop.

When an incident is resolved, IncidentMind can retain the post-mortem information in Hindsight. That information includes the symptoms, root cause, resolution, runbook information, and additional notes.

Later, another investigation can use RECALL to search that accumulated organizational knowledge.

This creates a simple learning cycle:

Investigate → Resolve → Retain → Recall → Investigate

The organization does not have to treat every new incident as completely independent.

How Hindsight Fits Into IncidentMind

Hindsight provides the persistent memory layer used by IncidentMind.

The application initializes a Hindsight client using the configured base URL and API key, and uses a memory bank for IncidentMind's incident knowledge.

The two operations we use are RECALL and RETAIN.

from hindsight_client import Hindsight

client = Hindsight(

base_url=os.environ["HINDSIGHT_BASE_URL"],

api_key=os.environ["HINDSIGHT_API_KEY"],
Enter fullscreen mode Exit fullscreen mode

)

BANK_ID = os.environ["HINDSIGHT_BANK_ID"]

def retain_incident(content: str):

return client.retain(

    bank_id=BANK_ID,

    content=content,

)
Enter fullscreen mode Exit fullscreen mode

def recall_incidents(query: str):

return client.recall(

    bank_id=BANK_ID,

    query=query,

)
Enter fullscreen mode Exit fullscreen mode

This separation is useful because the incident-response workflow has two different memory requirements.

During an investigation, the system needs to retrieve relevant knowledge.

After resolution, it needs to preserve what was learned.

Hindsight's Python client provides the RETAIN and RECALL operations used for this workflow. Hindsight GitHub: https://github.com/vectorize-io/hindsight
Hindsight documentation: https://hindsight.vectorize.io/
Agent memory: https://vectorize.io/what-is-agent-memory

A Concrete Incident Example

For the prototype demonstration, we used an incident with the following context:

**Incident: INC-1042

Service: payment-api

Severity: SEV-1

Deployment: v2.5.0

Symptoms: high latency, high error rate, and database connection exhaustion **

This gives IncidentMind enough context to construct a targeted memory query rather than performing a generic search.

The investigation service then combines the current incident information with retrieved organizational memory returned by Hindsight.

The AI model is instructed to reason from that supplied evidence and explicitly identify uncertainty when the evidence is insufficient.


Figure 2. IncidentMind investigation output showing retrieved organizational memory and the resulting investigation report.

This distinction is important: historical evidence is supporting context, not automatically the root cause.

IncidentMind can use retrieved information to formulate hypotheses and recommend investigation steps, but engineers still need to validate those hypotheses against the actual production environment.

The Architecture

At a high level, the workflow looks like this:

Figure 3. IncidentMind architecture connecting incident investigation with Hindsight persistent organizational memory.

An engineer sends the current incident to the IncidentMind API.

The investigation service prepares the incident context and calls Hindsight RECALL to retrieve potentially relevant historical knowledge.

That historical evidence is passed into the investigation process along with the current incident.

LLM generates the investigation report from the supplied context.

Once the incident is resolved, IncidentMind can record the post-mortem and use Hindsight RETAIN to add that knowledge to the organization's persistent memory.

The result is a feedback loop between incident investigation and organizational learning.

What We Learned

Building the prototype highlighted a few practical points.

  1. Memory is useful only when it is relevant

Persistent memory by itself is not enough. The investigation still depends on retrieving information that is actually related to the current incident.

The quality and completeness of the stored incident knowledge therefore matter.

  1. Evidence and reasoning should remain separate

An AI-generated hypothesis should not be presented as a confirmed root cause.

IncidentMind explicitly asks the model to distinguish historical evidence from hypotheses and to state when information is missing.

That makes the memory layer a source of investigation context rather than an authority that decides what happened.

  1. The learning loop is more important than a single response

The useful part of persistent memory appears over time.

A resolved incident becomes organizational knowledge. That knowledge can then become context for a future investigation.

This changes the role of the AI agent from simply answering a question to participating in a longer organizational learning process.

Current Limitations

IncidentMind is currently a prototype rather than a complete production incident-management platform.

It does not yet integrate directly with live observability systems, incident-ticketing platforms, deployment systems, or other production infrastructure.

The quality of investigation results also depends on the relevance and completeness of the incident knowledge available in memory.

The current implementation demonstrates the core memory-enabled investigation workflow: historical evidence retrieval, AI-generated investigation hypotheses and recommendations, and post-mortem retention.

The next step would be connecting that workflow to the systems engineers already use during real production incidents.


Closing

Incident response produces valuable knowledge, but that knowledge is often difficult to reuse when the next incident arrives.

IncidentMind explores a simple alternative: give the incident-response agent persistent organizational memory.

With Hindsight, the workflow becomes more than:

Investigate → Resolve → Forget

It becomes:

Investigate → Resolve → Retain → Recall → Learn

That is the idea behind IncidentMind: not replacing the engineer's judgment, but giving the investigation process access to what the organization has already learned.

GitHub: https://github.com/mallikarjun36/incidentmind

Top comments (0)