DEV Community

Mahalaxmi Kouchika
Mahalaxmi Kouchika

Posted on

I Gave My Incident Response Agent a Memory. Here's What Changed.

The first time I asked an AI agent to help debug a production incident, it gave me a perfectly reasonable, perfectly generic checklist: check connectivity, check credentials, check the connection pool.

The problem wasn't that the advice was wrong.

The problem was that an organization may have already solved the same failure before, and a stateless agent has no way to use that experience.

That gap is what I built IncidentMind to address.

What Is IncidentMind?

IncidentMind is an AI-powered Incident Response Agent designed for DevOps and SRE teams.

The core idea is simple:

LLM = THINK

Hindsight = REMEMBER

Instead of asking one model to handle both reasoning and long-term memory, I separated those responsibilities.

The current stack is:

  • Hindsight for persistent organizational memory
  • Groq with openai/gpt-oss-120b for reasoning and final response generation
  • n8n for AI-agent orchestration
  • FastAPI + PostgreSQL for structured incident and operational data
  • React + Vite for the interface engineers interact with

The high-level architecture looks like this:

Engineer
↓
React
↓
n8n Webhook
↓
IncidentMind AI Agent
├── Groq
├── Hindsight
└── Incident Tools
↓
Final Response
↓
React

The interesting part isn't the number of components.

It is what happens when the agent can use organizational memory while investigating a new incident.

The Problem With Stateless Incident Response

Incident response is rarely an isolated event.

When a production system fails, engineers usually need to investigate logs, deployments, configuration changes, dependencies, previous incidents, runbooks, and postmortems.

A language model can reason about the information provided to it.

But reasoning over the current incident is only part of the problem.

An organization also has history.

Maybe the same service failed previously.

Maybe the same dependency caused an outage.

Maybe a configuration change introduced a problem several months ago.

Maybe the team already discovered the correct fix.

That information is extremely valuable during the next incident.

A stateless agent starts again from zero.

IncidentMind is designed to avoid that.

The Core Idea: Retain, Recall, Reflect

The memory layer in IncidentMind follows three important operations:

Retain → Recall → Reflect

These operations have different purposes.

1. Retain

After an incident is investigated and resolved, useful information can be stored for future investigations.

The information can include:

  • What happened
  • What was investigated
  • Which troubleshooting steps were attempted
  • What actually fixed the problem
  • Lessons learned from the incident

The goal isn't to save every message from every conversation.

The goal is to preserve reusable operational knowledge.

An incident should become more than a closed ticket.

It should become experience that can help with the next incident.

[INSERT YOUR ACTUAL HINDSIGHT RETAIN CODE SNIPPET HERE]

2. Recall

When a new incident arrives, IncidentMind can retrieve relevant historical information from Hindsight.

For example:

User:

"We're getting repeated database connection failures."

Instead of immediately generating a generic troubleshooting checklist, the agent can first look for relevant historical experience.

The flow becomes:

Current incident
↓
Hindsight recall
↓
Relevant historical context
↓
Groq reasoning
↓
Final response

This gives the reasoning model information that it could never know from the current prompt alone.

[INSERT YOUR ACTUAL HINDSIGHT RECALL CODE SNIPPET HERE]

3. Reflect

Reflection is different from simply finding one similar incident.

Recall asks:

"Have we seen something relevant before?"

Reflect goes further:

"What patterns can we identify across the incidents we have experienced?"

For example, several incidents might individually look unrelated.

But when considered together, they could reveal a recurring dependency failure, configuration pattern, or deployment-related problem.

That is where persistent organizational memory becomes more interesting than simply searching old incident reports.

Before Memory vs After Memory

The easiest way to understand the difference is to look at the same incident without and with historical context.

Without memory

User:

"We're getting repeated database connection failures."

Agent:

"Check database connectivity, credentials, connection pool configuration, and database logs."

The response is reasonable.

But it is generic.

It could have been produced without knowing anything about the organization's history.

With IncidentMind

User:

"We're getting repeated database connection failures."

IncidentMind first retrieves relevant historical context from Hindsight.

Groq then reasons over the current incident together with that context.

The resulting response can be specific to what the organization has already experienced.

For example, if a previous incident showed that a similar failure followed a connection-pool configuration change, that historical information can influence the investigation.

The difference is not simply that the second response contains more information.

The difference is that it contains information that comes from the organization's own experience.

Why Separate Reasoning From Memory?

One design decision I wanted to make explicit was the separation between reasoning and memory.

Groq is responsible for reasoning and generating the final response.

Hindsight provides persistent memory.

n8n coordinates the agent workflow.

FastAPI and PostgreSQL handle structured operational information.

React provides the interface.

That creates a simple separation of responsibilities:

Current Incident
+
Historical Organizational Memory
↓
Groq
↓
Final Investigation Response

The LLM does not need to remember everything itself.

Instead, the memory layer can retrieve the information that is relevant to the current investigation.

This also gives the architecture a clear boundary between short-term reasoning and long-term organizational knowledge.

Why Hindsight Is Central to the Design

I didn't want memory to be a feature that was added to the agent after the main system was already designed.

The memory layer is part of the investigation workflow itself.

Without memory, the agent primarily reasons from the current incident.

With memory, the investigation can incorporate what the organization has already learned.

That creates a different lifecycle:

Incident
↓
Investigation
↓
Resolution
↓
Retain useful knowledge
↓
Future incident
↓
Recall previous experience
↓
New investigation

The agent can therefore build on previous operational experience instead of treating every incident as a completely new problem.

The Architecture in Practice

The workflow starts when an engineer submits an incident through the React interface.

The request reaches the n8n workflow through a webhook.

n8n orchestrates the agent workflow and connects the different components.

Hindsight provides the persistent memory layer.

Relevant historical information can be recalled and passed into the reasoning process.

Groq then generates the response using the current incident information together with the available historical context.

FastAPI and PostgreSQL provide the structured operational data layer.

Finally, the response is returned to the React interface.

This gives IncidentMind a clear separation:

Frontend → interaction

n8n → orchestration

Hindsight → memory

Groq → reasoning

FastAPI + PostgreSQL → structured data

What I Learned

1. Memory only matters if it changes the answer

Adding a memory system doesn't automatically make an agent useful.

The important question is whether the information retrieved from memory actually affects the investigation.

If the response would be exactly the same with or without memory, the memory layer isn't providing much value.

2. More memory isn't automatically better memory

An incident response agent doesn't need every previous conversation.

It needs the relevant information for the problem being investigated.

That makes retrieval quality important.

The goal is not to give the model everything.

The goal is to give it the right context.

3. Organizational knowledge is different from general knowledge

An LLM already knows generic troubleshooting techniques.

It can explain database connection errors, deployment failures, service outages, and many other technical problems.

But it doesn't automatically know how a specific organization solved an incident six months ago.

That knowledge has to come from somewhere.

For IncidentMind, that source is persistent organizational memory.

4. Reasoning and memory can have separate responsibilities

Groq reasons.

Hindsight remembers.

n8n orchestrates.

FastAPI and PostgreSQL handle structured operational data.

React provides the interface.

Giving each component a clear responsibility makes the overall architecture easier to reason about.

5. Repeated incidents are the real test

One good response isn't enough to prove that memory is useful.

The more interesting test is what happens when a related incident appears again.

The intended cycle is:

Incident
↓
Resolution
↓
Memory
↓
New Incident
↓
Recall
↓
Context-aware Investigation

That is where a memory-first agent can provide value over time.

Where This Goes Next

The long-term goal isn't simply to build an agent that answers incident questions faster.

It is to build an agent that becomes more useful as organizational experience accumulates.

Every resolved incident can potentially become useful context for a future investigation.

Instead of repeatedly asking:

"What should I check?"

the system can move toward:

"What do we already know about this?"

That is the direction behind IncidentMind.

The LLM handles the thinking.

Hindsight provides the remembering.

n8n connects the workflow.

The structured data layer preserves operational information.

And the result is an incident response system designed to use both the evidence from the incident happening now and the experience accumulated from incidents that happened before.

Resources

If you want to explore the memory layer used in IncidentMind:

Hindsight GitHub:
https://github.com/vectorize-io/hindsight

Hindsight Documentation:
https://hindsight.vectorize.io/

Vectorize Guide to Agent Memory:
https://vectorize.io/what-is-agent-memory

Source Code

IncidentMind:
[https://github.com/MahalaxmiKouchika/Incident-Mind]

The goal behind IncidentMind is simple:

Build an incident response agent that doesn't just answer the incident happening now, but can use what the organization learned from the incidents that happened before.

Top comments (0)