How persistent operational memory can help an AI incident-response agent learn from previous incidents instead of starting from scratch every time.
Production incidents are rarely completely new.
A Payment API may start returning 503 errors because of a Redis connection problem today, and a very similar incident may happen again weeks or months later.
An engineer who handled the previous incident might remember what worked.
But what happens when that engineer isn't available?
What happens when the knowledge is buried inside an old incident ticket, postmortem, dashboard, or someone's memory?
This is where the idea behind IncidentBrain began.
IncidentBrain is an AI-powered SRE incident-response agent designed around one simple idea:
An AI agent should be able to learn from previous production incidents and use that experience when investigating future incidents.
Instead of treating every incident as a completely new problem, IncidentBrain maintains persistent operational memory using Hindsight.
When an incident is resolved, the system remembers three things:
What happened
What action was taken
What the observed outcome was
Later, when a similar incident occurs, IncidentBrain can retrieve that previous experience and provide it as historical context during the investigation.
This creates a continuous learning loop:
Investigate → Resolve → Remember → Recall → Investigate Again
The goal isn't to replace the SRE.
The goal is to make the SRE's previous experience easier to retrieve when the next incident happens.
The Problem: Stateless AI Isn't Enough
A conventional AI assistant can analyze an incident report and suggest possible causes or remediation steps.
However, each interaction can effectively start from scratch unless relevant historical information is explicitly provided.
For an SRE, this creates a gap between AI reasoning and organizational experience.
Consider a simple scenario.
The Payment API is returning HTTP 503 errors, and Redis connections are approaching the configured connection limit.
An engineer may have encountered this exact pattern before.
Perhaps a previous incident revealed that Redis connection-pool exhaustion was responsible, and increasing the pool resolved the problem.
A stateless assistant doesn't automatically know that.
This is the problem IncidentBrain is designed to address.
Instead of treating every incident as an isolated conversation, IncidentBrain gives the agent a way to retrieve relevant experiences from previous incidents.
The goal is not simply to generate another AI recommendation.
The goal is to give the recommendation access to operational experience from what happened before.
The Core Idea: Persistent Operational Memory
IncidentBrain uses Hindsight as its persistent memory layer.
When an incident is resolved, the agent stores an experience containing three important pieces of information:
- What happened
- What action was taken
- What the observed outcome was
For example, an incident experience might look like this:
Incident: Payment API returned HTTP 503 errors.
Action: The Redis connection pool was increased from 100 to 250.
Outcome: The HTTP 503 error rate returned to normal.
This isn't simply a conversation transcript.
It becomes operational experience that can be retrieved when a future incident resembles the previous one.
The workflow is simple:
Incident → Remember → Future Incident → Recall Previous Experience → Investigate
This creates an important feedback loop.
The agent doesn't just generate an answer and forget it.
The outcome of an incident becomes part of its future context.
From Incident Investigation to Incident Learning
IncidentBrain's workflow has two major stages:
- Investigate
- Resolve and Learn
1. Investigate
When a new incident is submitted, IncidentBrain first searches Hindsight for relevant previous experiences.
For example, the agent can recall previous incident knowledge using:
memory_result = self.memory.recall(
bank_id=self.bank_id,
query=incident
)
memories = []
for result in memory_result.results:
memories.append(result.text)
The retrieved experiences are then provided to the LLM as historical context for the current investigation.
The current incident and the retrieved memories are combined before the model generates its recommendation.
The model is instructed to distinguish between four types of information:
- Current observations
- Historical evidence
- Recommendations
- Unknown information This distinction is important in production environments. A historical configuration value should not automatically be treated as the current configuration. For example, suppose a previous incident involved a Redis connection pool of 100. That value may no longer be valid when the next incident occurs. Instead, the historical value can be treated as evidence while the engineer verifies the current system state before making a change. This helps IncidentBrain use previous experience without blindly applying it to the current incident.
2. Resolve and Learn
Investigating an incident is only half of the learning process.
After the engineer investigates and resolves the incident, IncidentBrain provides a way to record two important pieces of information:
- The action that was taken
- The resulting outcome
For example, the agent can store the resolved incident in Hindsight:
experience = f"""
HISTORICAL SRE INCIDENT EXPERIENCE
Incident:
{incident.strip()}
Action that was taken:
{action_taken.strip()}
Observed outcome:
{outcome.strip()}
"""
result = self.memory.retain(
bank_id=self.bank_id,
content=experience
)
The incident, action, and observed outcome are now stored as a reusable experience.
This turns the resolution of one incident into operational knowledge that can be retrieved during a future investigation.
The next time a similar incident occurs, IncidentBrain can recall this experience from Hindsight and provide it as historical context.
This creates the complete learning cycle:
IncidentBrain Learning Loop
The complete workflow can be visualized as:
Incident → Recall → Recommend → Resolve → Learn

Investigate → Resolve → Remember → Recall → Investigate Again
The important part is that this learning process does not require retraining the underlying language model.
Instead, the agent improves its future context by accumulating relevant operational experiences.
A Simple Example
Let's walk through a realistic incident to see how IncidentBrain's memory loop works.
First Incident
Imagine that the Payment API starts returning HTTP 503 errors.
At the same time, Redis connections have increased significantly and are approaching the configured connection limit.
The engineering team investigates the incident and identifies Redis connection-pool exhaustion as the likely cause.
They increase the Redis connection pool from 100 to 250.
After the change, the HTTP 503 error rate returns to normal.
IncidentBrain then stores the experience:
Incident: Payment API returned HTTP 503 errors because Redis connections were approaching the configured limit.
Action: Increased the Redis connection pool from
100to250.Outcome: The HTTP 503 error rate returned to normal.
This experience is retained in Hindsight.
A Similar Incident Happens Again
Now imagine that weeks later, the Payment API starts returning HTTP 503 errors again.
Redis connections are once again approaching the configured connection limit.
The engineer doesn't need to manually search through old incident reports to find the previous resolution.
IncidentBrain can recall the previous experience from Hindsight.
Incident Investigation with Historical Memory
When a similar incident occurs, IncidentBrain retrieves relevant previous incident experience from Hindsight instead of starting with an empty context.
The retrieved experience becomes historical evidence for the new investigation.
The LLM can then consider:
- What is happening now?
- What happened during the previous incident?
- What action worked previously?
- Is the current situation similar enough to consider the previous resolution?
- What information still needs to be verified? ### AI Recommendation
The retrieved historical experience is used as evidence while the LLM analyzes the current incident and generates a recommendation.
This is the key difference between the first and second investigation.
The first investigation creates operational knowledge.
The second investigation can benefit from that knowledge.
The agent is no longer starting with an empty context.
Architecture
IncidentBrain uses a relatively focused architecture built around four main components:
- React — provides the incident investigation dashboard.
- FastAPI — exposes the backend APIs for investigation and resolution.
- IncidentBrain Agent — coordinates memory retrieval, LLM reasoning, and storing incident outcomes.
- Hindsight + Groq — Hindsight provides persistent operational memory, while Groq provides the LLM inference layer.
The overall flow looks like this:
React Dashboard
│
▼
FastAPI Backend
│
▼
IncidentBrain Agent
│
├──────────────► Hindsight Memory
│ │
│ ▼
│ Previous Incidents
│
▼
Groq LLM
│
▼
AI Recommendation
│
▼
Engineer Resolution
│
▼
Action + Outcome
│
└──────────────► Hindsight Memory
The React frontend provides the interface where an engineer can submit and investigate an incident.
The FastAPI backend exposes the APIs that connect the frontend with the IncidentBrain agent.
The agent acts as the coordinator. It retrieves relevant historical experiences from Hindsight, combines them with the current incident, sends the context to the LLM, and produces a recommendation.
After the incident is resolved, the agent stores the action and observed outcome back into Hindsight.
This creates the feedback loop that allows future investigations to benefit from previous incidents.
System Architecture
The system connects the incident dashboard, backend API, IncidentBrain agent, persistent Hindsight memory, and Groq LLM into a continuous incident-learning workflow.
Connecting IncidentBrain to the LLM
Once IncidentBrain retrieves relevant historical experiences from Hindsight, the current incident and that historical context are passed to the LLM.
In the current implementation, Groq is used as the LLM inference layer.
A simplified version of the integration looks like this:
response = self.groq.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[
{
"role": "system",
"content": "You are an experienced SRE incident response assistant."
},
{
"role": "user",
"content": prompt
}
]
)
recommendation = response.choices[0].message.content
The important part isn't simply calling an LLM.
The important part is what the model receives as context.
The prompt contains the current incident together with relevant historical experiences retrieved from Hindsight.
This allows the model to reason about the current situation while also considering what happened during similar incidents in the past.
In other words:
Current Incident + Historical Experience → LLM Reasoning → Recommendation
Hindsight provides the persistent memory, while Groq provides the LLM inference layer.
The two components therefore play different roles in the system:
- Hindsight: remembers previous operational experiences.
- Groq: reasons over the current incident and retrieved context. ## Evidence-Driven Recommendations
There is an important design consideration when using historical memory for production operations.
Past experience should be treated as evidence, not absolute truth.
A previous incident may have occurred under different:
- Traffic levels
- Infrastructure configurations
- Software versions
- Resource constraints
Because of this, IncidentBrain is designed to avoid blindly applying historical information to the current incident.
The agent is instructed to:
- Avoid inventing metrics or configuration values.
- Clearly distinguish historical facts from current observations.
- Verify the current configuration before making changes.
- Identify risks and uncertainty.
- Request additional investigation when the available evidence is insufficient.
For example, suppose a previous incident was resolved by increasing a Redis connection pool from 100 to 250.
That does not automatically mean the same change should be applied to every future Redis-related incident.
The current system may have a different configuration, traffic pattern, or underlying problem.
Instead, the previous resolution becomes a piece of historical evidence that the engineer can consider while investigating the current state.
This approach makes persistent memory useful without treating it as an unquestionable source of truth.
Why Persistent Memory Changes the Role of an AI Agent
The interesting part of IncidentBrain isn't simply that an LLM can analyze a Redis incident.
Modern LLMs can already explain technical problems and suggest possible solutions.
The more interesting capability is experience accumulation.
When an incident is resolved, the experience can be stored in persistent memory.
Later, when a similar incident occurs, that previous experience can be retrieved and used as context.
This creates a continuous learning cycle:
Incident → Resolution → Memory → Recall → Better Investigation
For example:
- An engineer investigates a production incident.
- A resolution is applied.
- The observed outcome is recorded.
- The experience is stored in Hindsight.
- A similar incident occurs later.
- IncidentBrain retrieves the previous experience.
- The LLM uses that experience while reasoning about the new incident.
As more relevant experiences are stored, more organizational knowledge becomes available to the agent.
This opens the possibility of AI systems that don't just answer questions, but become more useful over time because they can remember what happened before.
What's Next
IncidentBrain currently demonstrates the core idea of combining AI reasoning with persistent operational memory.
The architecture can be extended further by connecting the agent to real production systems.
Possible future improvements include:
- Monitoring systems — Automatically collect current metrics and alerts.
- Incident management systems — Connect incidents directly to the agent.
- Logs and traces — Use application logs and distributed traces as investigation evidence.
- Kubernetes — Allow the agent to reason about workloads, pods, deployments, and cluster events.
- Deployment systems — Correlate incidents with recent deployments and configuration changes.
- Richer incident summaries — Automatically generate structured incident reports.
- Similarity detection — Identify incidents that resemble previous production problems.
- Automated verification — Verify whether a recommended action actually improved the system.
- Team knowledge — Store operational knowledge that can be reused across engineers and incidents.
The long-term goal is to move from a standalone incident-response assistant toward an AI system that can participate more deeply in the production operations workflow.
Conclusion
IncidentBrain explores a simple idea:
AI incident response becomes more useful when the system can remember previous operational experiences.
The core workflow is:
Investigate → Resolve → Remember → Recall → Investigate Again
Instead of treating every production incident as an isolated problem, IncidentBrain creates a feedback loop where previous incidents can become useful evidence for future investigations.
Hindsight provides the persistent memory layer, while the LLM provides reasoning over the current incident and retrieved historical context.
The goal is not to replace SRE engineers.
The goal is to make previous operational experience easier to retrieve and use when the next incident happens.
That combination of AI reasoning + persistent operational memory is the foundation of IncidentBrain.



Top comments (0)