Introduction
Production incidents are rarely completely new.
A similar combination of high latency, resource exhaustion, configuration changes, or deployment issues may have happened before. However, a typical AI assistant analyzing an incident does not automatically remember what happened during previous incidents.
We built Incident Memory Agent to address this problem.
It is an AI-powered incident-response assistant that uses Hindsight persistent memory to recall previous production incidents, analyze new incidents using historical context, retain new experiences, and reflect on recurring patterns.
The central idea is:
AI should have not only intelligence, but also experience.
The Problem
A stateless incident-response assistant can analyze the information provided in the current incident, but it may not have access to the organization's previous incident experience.
That means every incident can effectively become a new problem.
For example, imagine a production incident with:
- API latency increasing to 5.1 seconds
- Redis memory reaching 94%
- A deployment occurring shortly before the incident
If the organization previously experienced a similar incident and discovered that a cache invalidation bug caused the problem, that historical experience can be extremely valuable.
The challenge is making that experience available to the AI when the next incident occurs.
Our Solution
We built Incident Memory Agent, an AI incident-response system with persistent memory.
The system follows a continuous learning loop:
Recall → Analyze → Retain → Reflect
1. Recall
When a new incident is submitted, the agent first queries Hindsight for relevant historical memories.
These memories can include:
- Previous incident symptoms
- Recent deployments
- Root causes
- Successful resolutions
- Incident outcomes
- Observed patterns
2. Analyze
The AI then analyzes the current incident together with the historical context retrieved from Hindsight.
This allows the model to reason from both:
Current incident + Previous experience
rather than analyzing the incident in isolation.
3. Retain
After analyzing the incident, the new incident and its outcome are stored back into Hindsight.
This means the current incident becomes part of the agent's future experience.
4. Reflect
Finally, Hindsight's reflection capability is used to identify higher-level patterns from the accumulated experience.
This helps move from individual incident memories toward reusable operational knowledge.
Example
Consider a previous incident, INC-001.
The system remembered that:
- Production API latency increased to approximately 5.2 seconds.
- Redis memory usage reached approximately 93%.
- The incident followed deployment v2.4.1.
- Investigation identified a cache invalidation bug.
- Rolling back the deployment restored normal API latency.
Later, a new incident, INC-002, occurs:
- API latency increases to 5.1 seconds.
- Redis memory reaches 94%.
- Deployment v2.4.2 occurred shortly before the incident.
The Incident Memory Agent recalls the previous experience from Hindsight.
Instead of starting from zero, the agent can use the historical evidence to identify the similarity and surface the previously successful resolution.
This is the behavior we wanted to demonstrate:
The agent learns from what happened before.
How Hindsight Is Used
Hindsight is the core memory layer of our application.
Our application uses Hindsight for three important operations:
RETAIN
Stores incident experiences and outcomes in persistent memory.
RECALL
Retrieves relevant historical memories when a new incident is analyzed.
REFLECT
Synthesizes recurring patterns from accumulated experiences.
This creates a feedback loop:
text
Current Incident
↓
RECALL
↓
Historical Experience
↓
AI Analysis
↓
RETAIN
↓
New Experience
↓
REFLECT
↓
Learned Pattern
Top comments (1)