Finding a similar incident is easy. Deciding whether that incident is actually relevant to the current outage is the harder engineering problem.
When I built IncidentMind, I initially thought the most important operation would be memory retrieval: give Hindsight the current incident, retrieve similar incidents, and pass them to the model.
That worked as a starting point, but it exposed a more important problem.
Two incidents can look similar at the symptom level and still require completely different responses.
I designed the analysis path around that distinction:
Hindsight Recall → Hindsight Reflect → LLM Reasoning
Recall finds historical experience. Reflect compares that experience with the current incident. The language model then uses the comparison to produce a structured investigation path.
That separation became one of the most important design decisions in the system.
🔍 Similar Symptoms Do Not Mean the Same Cause
Consider two payment incidents.
The first produced 503 errors during a traffic spike because database connections were exhausted. The resolution involved increasing the database pool and restarting affected nodes.
The second also produced failures in the payment API, but the root cause was upstream payment-provider latency. The appropriate response was related to the external provider and circuit-breaking behavior rather than database capacity.
A similarity search can reasonably retrieve both.
That is useful, but it is not enough.
If an incident agent simply copied the resolution from the nearest historical incident, memory would become a source of dangerous confidence.
I wanted historical memory to be evidence for reasoning rather than a replacement for reasoning.
🧠 The Three-Stage Analysis Path
The analysis pipeline is deliberately explicit:
Current Incident
|
v
Hindsight RECALL
|
v
Historical Experiences
|
v
Hindsight REFLECT
|
v
Comparison & Differences
|
v
LLM Reasoning
|
v
Investigation Recommendation
The responsibilities are separated:
| Stage | Responsibility |
|---|---|
| RECALL | Retrieve relevant organizational experiences |
| REFLECT | Compare historical experiences with the current incident |
| LLM Reasoning | Turn available evidence into an investigation path |
| Engineer | Evaluate the recommendation and make the final decision |
This separation also makes the system easier to debug.
If a recommendation looks wrong, I can ask three separate questions:
- Did Recall retrieve the right experiences?
- Did Reflect distinguish the important differences?
- Did the final reasoning use that evidence correctly?
Without those boundaries, every bad recommendation would simply look like:
“The AI got it wrong.”
🔎 Recall Is Intentionally Broad
The Recall query uses the current incident's actual context:
const query = `Service: ${service}. Title: ${title}. Symptoms: ${symptoms}.
Find past engineering incidents with similar service, symptoms,
root causes, decisions, and successful fixes.`;
const recallResult = await this.client!.recall(this.bankId, query, {
budget: 'mid'
});
I want Recall to retrieve relevant experiences, not make the final decision.
The result can contain several historical incidents:
- strong matches
- partial matches
- useful contrasts
That last category matters.
A historical incident does not have to be a direct match to be useful. Sometimes the most valuable memory is the one that looks similar initially but has a different root cause.
🧩 Reflect Turns Retrieval Into Comparison
The Reflect step asks Hindsight to reason over the current incident and recalled memories:
const query = `Current incident: "${incident.title}" on service "${incident.service}".
Symptoms: "${incident.symptoms}".
Recalled Historical Memories from Hindsight:
${recalledSummary}
Analyze and compare the current incident against the recalled historical memories:
1. Identify relevant historical patterns and key similarities.
2. Note important differences between past incidents and current symptoms.
3. Highlight which past resolutions apply and which should not be blindly reused.
4. Recommend prioritized next investigation steps based on organizational memory.`;
The instruction not to blindly reuse a historical resolution is intentional.
A memory system should preserve experience without turning experience into a rigid runbook.
The result is a useful middle layer between raw retrieval and final generation.
🤖 The Language Model Gets Evidence, Not Just a Prompt
After Recall and Reflect, the current telemetry, historical matches, and reflection are provided to the reasoning layer.
That gives the model a much richer input than a generic question such as:
“What is wrong with this incident?”
The model can reason over:
- current symptoms
- current service and environment
- severity
- historical root causes
- previous resolutions
- differences between incidents
- recommended investigation areas
The output is then represented as structured incident analysis rather than simply generating a free-form chatbot response.
That matters during an outage.
An engineer needs to know what to investigate next, not just read a convincing paragraph.
🧪 A Concrete Example
Suppose a new production incident reports:
Service: Payment API
Severity: SEV-1
Symptoms:
- 503 responses increasing
- latency elevated
- database pool at 100%
Recall can surface a previous payment incident involving connection-pool exhaustion.
It may also surface another payment incident involving an external provider.
Reflect then has an important job:
Compare the evidence rather than choosing the first similar incident.
If the current telemetry shows database pool saturation and no corresponding upstream-provider signal, the database-related incident becomes more relevant.
If the database looks healthy but upstream calls are timing out, the historical provider incident becomes more useful.
The system therefore does not need to pretend that one historical memory is always correct.
🛡️ This Also Improves Failure Behavior
One of my strongest requirements became:
An empty Recall result must remain empty.
If Hindsight has no relevant historical experience, IncidentMind should say so.
I deliberately removed fallback behavior that silently injected local historical incidents whenever Recall returned no matches.
That shortcut made the interface look more intelligent, but it undermined the central claim of the system.
Now Hindsight is the source of truth for recalled operational memory.
The application can therefore represent two honest states:
Relevant Memory Found
|
v
Compare + Reason
or:
No Relevant Memory Found
|
v
Reason From Current Evidence
|
v
Report Limited Historical Context
This is more useful than fabricating certainty.
⚖️ Before vs. After
Before the Separation
Current Incident
↓
Find Similar Incident
↓
Reuse Its Fix
After Separating Recall, Reflect, and Reasoning
Current Incident
↓
Recall Historical Experience
↓
Reflect on Similarities & Differences
↓
Reason From Current + Historical Evidence
↓
Recommend Investigation
The diagram is only slightly different.
The behavior is significantly different.
Historical similarity becomes evidence instead of an automatic answer.
📚 What I Learned
1. Similarity Is a Retrieval Mechanism, Not a Decision Mechanism
The nearest historical incident is not automatically the right incident.
Retrieval should narrow the search space; reasoning still has to determine applicability.
2. Contradictory Memories Are Useful
Two incidents with similar symptoms but different root causes can teach the system what evidence separates them.
A memory system should preserve those differences rather than flattening everything into:
“Similar incident.”
3. Reflection Deserves Its Own Boundary
Putting comparison into a separate step made the architecture easier to understand and debug.
It also made Hindsight more than simply a storage layer.
4. Empty Memory Is a Valid Result
A system that admits it has no relevant history is more trustworthy than one that always produces a historical match.
5. Human Judgment Remains Important
IncidentMind recommends investigation paths.
It does not automatically change production systems.
Historical experience informs the responder; it does not replace the responder.
🚀 Project and References
The complete IncidentMind project is available in the IncidentMind GitHub Repository.
For implementation details and background:
- Hindsight GitHub Repository
- Hindsight Documentation
- Vectorize — What Is Agent Memory? ### Architecture




Top comments (0)