An incident post-mortem is only useful if someone can find and apply it during the next outage.
I built IncidentMind so that resolving an incident does not end the workflow. Instead, it creates memory that can become useful during the next incident.
The key operation is Hindsight RETAIN.
Instead of treating memory as a collection of documents, I treat a resolved incident as an experience:
- What happened?
- What did engineers try?
- What worked?
- What failed?
- Why did the final resolution work?
- What was the outcome?
That became the learning loop at the center of the system.
π The Old Lifecycle Ended Too Early
A conventional incident workflow looks roughly like this:
Incident
β
Investigation
β
Root Cause
β
Resolution
β
Post-mortem
The problem is what happens after the post-mortem.
The document gets stored somewhere. Months later, another incident occurs. Someone searches for it, remembers part of it, or starts investigating from scratch.
I wanted the lifecycle to continue:
Incident
β
Investigation
β
Resolution
β
Retain Experience
β
Future Incident
β
Recall Experience
β
Better Investigation
The important part is that the memory is created from the resolution itself.
π§ What I Retain
I did not want to store only the incident title and a few keywords.
The retained record contains operational context:
const structuredContent = `
INCIDENT DECISION MEMORY:
Incident ID: ${incident.id}
Title: ${incident.title}
Service: ${incident.service}
Severity: ${incident.severity}
Environment: ${incident.environment}
Symptoms: ${incident.symptoms}
Root Cause: ${incident.actualRootCause || 'Under investigation'}
Resolution Steps: ${incident.resolutionSteps || 'Applied recovery'}
What Worked: ${incident.whatWorked || 'Applied mitigation'}
What Failed: ${incident.whatFailed || 'None'}
Runbook Used: ${incident.runbookUsed || 'Standard Triage'}
Resolution Time: ${incident.resolutionTimeMinutes || 10} minutes
Outcome: Resolved successfully
`.trim();
await this.client!.retain(this.bankId, structuredContent, {
context: `Post-Mortem for ${incident.title}`,
tags: [
incident.service.toLowerCase().replace(/\s+/g, '-'),
incident.severity.toLowerCase()
],
metadata: {
incidentId: incident.id,
service: incident.service,
severity: incident.severity,
rootCause: incident.actualRootCause || 'Under investigation'
}
});
There are two important ideas here.
First, the content is written for future retrieval, not merely for today's UI.
Second, metadata preserves identity. The incident ID and service make it possible to understand where the memory came from without relying on an opaque generated response.
β Why βWhat Failedβ Belongs in Memory
It is tempting to retain only successful actions.
I think that loses valuable operational knowledge.
Suppose an engineer tries restarting a service, but the restart has no effect. They then discover a saturated database connection pool and fix the actual problem.
A future responder benefits from knowing both facts.
βRestart did not address the causeβ is valuable operational knowledge.
That is why the incident decision record includes:
- What worked
- What failed
- Root cause
- Resolution
- Runbook
- Outcome
The goal is not to create perfect documentation.
The goal is to preserve the decisions that another engineer would otherwise have to rediscover.
π Hindsight Becomes the Organizational Memory
The application uses a Hindsight Cloud bank for these experiences.
The workflow is:
Resolve Incident
β
Build Structured Experience
β
Hindsight RETAIN
β
Memory Bank
β
Future Incident
β
Hindsight RECALL
On the next incident, Recall uses the current service, title, and symptoms to search the Hindsight bank rather than querying a fixed incident list:
const query =
`Service: ${service}. Title: ${title}. Symptoms: ${symptoms}.
Find past engineering incidents with similar service, symptoms,
root causes, decisions, and successful fixes.`;
const recallResult = await this.client!.recall(
this.bankId,
query,
{ budget: 'mid' }
);
That makes Hindsight part of the operational lifecycle instead of an isolated AI feature.
The Hindsight Documentation describes persistent agent memory through operations such as retaining and recalling experience. I apply that idea to incident response by making the unit of memory an engineering decision record.
π§ͺ The Strongest Test Was a Future Incident
The useful test was not simply checking that RETAIN returned successfully.
A stronger validation path is:
- Resolve an incident.
- Retain its experience.
- Create a later incident with related symptoms.
- Check whether Recall can retrieve the earlier experience.
The intended behavior is:
Incident A
β
Resolved
β
Retained in Hindsight
β
Incident B
β
Recall
β
Relevant Experience
This is important because successful storage alone does not prove that the system has created useful future context.
The real test of memory is whether future behavior can use what was stored previously.
β±οΈ Retention Has to Happen at the Right Boundary
I also learned that RETAIN should not happen on every interaction.
If every analysis request created a memory, the memory bank could quickly fill with:
- transient reasoning
- repeated observations
- incomplete conclusions
- temporary investigation states
The useful boundary is resolution.
During investigation, the system is still uncertain.
After resolution, the incident contains an outcome that another responder can learn from.
That gives the memory lifecycle a clean semantic boundary:
Investigating = Evidence
Resolved = Experience
This distinction helps keep the memory corpus focused.
π€ Hindsight Is Not the Agent
Another architectural decision was keeping Hindsight separate from the reasoning model.
The responsibilities are different:
| Component | Responsibility |
|---|---|
| Hindsight | Persistent organizational experience |
| LLM | Reasoning and structured recommendations |
| IncidentMind | Connects memory, reasoning, and incident workflow |
| Engineer | Final investigation and production decision |
That separation matters because the model should not be responsible for remembering the organization.
A model invocation can reason about the context it receives. It should not be treated as the durable source of truth for what happened weeks or months ago.
This is why I describe IncidentMind as an agent with memory, rather than simply a memory system that happens to generate text.
The Vectorize explanation of agent memory helped frame this distinction: memory gives an agent continuity across interactions instead of forcing every interaction to begin with an empty context.
π‘οΈ Memory Has to Be Honest
One of the easiest mistakes in a memory-backed application is pretending that memory exists when it does not.
I removed a fallback that could inject predefined historical incidents after Hindsight returned no relevant matches.
The new behavior is simpler:
Hindsight returns relevant memories
β
Use them as historical evidence
or:
Hindsight returns no relevant memories
β
Report no direct historical evidence
β
Reason from current evidence
That makes testing more meaningful too.
If a completely new incident produces no historical match, that is a valid result.
It tells me the agent is reasoning from current evidence rather than inventing a history.
βοΈ Before and After
Before Persistent Retention
Resolve Incident
β
Write Post-mortem
β
Future Incident Starts Again
With Hindsight in the Lifecycle
Resolve Incident
β
Retain Structured Experience
β
Future Incident
β
Recall Previous Experience
β
Investigate With Historical Context
The second flow gives the next incident a chance to benefit from the previous one.
π What I Learned
1. Retain Decisions, Not Conversations
A long conversation contains plenty of information that will never help another responder.
A compact incident decision record is easier to retrieve, interpret, and reuse.
2. Resolution Is the Natural Memory Boundary
An unresolved incident is still changing.
A resolved incident has an outcome.
That makes resolution a natural point to create durable organizational experience.
3. Learning Requires a Future Retrieval Test
A successful RETAIN call proves storage.
It does not prove learning.
The stronger validation is checking whether a later incident can retrieve relevant information from an earlier resolved incident.
4. Failed Actions Are Operational Knowledge
What did not work can be as useful as what did.
Future responders should inherit both successful and unsuccessful investigation paths.
5. Memory Should Not Manufacture Certainty
If no relevant historical experience exists, the system should say so.
An honest empty result is better than a fabricated historical match.
π Project and References
The complete IncidentMind project is available in the IncidentMind GitHub Repository.
For implementation details and background:
- Hindsight GitHub Repository
- Hindsight Documentation
- Vectorize β What Is Agent Memory? ### Architecture




Top comments (0)