I Gave Incident Response a Memory With Hindsight
The first useful question during an outage is often not “what could be wrong?” but “have we seen this before?”
I built Incident-Memory-Copilot around that question. The system combines an incident-response workflow with persistent organizational memory so that a new incident can be investigated using what the organization learned from previous incidents.
What the system does
The application is an incident operations console. It gives an engineer a place to inspect active incidents, search historical memory, review runbooks and postmortems, investigate a new incident, and explicitly teach the system what was learned after resolution.
The important architectural decision is that Hindsight is not treated as a secondary search box. It sits inside the incident lifecycle.
At a high level, the flow is:
DATA SOURCES
│
┌─────────────┼──────────────┐
▼ ▼ ▼
Rootly PagerDuty PagerDuty
Logs Incident Docs Postmortems
│ │ │
└─────────────┼──────────────┘
│
▼
SYNTHETIC ORGANIZATIONAL DATA
│
┌─────────────┼──────────────┐
▼ ▼ ▼
100–150 50–100 50–100
Incidents Runbooks Postmortems
│ │ │
└─────────────┼──────────────┘
▼
HINDSIGHT CLOUD
│
┌──────────┼──────────┐
▼ ▼ ▼
RETAIN RECALL REFLECT
│ │ │
└──────────┼──────────┘
▼
INCIDENT RESPONSE AGENT
│
▼
EVIDENCE-BACKED ACTIONS
│
▼
HUMAN ENGINEER
│
▼
POSTMORTEM
│
└──────→ RETAIN
The data foundation combines operational incident information, incident-response knowledge, postmortems, realistic synthetic incidents, runbooks, and postmortems. Hindsight becomes the layer that turns that accumulated information into reusable memory.
The application then has a simple loop:
Recall → Investigate → Human decision → Resolve → Retain
That loop is more important to me than any individual UI screen.
Why I wanted memory in the incident workflow
An LLM can already explain an HTTP 502, list possible database problems, or suggest checking a deployment.
That is not the difficult part.
The difficult part is knowing what happened in this environment before.
Suppose a payment service is returning HTTP 502 responses. A generic assistant might recommend checking the gateway, application health, upstream dependencies, database connectivity, and recent deployments.
Those are reasonable checks.
But suppose the organization previously had a related incident where a deployment changed database connection-pool limits. The team discovered that restarting the API gateway made the situation worse because it created a traffic spike, while rolling back the deployment restored the correct connection-pool configuration.
That historical experience is much more useful than a generic list of possible causes.
The incident console captures exactly this kind of information.
A concrete incident flow
The current incident screen represents an example around a Payment API returning HTTP 502 errors.
The incident contains operational context such as the service, severity, error, current CPU and memory utilization, recent deployment, and explanatory description.
The investigation then moves through three memory-oriented stages:
CURRENT INCIDENT
│
▼
HINDSIGHT RECALL
│
▼
HISTORICAL INCIDENTS
│
▼
HINDSIGHT REFLECTION
│
▼
RECOMMENDED INVESTIGATION
│
▼
HUMAN REVIEW
The key point is that historical information is presented alongside the current evidence.
The agent is not simply saying:
“This happened before, so do the same thing.”
Instead, the previous incident becomes evidence that an engineer can compare against the current situation.
That distinction matters in production systems because infrastructure changes over time. The same symptom can have a different cause.
Retain, Recall, and Reflect
The integration with Hindsight follows three operations.
Retain
When an incident has been resolved and the engineer knows the actual root cause, what worked, what failed, and what should be prevented can be retained as organizational memory.
Conceptually, the operation looks like:
client.retain(
bank_id=BANK_ID,
content=incident_learning,
context="resolved production incident"
)
The important part is not the API call itself. It is the quality of what gets retained.
A useful memory record should capture the operational lesson rather than merely saying that an incident was closed.
Recall
When a new incident arrives, the current incident becomes the basis for retrieving relevant historical knowledge:
memories = client.recall(
bank_id=BANK_ID,
query=current_incident
)
This is where the system moves beyond a stateless interaction.
The engineer is no longer asking the model to reason only from the current error message. The model can receive historical experiences that are relevant to the current investigation.
Hindsight's recall operation is designed to combine multiple retrieval signals rather than relying only on one simple similarity lookup. The official documentation describes semantic, keyword, graph, and temporal retrieval being combined and reranked. citeturn0search0turn0search2
Reflect
Recall gives the agent relevant memories. Reflection is useful when the question requires a broader synthesis across those memories.
Conceptually:
reflection = client.reflect(
bank_id=BANK_ID,
query="What patterns and failed fixes should I consider?"
)
For incident response, that means the system can move from:
“Here are some similar incidents.”
toward:
“Here is the recurring pattern that appears across those incidents.”
That distinction is why I wanted Hindsight to be central to the design rather than adding a generic vector search layer and calling it memory.
The second architecture: where the data comes from
The data foundation is deliberately broader than one incident table.
DATA FOUNDATION
Rootly Logs
+
PagerDuty Incident Response knowledge
+
PagerDuty Postmortem knowledge
+
100–150 realistic synthetic incident records
+
50–100 realistic runbooks
+
50–100 realistic postmortems
↓
HINDSIGHT CLOUD
↓
INCIDENT RESPONSE AGENT
The purpose is to give the memory layer enough operational context to make historical recall meaningful.
A memory system is only as useful as the information that enters it.
If every retained record says only “incident resolved successfully,” there is very little for a future investigation to learn.
The useful record is closer to:
symptom → investigation → failed action → successful action → root cause → lesson → prevention
That is the shape of knowledge an incident responder can reuse.
The moment that changed the design
The most interesting part of the interface is the resolved-incident screen.
After the incident is resolved, the engineer is asked to capture:
Root cause
What worked
What failed
Lesson learned
Prevention
In the example shown in the application, the root cause is a connection-pool configuration problem introduced in deployment v2.8.0. Rolling back to v2.7.9 restored the connection-pool limits. Restarting the API gateway failed because it caused a traffic spike that made the database problem worse.
The lesson is explicit:
Check connection limits before restarting the gateway during a database-exhaustion event.
The prevention step is also explicit:
Add automated tests that verify connection limits before deployment.
Then there is a Teach Organizational Memory action.
That button represents the part of the design I care about most.
The incident is not finished from the system's perspective when the service recovers. The resolution becomes future context.
Why failed actions belong in memory
One design decision I would keep even if I rebuilt the system from scratch is retaining failed actions.
Incident documentation often emphasizes the final fix.
That is understandable, but it loses valuable information.
Imagine a future engineer sees:
“Rollback deployment v2.8.0.”
That is useful.
But this is more useful:
“Rollback deployment v2.8.0 restored connection-pool limits. Restarting the gateway first caused a traffic spike and made the database exhaustion worse.”
The second record prevents a future engineer from repeating a known mistake.
For incident response, what did not work can be as valuable as what did.
This is also where persistent memory is different from a static runbook. A runbook describes a known procedure. Incident memory can preserve the experience around that procedure.
Making memory visible
I also wanted the system to make memory operations visible to the engineer.
The overview screen exposes memory activity as:
RECALL
8 memories retrieved
REFLECT
Historical pattern synthesized
RETAIN
New learning stored
The dashboard also exposes the relationship between active incidents, historical memory, memory records, and the Hindsight connection.
That visibility matters because an incident recommendation should not feel like an unexplained answer from an LLM.
An engineer should be able to ask:
What historical information influenced this?
Was the memory retrieved successfully?
Did the system synthesize a broader pattern?
What will be remembered after this incident?
The UI is therefore part of the trust model, not just presentation.
Keeping the engineer in control
I deliberately kept a human review step in the incident workflow.
The system can retrieve historical incidents and produce recommended investigation steps, but it does not get to silently decide that a production change should happen.
The workflow is:
Current incident
↓
Historical memory
↓
AI investigation
↓
Evidence-backed recommendation
↓
Human review
↓
Resolution
↓
Postmortem
↓
Organizational memory
That separation is important.
Historical memory can be wrong, incomplete, or no longer applicable. Infrastructure changes. Services are upgraded. Architecture changes. A recommendation based on a six-month-old incident should therefore be treated as evidence, not an instruction that bypasses engineering judgment.
What I learned
- Storage is not memory
Putting incident documents into a database does not automatically give an agent useful memory.
Memory becomes useful when information can be retrieved in the context of a later decision.
- Recall quality depends on what gets retained
If the retained information is vague, future recall will also be vague.
The most useful incident memories contain concrete symptoms, services, actions, outcomes, root causes, and lessons.
- Failed fixes deserve first-class treatment
The final successful remediation is only part of the incident story.
Knowing what made the situation worse can prevent repeated mistakes.
- Historical context should inform, not override
A similar incident is evidence.
It is not proof that the current incident has the same root cause.
That is why the current incident, historical memories, and human review remain separate parts of the workflow.
- The learning loop is the real product
The interesting behavior is not a single successful recall.
It is:
Incident → Memory → Investigation → Resolution → New Memory
Every resolved incident can make the next relevant incident easier to investigate.
Where Hindsight fits
I used Hindsight because its memory model maps naturally onto this workflow.
The official Hindsight GitHub repository describes memory around retain, recall, and reflect, while the Hindsight documentation explains the underlying concepts and APIs. The broader Vectorize explanation of agent memory is also useful for understanding why persistent memory is different from simply keeping more text in an LLM context window.
For this system, the division is straightforward:
Retain the validated lesson.
Recall relevant experience during the next investigation.
Reflect when several memories need to be synthesized into a broader pattern.
That gives the incident agent something a normal one-shot prompt does not have: organizational experience that can persist across incidents.
The part I would measure next
The next engineering question is not whether the dashboard looks convincing.
It is whether memory changes incident investigation in measurable ways.
I would evaluate questions such as:
How often does recall surface a genuinely relevant prior incident?
How often does the historical recommendation match the eventual root cause?
Which failed actions are successfully avoided?
How does recall quality change as the memory bank grows?
Which retained fields contribute most to useful future investigations?
When does reflection provide information that recall alone does not?
Those measurements would tell me whether the memory layer is actually improving the workflow rather than simply adding another component.
Closing the loop
The architecture ultimately comes down to a simple idea:
REMEMBER
↑
│
INCIDENT → INVESTIGATE → RESOLVE
│
↓
LEARN
│
└────────→ REMEMBER
I did not want to build another assistant that gives an answer and forgets it.
I wanted an incident-response system where a resolved outage becomes useful evidence for the next one.
That is what Incident-Memory-Copilot is designed around: turning incident history into organizational memory, and organizational memory into better-informed incident response.
Top comments (0)