The easiest way to build an incident agent is to make it excellent at one demo scenario.
The harder problem is making it useful when the next incident looks nothing like the previous one.
I designed IncidentMind around incident scenarios across payment, authentication, orders, notifications, web services, databases, queues, and infrastructure.
That forced me to remove a class of shortcuts that made early versions look good but would not survive outside one scenario.
The central design goal became simple:
Hindsight should provide the memory, while the reasoning path should work from whatever incident the engineer actually reports.
β οΈ The Danger of Scenario-Specific Intelligence
Payment failures were useful while developing the system because they provided concrete examples.
But payment incidents also made it easy to accidentally encode assumptions such as:
```text id="z4v3uk"
Payment API
β database pool
β external payment provider
That is not an incident-response system.
It is a **payment troubleshooting workflow**.
A real incident responder may receive:
* Authentication failures
* Order API latency
* Notification queue backlog
* Kubernetes resource exhaustion
* Certificate failures
* DNS problems
* Database saturation
* Third-party API timeouts
* A completely new failure mode
The architecture therefore has to treat the **incident itself as data**.
---
## π§© The Memory Interface Is Deliberately Generic
The backend isolates memory operations behind a service abstraction.
The important operations are conceptually:
```ts id="r9u4kv"
interface IMemoryService {
recallSimilarIncidents(...): Promise<...>;
retrieveRelevantMemories(...): Promise<...>;
storeIncidentExperience(...): Promise<...>;
getIncidentHistory(...): Promise<...>;
}
The incident service does not need to know how Hindsight stores memories.
That gives me two useful properties.
1. The incident workflow remains domain-focused
The incident service works with incidents rather than with provider-specific memory structures.
2. Memory behavior can be tested independently
I can test retrieval and storage behavior without coupling every test to the UI.
The production implementation is the Hindsight adapter.
π§ Hindsight Receives the Incident, Not a Fixed Scenario
The Recall query is generated from the current incident:
``ts id="74y9dj"Service: ${service}. Title: ${title}. Symptoms: ${symptoms}. Find past engineering incidents with similar service, symptoms, root causes, decisions, and successful fixes.`;
const query =
There is no:
```text
if payment
β use these two memories
branch.
That is important because the memory system should be able to return relevant experiences for:
- Authentication
- Notifications
- Order processing
- Infrastructure
- An unseen service
The Hindsight GitHub repository provides the underlying memory system; IncidentMind uses it as a general organizational memory rather than a payment-specific lookup table.
π The Parser Had to Become Generic Too
Retrieval alone was not enough.
Hindsight returns memory content and metadata. IncidentMind needs to translate that into a domain model that the UI and reasoning layer understand.
The parser therefore extracts actual fields from recalled memory content:
```text id="g8qu6h"
Incident ID
Title
Service
Root Cause
Resolution
Outcome
It does not assume that a particular memory must represent a particular incident.
This mattered because early demo-oriented mappings could make several different historical memories appear to be copies of the same incident.
That kind of shortcut is dangerous in an operational tool.
It creates **false history**.
The current approach is simpler:
> **Parse what Hindsight actually returned.**
---
## π€ General Reasoning Matters More Than a Bigger Prompt
The reasoning layer receives the current incident and historical context.
It does not receive a payment-specific instruction such as:
> βCompare database pool exhaustion with payment-provider latency.β
Instead, it is asked to analyze the actual evidence.
Conceptually:
```text id="vtx9c5"
Current incident
+
Recalled historical memories
+
Historical comparison
+
Current telemetry
β
Structured triage recommendation
This means an authentication incident can follow the same architecture as a payment incident.
The domain-specific content comes from the incident and its memories, not from hardcoded application logic.
π A New Incident Should Be Allowed to Have No Match
I consider this one of the most important generalization tests.
Suppose the system receives an incident involving a streaming pipeline and Kafka rebalance behavior, but the memory bank contains no relevant historical experience.
A weak implementation might find a vaguely similar infrastructure incident and present it as a historical match.
IncidentMind should not do that.
The expected behavior is:
```text id="7qg6yd"
Unseen incident
β
Hindsight Recall
β
No direct historical match
β
Reason from current evidence
β
Report limited historical context
That is a **successful outcome**.
A memory system does not fail because it cannot remember something that never happened before.
---
## π§ͺ Strong Matches and Unseen Incidents Test Different Things
I use multiple categories of test cases.
### Strong Match
Checks whether Hindsight can connect a new incident to an existing experience.
### Partial Match
Checks whether the system can identify useful similarities without ignoring important differences.
### Unseen Incident
Checks whether the system can operate without inventing history.
Those tests are more valuable than repeatedly testing the same payment incident.
For example:
```text id="j0f0m5"
Payment 503 spike
β historical database-pool incident
β strong match
Authentication outage after credential rotation
β historical authentication incident
β strong match
Order query slowdown
β historical database incidents
β compare carefully
New streaming pipeline failure
β no direct match
β current-evidence reasoning
The system should behave differently in each case.
π Hindsight Gives the Architecture Continuity
Without persistent memory, every incident starts with:
```text id="2t7xne"
Current incident
β
LLM
β
Answer
With Hindsight, the workflow becomes:
```text id="4p3x1h"
Past incidents
β
Hindsight memory
β
Current incident
β
Recall + Reflect
β
Reasoning
β
Resolution
β
New retained experience
That loop is what lets the system improve its available context over time.
The Hindsight documentation was particularly relevant here because the useful unit is not simply retrieval.
The system needs persistent experience that can be recalled and incorporated into later decisions.
π‘οΈ I Deliberately Kept Human Control
Generalization does not mean automation should make production changes.
The system recommends what to investigate.
An engineer decides what evidence to collect and what remediation to apply.
Once the incident is resolved, the outcome becomes a candidate for organizational memory.
This gives the system a useful boundary:
```text id="f0x4f8"
Agent:
I found these historical experiences and recommend investigating X.
Engineer:
Here is what I observed. I will decide the production action.
System:
After resolution, I will retain what we learned.
That is more realistic for incident response than allowing an agent to turn historical similarity directly into infrastructure changes.
---
## π Before and After
### Before Generalization
```text id="k0z1pr"
Known payment scenario
β
Fixed assumptions
β
Expected historical incident
β
Recommendation
After Generalization
```text id="5y6bde"
Any incident
β
Build query from actual incident data
β
Recall relevant experience
β
Compare evidence
β
Reason from current context
β
Allow a no-match result
This makes the architecture usable across incident categories rather than tying the memory layer to one demonstration scenario.
---
## π‘ What I Learned
### Generalization Starts With Removing Assumptions
The biggest improvements did not come from adding more prompts.
They came from removing payment-specific branches and making the incident itself the input.
### A Parser Is Part of the Trust Boundary
If recalled memory is mapped to the wrong incident identity, every later reasoning step is contaminated.
### No-Match Cases Are First-Class Tests
A system that only works when it recognizes a known scenario is not general.
### Memory and Reasoning Should Stay Separate
Hindsight remembers organizational experience.
The LLM reasons over that experience and current evidence.
### Production Systems Should Preserve Uncertainty
If the system has no relevant history, it should say so.
That is better than pretending a weak similarity is a known solution.
## π Project and References
The complete IncidentMind project is available in the [IncidentMind GitHub repository](https://github.com/gunja-ramana/incidentmind).
For implementation details and background, see:
* [Hindsight GitHub repository](https://github.com/vectorize-io/hindsight)
* [Hindsight Documentation](https://hindsight.vectorize.io/)
* [Vectorize: What Is Agent Memory?](https://vectorize.io/what-is-agent-memory)




Top comments (0)