DEV Community

Vishnu vardhan
Vishnu vardhan

Posted on

I Designed Incident Memory to Work Beyond Payment Failures

The easiest way to build an incident agent is to make it excellent at one demo scenario.

The harder problem is making it useful when the next incident looks nothing like the previous one.

I designed IncidentMind around incident scenarios across payment, authentication, orders, notifications, web services, databases, queues, and infrastructure.

That forced me to remove a class of shortcuts that made early versions look good but would not survive outside one scenario.

The central design goal became simple:

Hindsight should provide the memory, while the reasoning path should work from whatever incident the engineer actually reports.


⚠️ The Danger of Scenario-Specific Intelligence

Payment failures were useful while developing the system because they provided concrete examples.

But payment incidents also made it easy to accidentally encode assumptions such as:

```text id="z4v3uk"
Payment API
β†’ database pool
β†’ external payment provider




That is not an incident-response system.

It is a **payment troubleshooting workflow**.

A real incident responder may receive:

* Authentication failures
* Order API latency
* Notification queue backlog
* Kubernetes resource exhaustion
* Certificate failures
* DNS problems
* Database saturation
* Third-party API timeouts
* A completely new failure mode

The architecture therefore has to treat the **incident itself as data**.

---

## 🧩 The Memory Interface Is Deliberately Generic

The backend isolates memory operations behind a service abstraction.

The important operations are conceptually:



```ts id="r9u4kv"
interface IMemoryService {
  recallSimilarIncidents(...): Promise<...>;
  retrieveRelevantMemories(...): Promise<...>;
  storeIncidentExperience(...): Promise<...>;
  getIncidentHistory(...): Promise<...>;
}
Enter fullscreen mode Exit fullscreen mode

The incident service does not need to know how Hindsight stores memories.

That gives me two useful properties.

1. The incident workflow remains domain-focused

The incident service works with incidents rather than with provider-specific memory structures.

2. Memory behavior can be tested independently

I can test retrieval and storage behavior without coupling every test to the UI.

The production implementation is the Hindsight adapter.


🧠 Hindsight Receives the Incident, Not a Fixed Scenario

The Recall query is generated from the current incident:

``ts id="74y9dj"
const query =
Service: ${service}. Title: ${title}. Symptoms: ${symptoms}. Find past engineering incidents with similar service, symptoms, root causes, decisions, and successful fixes.`;




There is no:



```text
if payment
    β†’ use these two memories
Enter fullscreen mode Exit fullscreen mode

branch.

That is important because the memory system should be able to return relevant experiences for:

  • Authentication
  • Notifications
  • Order processing
  • Infrastructure
  • An unseen service

The Hindsight GitHub repository provides the underlying memory system; IncidentMind uses it as a general organizational memory rather than a payment-specific lookup table.


πŸ” The Parser Had to Become Generic Too

Retrieval alone was not enough.

Hindsight returns memory content and metadata. IncidentMind needs to translate that into a domain model that the UI and reasoning layer understand.

The parser therefore extracts actual fields from recalled memory content:

```text id="g8qu6h"
Incident ID
Title
Service
Root Cause
Resolution
Outcome




It does not assume that a particular memory must represent a particular incident.

This mattered because early demo-oriented mappings could make several different historical memories appear to be copies of the same incident.

That kind of shortcut is dangerous in an operational tool.

It creates **false history**.

The current approach is simpler:

> **Parse what Hindsight actually returned.**

---

## πŸ€– General Reasoning Matters More Than a Bigger Prompt

The reasoning layer receives the current incident and historical context.

It does not receive a payment-specific instruction such as:

> β€œCompare database pool exhaustion with payment-provider latency.”

Instead, it is asked to analyze the actual evidence.

Conceptually:



```text id="vtx9c5"
Current incident
       +
Recalled historical memories
       +
Historical comparison
       +
Current telemetry
       ↓
Structured triage recommendation
Enter fullscreen mode Exit fullscreen mode

This means an authentication incident can follow the same architecture as a payment incident.

The domain-specific content comes from the incident and its memories, not from hardcoded application logic.


πŸ†• A New Incident Should Be Allowed to Have No Match

I consider this one of the most important generalization tests.

Suppose the system receives an incident involving a streaming pipeline and Kafka rebalance behavior, but the memory bank contains no relevant historical experience.

A weak implementation might find a vaguely similar infrastructure incident and present it as a historical match.

IncidentMind should not do that.

The expected behavior is:

```text id="7qg6yd"
Unseen incident
↓
Hindsight Recall
↓
No direct historical match
↓
Reason from current evidence
↓
Report limited historical context




That is a **successful outcome**.

A memory system does not fail because it cannot remember something that never happened before.

---

## πŸ§ͺ Strong Matches and Unseen Incidents Test Different Things

I use multiple categories of test cases.

### Strong Match

Checks whether Hindsight can connect a new incident to an existing experience.

### Partial Match

Checks whether the system can identify useful similarities without ignoring important differences.

### Unseen Incident

Checks whether the system can operate without inventing history.

Those tests are more valuable than repeatedly testing the same payment incident.

For example:



```text id="j0f0m5"
Payment 503 spike
  β†’ historical database-pool incident
  β†’ strong match

Authentication outage after credential rotation
  β†’ historical authentication incident
  β†’ strong match

Order query slowdown
  β†’ historical database incidents
  β†’ compare carefully

New streaming pipeline failure
  β†’ no direct match
  β†’ current-evidence reasoning
Enter fullscreen mode Exit fullscreen mode

The system should behave differently in each case.


πŸ”„ Hindsight Gives the Architecture Continuity

Without persistent memory, every incident starts with:

```text id="2t7xne"
Current incident
↓
LLM
↓
Answer




With Hindsight, the workflow becomes:



```text id="4p3x1h"
Past incidents
      ↓
Hindsight memory
      ↓
Current incident
      ↓
Recall + Reflect
      ↓
Reasoning
      ↓
Resolution
      ↓
New retained experience
Enter fullscreen mode Exit fullscreen mode

That loop is what lets the system improve its available context over time.

The Hindsight documentation was particularly relevant here because the useful unit is not simply retrieval.

The system needs persistent experience that can be recalled and incorporated into later decisions.


πŸ›‘οΈ I Deliberately Kept Human Control

Generalization does not mean automation should make production changes.

The system recommends what to investigate.

An engineer decides what evidence to collect and what remediation to apply.

Once the incident is resolved, the outcome becomes a candidate for organizational memory.

This gives the system a useful boundary:

```text id="f0x4f8"
Agent:
I found these historical experiences and recommend investigating X.

Engineer:
Here is what I observed. I will decide the production action.

System:
After resolution, I will retain what we learned.




That is more realistic for incident response than allowing an agent to turn historical similarity directly into infrastructure changes.

---

## πŸ“Š Before and After

### Before Generalization



```text id="k0z1pr"
Known payment scenario
      ↓
Fixed assumptions
      ↓
Expected historical incident
      ↓
Recommendation
Enter fullscreen mode Exit fullscreen mode

After Generalization

```text id="5y6bde"
Any incident
↓
Build query from actual incident data
↓
Recall relevant experience
↓
Compare evidence
↓
Reason from current context
↓
Allow a no-match result




This makes the architecture usable across incident categories rather than tying the memory layer to one demonstration scenario.

---

## πŸ’‘ What I Learned

### Generalization Starts With Removing Assumptions

The biggest improvements did not come from adding more prompts.

They came from removing payment-specific branches and making the incident itself the input.

### A Parser Is Part of the Trust Boundary

If recalled memory is mapped to the wrong incident identity, every later reasoning step is contaminated.

### No-Match Cases Are First-Class Tests

A system that only works when it recognizes a known scenario is not general.

### Memory and Reasoning Should Stay Separate

Hindsight remembers organizational experience.

The LLM reasons over that experience and current evidence.

### Production Systems Should Preserve Uncertainty

If the system has no relevant history, it should say so.

That is better than pretending a weak similarity is a known solution.


## πŸš€ Project and References

The complete IncidentMind project is available in the [IncidentMind GitHub repository](https://github.com/gunja-ramana/incidentmind).

For implementation details and background, see:

* [Hindsight GitHub repository](https://github.com/vectorize-io/hindsight)
* [Hindsight Documentation](https://hindsight.vectorize.io/)
* [Vectorize: What Is Agent Memory?](https://vectorize.io/what-is-agent-memory)


Enter fullscreen mode Exit fullscreen mode

Architecture

IncidentMind Architecture

UI

IncidentMind Screenshot

recall

Hindsight Recall Screenshot

retain

Hindsight Retain Screenshot

Top comments (0)