DEV Community

Sanjay Kumar
Sanjay Kumar

Posted on

From Stateless Incident Response to Persistent AI Memory

Modern AI agents can analyze an incident and generate recommendations, but a major limitation appears when the agent has no memory of previous incidents. Every new incident can effectively become a fresh investigation, even when a similar problem has already been solved.

Our Incident Response Agent addresses this by introducing a persistent memory layer using Hindsight. The memory layer allows the system to retain information from resolved incidents and recall relevant historical incidents when a new incident occurs.

The basic learning loop is:

Incident → Investigation → Resolution → Retain → New Incident → Recall → AI Agent

This creates a connection between past incident experience and future investigations.

Hindsight Retain: Storing Incident Experience

When an incident is resolved, important information should not be discarded. We use Hindsight Retain to store the incident as a memory.

The information retained includes:

Incident ID
Service
Severity
Timestamp
Symptoms
Logs
Root cause
Resolution
Lessons learned

For example, consider a payment API incident:

Incident ID: INC-001
Service: payment-api
Severity: HIGH

Symptoms:

  • high latency
  • request timeouts

Logs:

  • ConnectionPoolTimeout
  • DB connection limit reached

Root Cause:
Database connection pool exhausted

Resolution:
Increased database connection pool size

Lessons Learned:
Check connection pool when payment requests timeout

The incident is stored using Hindsight:

result = hindsight.retain(
bank_id=BANK_ID,
content=content,
context="production incident and resolution",
metadata={
"incident_id": incident["id"],
"service": incident["service"],
"severity": incident["severity"],
"type": "incident"
},
document_id=f"incident_{incident['id']}"
)

The content contains the actual incident experience, while metadata helps associate the memory with information such as the incident ID, service, and severity.

Retain Output:

Our implementation produces an output confirming that the incident was successfully stored:

STEP 1: RETAIN INCIDENT 1

Incident 1 retained!
success=True

This demonstrates the first part of the memory pipeline: converting a resolved incident into persistent historical knowledge.

Hindsight Recall: Finding Relevant Historical Incidents

Storing information is only useful if the system can retrieve it when needed.

When a new incident occurs, our system sends its symptoms, service information, and logs to Hindsight Recall. The recall query asks Hindsight to find previous incidents with similar characteristics and useful resolutions.

For example, a second payment API incident might contain:

Service:
payment-api

Symptoms:

  • slow payment requests
  • request timeouts
  • high database connections

Logs:

  • PaymentRequestTimeout
  • DB connection usage high

The system constructs a recall request asking for previous incidents with similar symptoms, logs, service behavior, failures, resolutions, and lessons learned.

result = hindsight.recall(
bank_id=BANK_ID,
query=query
)

Hindsight can then return relevant historical memories.

In our example, the recall operation retrieves the earlier incident:

STEP 2: RECALL SIMILAR INCIDENTS

Historical matches found:

Memory ID: ...
Type: ...
Context: ...

Incident ID: INC-001

Root Cause:
Database connection pool exhausted

Resolution:
Increased database connection pool size

Lessons Learned:
Check connection pool when payment requests timeout

This demonstrates the second part of the memory pipeline: retrieving previous incident experience when a new incident needs investigation.

Before and After: Stateless vs Memory-Aware Investigation

Without persistent memory, the AI agent primarily receives information about the current incident:

New Incident
↓
AI Agent
↓
Investigation
↓
Recommendation

The previous incident may have already been solved, but its experience is not automatically available to the next investigation.

With Hindsight, the flow becomes:

Previous Incident
↓
Resolution
↓
Hindsight Retain
↓
Persistent Memory
↓
Hindsight Recall
↑
│
New Incident
↓
AI Agent
↓
Investigation + Recommendation

This means the AI agent can receive both current incident information and relevant historical context.

For example, when INC-002 has symptoms related to database connections and request timeouts, the memory layer can retrieve INC-001 and provide its previous root cause, resolution, and lesson learned.

The historical incident does not automatically prove that INC-002 has the same root cause. Instead, it gives the AI agent additional evidence to consider during investigation.

Connecting Hindsight to the AI Agent

The memory layer is connected to the investigation pipeline through the backend.

When the investigation endpoint is called, the backend first performs a Hindsight recall:

historical_memories = recall_incidents(incident_data)

The returned memories are then passed to the AI investigation agent:

result = run_agent_investigation(
incident=incident_data,
historical_memories=historical_memories,
)

The AI agent therefore receives two important sources of information:

The current incident
Relevant historical memories

The agent can use both sources when generating its structured investigation result, including:

Summary
Root cause
Evidence
Historical matches
Recommended action
Confidence

This creates the connection between persistent memory and AI-powered investigation.

The Complete Memory Learning Loop

The complete process can be summarized as:

┌─────────────────────┐
│ Incident 1 │
└──────────┬──────────┘
↓
Investigation
↓
Resolution
↓
┌─────────────────────┐
│ Hindsight Retain │
└──────────┬──────────┘
↓
Persistent Memory
↓
┌─────────────────────┐
│ New Incident │
└──────────┬──────────┘
↓
┌─────────────────────┐
│ Hindsight Recall │
└──────────┬──────────┘
↓
Historical Context
↓
┌─────────────────────┐
│ AI Agent │
└──────────┬──────────┘
↓
Investigation +
Recommendation

The important idea is that resolved incidents become reusable experience rather than disappearing after the incident is closed.

Limitation and Lesson Learned

A historical match should not be treated as proof that the current incident has exactly the same root cause.

Two incidents can have similar symptoms but different underlying causes. Therefore, recalled memories should be treated as evidence and context, while the AI agent should continue evaluating the current incident's logs, symptoms, and other available evidence.

This was an important design consideration for our memory implementation: memory should support investigation, not replace investigation.

Conclusion

Persistent memory changes how an incident response agent can work with historical experience. Hindsight provides the Retain and Recall capabilities needed to preserve resolved incidents and retrieve relevant information later.

In our implementation, a resolved incident such as INC-001 can be retained with its symptoms, logs, root cause, resolution, and lessons learned. When a new incident occurs, Hindsight Recall can retrieve that historical information and make it available to the AI investigation agent.

The resulting system connects past incident experience with present investigation, creating a continuous loop:

Retain → Recall → Investigate → Resolve → Retain

This memory layer is therefore an important part of making an incident response agent more context-aware and capable of using previous operational experience.

Resources
Hindsight GitHub: github.com/vectorize-io/hindsight
Hindsight Documentation: https://hindsight.vectorize.io/?utm_source=chatgpt.com
Agent Memory: https://vectorize.io/what-is-agent-memory?utm_source=chatgpt.com
Project GitHub: https://github.com/DivyaSree0912/incident-response-agent.git?utm_source=chatgpt.com

Top comments (0)