DEV Community

Nayaz Bhanu
Nayaz Bhanu

Posted on

How I Built a Memory-First Incident Response Agent with Hindsight

Production incidents rarely happen in isolation. A service becomes slow, an API starts returning errors, or a database suddenly becomes a bottleneck. Engineers investigate the problem, find the root cause, fix it, and write down what happened.

The problem is what happens the next time a similar incident occurs.

The previous solution may exist somewhere in a post-mortem, ticket, Slack conversation, runbook, or someone's personal notes. An engineer can spend valuable time rediscovering information that the organization has already learned.

I wanted to explore a different approach: what if an incident response agent could remember previous incidents and use that experience when investigating the next one?

That idea led me to build Incident Experience AI, a memory-first incident response agent powered by Hindsight.

The idea: turn incident history into operational memory

The central idea is simple.

Instead of treating every production incident as a completely new problem, the system stores the experience gained from previous incidents and recalls relevant experience when a new incident arrives.

The workflow is:

NEW INCIDENT
     ↓
HINDSIGHT RECALL
     ↓
AI REASONING
     ↓
INVESTIGATION RECOMMENDATION
     ↓
RESOLUTION
     ↓
CAPTURE OUTCOME
     ↓
HINDSIGHT RETAIN
     ↓
FUTURE INCIDENT
Enter fullscreen mode Exit fullscreen mode

This creates a learning loop.

A resolved incident is not just something that disappears from the current investigation. Its symptoms, investigation path, root cause, resolution, outcome, and lessons learned can become useful context for a future incident.

For example, imagine that a previous Production Orders API incident involved high latency. Engineers investigated application logs, database metrics, CPU, memory, and network performance and discovered that a frequently used database query was missing an index. Adding the index restored normal latency and the incident was resolved in 18 minutes.

If a similar Orders API latency problem happens later, the agent can recall that experience and use it as evidence when suggesting where engineers should investigate first.

Why Hindsight is important

The memory layer is the central part of the system.

I used Hindsight as the persistent memory system for incident experience. The application connects to a Hindsight bank named:

HINDSIGHT_BANK_ID = "Incident Experience"
Enter fullscreen mode Exit fullscreen mode

The agent uses two important operations: retain and recall.

Retain allows the application to store an incident experience. Recall allows it to search the accumulated experience using a natural-language query.

A simplified example from my implementation is:

client.retain(
    bank_id=BANK_ID,
    content=incident
)
Enter fullscreen mode Exit fullscreen mode

The incident stored in this test contains structured information such as the affected service, symptoms, investigation, root cause, resolution, outcome, resolution time, successful action, and lesson learned.

For example:

Incident ID: INC-001
Service: Production Orders API

Symptom:
API latency increased significantly...

Root Cause:
The orders database table was missing an index...

Resolution:
The team added the missing database index.

Outcome:
API latency returned to normal.

Resolution Time:
18 minutes.
Enter fullscreen mode Exit fullscreen mode

The important part is that the memory is not just a label such as "Orders API was slow." It contains the experience surrounding the incident.

Recalling previous experience

When a new incident arrives, the application constructs a recall query describing the incident and the type of experience that would be useful.

My implementation asks Hindsight to focus on:

  • similar symptoms
  • affected services
  • root causes
  • investigation approaches
  • successful remediation
  • failed approaches
  • resolution time
  • lessons learned

The actual recall call is:

result = client.recall(
    bank_id=HINDSIGHT_BANK_ID,
    query=query,
    max_tokens=5000,
    budget="high"
)
Enter fullscreen mode Exit fullscreen mode

The returned memories are then passed into the AI reasoning stage.

This distinction is important: Hindsight provides historical incident experience; the language model reasons over that experience in the context of the new incident.

Combining memory with AI reasoning

For the reasoning layer, I used Groq with the openai/gpt-oss-120b model.

The application sends the new incident together with the historical memories returned by Hindsight.

The reasoning prompt explicitly tells the model not to invent historical incidents and to distinguish historical evidence from its own inference.

For example, the prompt includes instructions such as:

- Do not invent historical incidents.
- Do not claim something is certain unless evidence supports it.
- Clearly distinguish historical evidence from your own inference.
- Prefer actions that succeeded in similar historical incidents.
- Give practical investigation steps.
Enter fullscreen mode Exit fullscreen mode

The model returns structured information including:

summary
severity
likely_root_cause
investigation_steps
recommended_action
confidence
evidence
Enter fullscreen mode Exit fullscreen mode

This makes the result more useful during an incident because the agent does not simply produce a paragraph of text. It produces an investigation-oriented assessment.

The application then displays the historical evidence separately from the AI's recommendation.

It also explicitly tells the responder that similarity is not proof of a shared root cause. The recommendation still needs to be validated against live service telemetry.

The investigation workflow

The Streamlit application provides an incident investigation interface.

An engineer enters:

  • Incident ID
  • affected service
  • symptoms
  • business or customer impact
  • recent changes or context

The application converts these details into an incident record and starts the investigation.

The first step is Hindsight recall.

The recalled experiences are displayed in the interface, including information such as the previous root cause, resolution, resolution time, and outcome.

The AI then generates an investigation path based on both the current incident and the recalled historical experience.

This produces a flow like:

Current Incident
       ↓
Similar Experience
       ↓
Recalled Memory
       ↓
Investigation Path
       ↓
Recommended Action
Enter fullscreen mode Exit fullscreen mode

The UI also exposes the number of recalled experiences and the confidence of the AI assessment, making the role of memory visible rather than hiding it inside the application.

[Screenshot: Incident investigation dashboard showing recalled memories and AI recommendation]

The most important step: learning from the resolution

The learning loop does not stop when the AI produces a recommendation.

After the incident is actually resolved, the engineer records what fixed it, the outcome or lesson learned, and the resolution time.

The application then creates a new incident response record and stores it back in Hindsight:

client.retain(
    bank_id=HINDSIGHT_BANK_ID,
    content=content,
    context="production incident postmortem"
)
Enter fullscreen mode Exit fullscreen mode

The stored experience includes the original incident, AI analysis, actual resolution, outcome, resolution time, and instructions to preserve useful information such as symptoms, root cause, investigation path, remediation, successful actions, unsuccessful actions, and lessons learned.

This creates the core feedback loop:

Recall → Investigate → Resolve → Retain → Recall Again
Enter fullscreen mode Exit fullscreen mode

The important idea is that every resolved incident has the potential to become useful context for the next investigation.

[Screenshot: Resolve & Teach Hindsight section]

Demonstrating the difference memory can make

The application includes a Demo Mode specifically to make this concept visible.

A new incident can be submitted to the system, after which the application searches the configured Hindsight memory bank and compares the available historical experience with the new incident.

If relevant memories are returned, the interface displays them alongside the resulting recommendation.

If no relevant memory is returned, the application does not fabricate one. Instead, it tells the user that no matching historical experience was found.

This is important because an incident response system should not pretend that historical evidence exists when it does not.

The intended demonstration is straightforward:

Without experience: investigate a new incident with only the current incident details.

After experience has been retained: investigate a similar incident while using relevant historical evidence to guide the investigation.

The value is not that the agent automatically knows the answer. The value is that previous operational experience becomes available at the moment it can be useful.

What I learned

Building this system changed how I think about AI agents.

A language model can generate an answer to an incident, but generation alone does not give the system organizational experience. The memory layer changes the interaction because previous outcomes can become part of future reasoning.

I also learned that memory quality matters. Storing an incident without its context would provide limited value. The useful information is the combination of symptoms, investigation, root cause, remediation, outcome, and lessons learned.

Another important lesson was to keep historical evidence and AI inference separate. A previous incident can be highly relevant without proving that the current incident has the same root cause. The system therefore presents recalled experience as evidence that should guide investigation, not as certainty.

Limitations and next steps

The current implementation is a focused prototype rather than a complete production incident-management platform.

It currently relies on manually entered incident information and captured resolutions. A production version could integrate with monitoring systems, logs, incident-management platforms, tickets, and deployment systems so that incident context is collected automatically.

The system would also benefit from more extensive evaluation of recall quality, recommendation accuracy, and whether memory actually reduces investigation time across a larger set of historical incidents.

Access control, auditability, richer observability, and stronger evaluation would also be necessary before using such a system with sensitive production data.

Conclusion

Incident Experience AI explores a simple idea: production incidents should produce reusable experience, not just temporary fixes.

By combining Hindsight persistent memory with an LLM-based reasoning layer, the system can recall previous incident experiences, use them as evidence during a new investigation, recommend practical investigation steps, and retain the final outcome for future use.

The resulting loop is:

INCIDENT
   ↓
RECALL EXPERIENCE
   ↓
REASON
   ↓
INVESTIGATE
   ↓
RESOLVE
   ↓
RETAIN EXPERIENCE
   ↓
LEARN FOR THE NEXT INCIDENT
Enter fullscreen mode Exit fullscreen mode

The goal is not to replace the engineer responding to an incident. It is to make the organization's accumulated incident experience available when it matters most.



https://github.com/kbasheerbasha027/incident-response-agent?utm_source=chatgpt.com

Top comments (1)

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

plot twist: green settlement tiles are not a signed hop tip.

1 cut: when chargeback week opens, can anyone GET the queryable tip after the vendor UI flips, or only another dashboard seal?

receipts > seals. marker0929h2228-dt