DEV Community

Cover image for An Incident-Response Agent Should Remember What Failed
Sharon Medithi
Sharon Medithi

Posted on

An Incident-Response Agent Should Remember What Failed

Production incidents generate enormous amounts of organizational knowledge.

The frustrating part is that this knowledge often gets trapped inside postmortems, tickets, Slack threads, and the memory of the engineer who handled the incident.

An incident-response agent can retrieve logs, inspect metrics, and suggest troubleshooting steps. But one important question can still remain unanswered:

What did we try last time, what failed, and what should we avoid repeating?

That's the problem we built around.

1. The Problem with One-Shot Incident Agents

A typical incident-response workflow might look like this:

Current incident
      ↓
Logs + metrics + deployment information
      ↓
Generic troubleshooting plan
Enter fullscreen mode Exit fullscreen mode

An agent can reason about what it sees now. But production engineering is rarely about the current incident alone.

Organizations accumulate experience from previous incidents:

Incident
   ↓
Hypothesis
   ↓
Action
   ↓
Result
   ↓
Root cause
   ↓
Successful mitigation
Enter fullscreen mode Exit fullscreen mode

That experience can contain something extremely valuable: not only what worked, but also what didn't work.

Without access to that experience, an agent may repeatedly investigate the same failed hypothesis or recommend an action that engineers have already tried unsuccessfully.

The problem, therefore, isn't simply giving an agent more information.

It's giving the agent access to organizational experience that can influence its next decision.

2. Memory Has to Change Decisions

This is the central idea behind our system.

Consider a historical incident: INC-1041.

Historical incident — INC-1041

  • Checkout latency: 4.8s
  • Deployment: v3.8.2
  • Initial hypothesis: database overload
  • Action: increase database capacity
  • Result: failed
  • Actual cause: connection-pool exhaustion
  • Successful mitigation: rollback pool configuration

Now consider a new incident.

Current incident — INC-1187

  • Checkout latency: 5.1s
  • Deployment: v3.9.0
  • Intermittent 504s

With Memory OFF, the agent starts with generic database investigation and database scaling.

With Memory ON, Hindsight recalls INC-1041 and the investigation changes:

CHECK CONNECTION POOL FIRST
          ↓
COMPARE POOL CONFIGURATION
          ↓
REVIEW DEPLOYMENT
          ↓
DATABASE SCALING DEMOTED
Enter fullscreen mode Exit fullscreen mode

The important result isn't simply that the agent remembered INC-1041.

The important result is that the memory changed what the agent did next.

That distinction is fundamental.

A memory system is much more useful when recalled information changes the current decision rather than simply appearing as additional context.

3. Why Failed Actions Are Valuable Memories

Incident response has an unusual relationship with failure.

A failed action during an outage isn't necessarily wasted effort. Once recorded, it becomes operational knowledge.

Consider the historical experience:

Context:
Checkout latency increased

Hypothesis:
Database overload

Action:
Increase database capacity

Outcome:
No improvement

Root cause:
Connection-pool exhaustion

Mitigation:
Rollback pool configuration
Enter fullscreen mode Exit fullscreen mode

A future engineer shouldn't have to rediscover the same lesson during another outage.

Most systems naturally encourage us to preserve successful procedures:

"What worked last time?"

Incident response also needs to preserve another question:

"What didn't work last time?"

That is why we treat an incident as an experience record rather than simply storing its transcript.

Context
   ↓
Hypothesis
   ↓
Action
   ↓
Outcome
   ↓
Root cause
   ↓
Mitigation
Enter fullscreen mode Exit fullscreen mode

A transcript tells us what people said.

An experience record tells us what happened when they acted.

That distinction matters because an unsuccessful action can be just as useful as a successful one when deciding what to investigate next.

4. Where Hindsight Fits

Our application uses a FastAPI backend and a browser-based interface, with Hindsight providing persistent memory for incident experiences.

The overall workflow is:

Current Incident
       ↓
Investigation
       ↓
Retain Incident Experience
       ↓
Hindsight
       ↓
Recall Relevant Experience
       ↓
Current Investigation
       ↓
Changed Priorities
Enter fullscreen mode Exit fullscreen mode

The important design requirement is that memory should be causal to the changed investigation plan.

In other words, the system shouldn't merely retrieve a historical incident and display it to the engineer.

The engineer—or the agent—should be able to answer:

Why did this investigation step move above another one?

The answer should be grounded in the recalled experience:

Because a previous incident with a similar pattern showed that the earlier approach failed, while connection-pool exhaustion was ultimately identified as the cause.

This creates a trace from:

Historical experience
        ↓
Recalled evidence
        ↓
Changed priority
        ↓
Current investigation
Enter fullscreen mode Exit fullscreen mode

That connection is what makes memory operationally meaningful.

5. Why This Isn't Just RAG

Retrieval-Augmented Generation can retrieve useful information.

For example, a runbook might contain instructions such as:

Increase database capacity when database utilization is high.

That's useful operational knowledge.

But our problem is slightly different.

We want to know:

What happened when this organization tried something before?

A runbook might tell an agent what to do.

Incident memory can tell it what happened when someone did it.

For example:

Runbook:

High database load
        ↓
Consider increasing database capacity
Enter fullscreen mode Exit fullscreen mode

versus:

Incident memory:

INC-1041
        ↓
Database capacity increased
        ↓
Incident persisted
        ↓
Actual cause: connection-pool exhaustion
Enter fullscreen mode Exit fullscreen mode

The distinction is important.

The system isn't only retrieving instructions.

It is retrieving experience, actions, and outcomes.

RAG can answer:

"What does the documentation say we should try?"

Incident memory can additionally answer:

"What happened the last time we tried something similar?"

Those are different questions, and both can be useful during incident response.

6. Evaluation

We evaluated the system behaviorally rather than claiming a numerical performance improvement.

The key question was:

Does recalled experience change what the agent investigates first?

Memory OFF

The current incident is investigated primarily from its present evidence.

The agent can identify plausible causes and generate a generic investigation plan, including database-related investigation and scaling.

Memory ON

The system recalls the relevant historical incident.

The previous failed database-scaling attempt becomes part of the reasoning context, and connection-pool investigation is prioritized instead.

The important observable difference is therefore:

Memory OFF

Current incident
      ↓
Generic investigation
Enter fullscreen mode Exit fullscreen mode

versus:

Memory ON

Current incident
      ↓
Historical experience recalled
      ↓
Previous failed action considered
      ↓
Investigation priorities change
Enter fullscreen mode Exit fullscreen mode

This gives us a concrete behavioral test without inventing an accuracy percentage that the current demonstration cannot justify.

The goal isn't to claim that memory automatically produces the correct diagnosis.

The goal is to demonstrate that relevant experience can influence the investigation strategy.

7. Limitations

There are several important limitations.

Historical similarity is evidence, not proof

Two incidents can look similar while having completely different root causes.

A recalled incident should therefore influence investigation priorities, not be treated as definitive evidence of the current root cause.

Current telemetry still matters

Historical experience cannot replace current logs, metrics, traces, deployment information, or other production evidence.

The system still needs to validate historical lessons against what is happening now.

The system recommends rather than autonomously changing production

The agent is intended to support incident investigation and decision-making rather than independently modifying production infrastructure.

This keeps the human engineer in control of production actions.

The demonstration uses synthetic incident data

The incidents used in the demonstration are synthetic.

Therefore, the demo demonstrates the memory mechanism and its effect on investigation behavior, rather than establishing production-scale effectiveness.

Memory quality matters

The usefulness of memory depends on the quality of the experiences being retained.

Incomplete, incorrect, or poorly structured incident records can lead to less useful recollections.

Memory is not automatically valuable simply because more of it exists.

Conclusion

The most interesting part of building this system wasn't making an agent remember more information.

It was making memory affect what the agent did next.

An incident isn't just a collection of logs.

It is a record of hypotheses, failed decisions, successful mitigations, root causes, and lessons learned.

A useful incident-response agent shouldn't have to rediscover those lessons every time a similar problem appears.

The goal isn't simply to remember what happened.

It is to preserve the experience of what happened—and use that experience when deciding what to investigate next.

Not just what happened.

What we learned from it.

Top comments (0)