DEV Community

Nihas Ram
Nihas Ram

Posted on

How I Stopped My Team From Repeating The Same Outages

I Built an Incident Copilot That Remembers What Happened Last Time

An incident happens. Someone investigates it, finds the root cause, applies a fix, and writes a postmortem.

Then, months later, a similar incident happens again.

The engineer investigating it may know that something like this happened before, but that knowledge is usually buried inside old postmortems, tickets, chat messages, or documentation. The information exists, but it isn't necessarily available at the moment it matters.

That led me to a simple question:

What if an incident-response copilot could remember what happened during previous incidents and use that experience when investigating the next one?

That's the idea behind Incident-Memory-Copilot.

Instead of treating every incident as an isolated question for an LLM, I built a workflow around persistent agent memory:

Recall → Reason → Resolve → Retain

The goal isn't simply to generate another AI-generated troubleshooting checklist. It's to make previous incident experience available as context for future investigations.

Incident-Memory-Copilot main interface

Incident-Memory-Copilot provides an interface for investigating operational incidents with persistent agent memory.

The problem with starting every incident from zero

Consider a payment service returning intermittent 502 Bad Gateway errors.

A general-purpose LLM can produce a reasonable list of possibilities:

  • Check upstream services.
  • Inspect application logs.
  • Review connection pools.
  • Check recent deployments.
  • Look at latency and error rates.

None of those suggestions are necessarily wrong.

But they aren't specific to the environment where the incident is happening.

Suppose a previous incident showed that the same service experienced connection-pool exhaustion, and that a particular attempted fix made the situation worse.

That information is much more useful than a generic troubleshooting checklist.

The challenge is therefore not only generating an answer.

It is bringing the right previous experience into the current investigation.

That's where Hindsight becomes the memory layer of the system.

How Incident-Memory-Copilot works

The system follows a simple lifecycle:

CURRENT INCIDENT
       │
       ▼
  Hindsight Recall
       │
       ▼
Relevant Past Experience
       │
       ▼
   LLM Reasoning
       │
       ▼
Investigation Recommendation
       │
       ▼
    Resolution
       │
       ▼
    Postmortem
       │
       ▼
  Hindsight Retain
       │
       └──────────► Future Incidents
Enter fullscreen mode Exit fullscreen mode

When a new incident is submitted, the system first looks for relevant historical memory.

The retrieved context is then incorporated into the reasoning process used to generate the investigation and recommendation.

After the incident is resolved, the resulting knowledge can be retained so that a future incident has access to it.

This creates a feedback loop instead of a one-way question-and-answer system.

The incident flows through recall, reasoning, resolution, and retention, creating a persistent memory loop.

Recall before reasoning

One of the most important design decisions was to make memory retrieval an explicit part of the incident workflow.

Instead of immediately sending the incident description to an LLM, the application first constructs a query from the incident information and searches the Hindsight memory.

Conceptually:

Current incident
      ↓
Extract useful symptoms
      ↓
Search historical memory
      ↓
Retrieve relevant experiences
      ↓
Give that context to the reasoning step
      ↓
Generate recommendation
Enter fullscreen mode Exit fullscreen mode

Hindsight provides the memory operations needed for this workflow, including retaining information and recalling relevant memories through natural-language queries.

This changes the question the model is answering.

Instead of:

"What could cause this error?"

the system can reason with additional context:

"What could cause this error, and what have we learned from similar incidents before?"

That distinction is the core idea behind the project.

Hindsight recall result

Hindsight recall retrieves relevant historical incident experience before the copilot reasons about the current incident.

Retaining what we learned

Recall is only useful if the system has something meaningful to recall.

That's why the other half of the workflow is retention.

After an incident is resolved, the important information can be captured as an incident memory containing things such as:

  • The affected service
  • The symptoms
  • The root cause
  • What worked
  • What didn't work
  • The lesson learned

That information can then become available to later investigations.

The important design principle is that memory should be part of the existing incident workflow rather than a separate knowledge-management task.

Hindsight provides dedicated memory operations for storing information for later retrieval and searching previously stored memory using natural-language queries.

Postmortem and Hindsight retain workflow

Memory Updated

A resolved incident is retained as reusable memory so that future investigations can benefit from the experience.

The interesting part: the second incident

The most useful demonstration of the system isn't the first incident.

It's the second one.

Imagine the system has already retained knowledge from a previous payment-service incident.

Later, another incident occurs:

Checkout requests are intermittently failing
with gateway errors during increased traffic.

Latency has increased and upstream
connection failures are being reported.
Enter fullscreen mode Exit fullscreen mode

The wording is different.

The incident isn't simply copied from the previous example.

The system recalls relevant historical experience and makes that information available to the reasoning stage.

Conceptually:

NEW INCIDENT
     +
HISTORICAL EXPERIENCE
     ↓
   REASONING
     ↓
INVESTIGATION PLAN
Enter fullscreen mode Exit fullscreen mode

This is where persistent memory becomes useful.

The model isn't expected to remember the previous incident by itself. Instead, the memory system provides relevant historical context when it is needed.

A new incident triggers recall of relevant historical experience, allowing the previous incident to become part of the current investigation context.

Why memory is different from a normal chatbot

A normal chatbot conversation is primarily focused on the current interaction.

An incident-memory system has a longer lifecycle:

Incident
   ↓
Investigation
   ↓
Resolution
   ↓
Learning
   ↓
Memory
   ↓
Future Incident
   ↓
Recall
   ↓
Investigation
Enter fullscreen mode Exit fullscreen mode

The previous incident becomes part of the context for the next one.

That means the system can potentially accumulate operational knowledge over time instead of treating every request as independent.

This is also why I chose Hindsight rather than implementing the memory layer as a simple collection of prompt text. Hindsight provides dedicated memory capabilities for agent workflows.

Recall, reflect, and reason

Another important distinction in the architecture is between retrieving memories and reasoning over them.

Recall gives the system relevant historical information.

Reflection can then turn retrieved information into a more useful higher-level context for reasoning.

The overall flow becomes:

Incident
   ↓
Recall
   ↓
Relevant memories
   ↓
Reflect
   ↓
Historical insight
   ↓
LLM reasoning
   ↓
Recommendation
Enter fullscreen mode Exit fullscreen mode

This separation is useful because raw retrieved memories aren't necessarily the best format for a reasoning prompt.

The system can first identify relevant historical context and then use that context when producing the final investigation output.

What I learned building it

1. Retrieval is only useful when the context is relevant

Adding memory to an agent doesn't automatically make its answers better.

The important question is:

Did we retrieve something that actually matters to this incident?

Poor retrieval can add noise instead of useful context.

That makes the design of the incident query and the structure of retained information important parts of the system.

2. The memory loop matters more than a single prompt

It's easy to build an LLM prompt that generates an incident checklist.

The more interesting engineering problem is building the complete lifecycle:

Recall
  ↓
Reason
  ↓
Resolve
  ↓
Retain
  ↓
Recall again
Enter fullscreen mode Exit fullscreen mode

The value appears when information from one incident becomes useful during another.

3. Failed actions are valuable knowledge

Incident documentation shouldn't only record what worked.

Knowing what didn't work can also be extremely useful.

A future investigation may encounter the same symptoms and consider the same tempting action.

Historical evidence that an action failed under similar conditions can become an important part of the investigation context.

That is one of the reasons I designed the memory around incident experiences rather than simply storing generic documentation.

4. Memory should fit naturally into the workflow

If engineers have to perform a completely separate knowledge-management process after every incident, memory will quickly become stale.

The better approach is to make retention part of the normal resolution and postmortem workflow.

The incident gets resolved.

The useful knowledge gets captured.

That knowledge becomes available to future investigations.

The bigger idea

The interesting thing about incident response isn't that engineers encounter problems.

It's that organizations repeatedly encounter related problems.

Every resolved incident contains some amount of information that could be useful later:

What happened?
Why did it happen?
What fixed it?
What didn't work?
What should we remember?
Enter fullscreen mode Exit fullscreen mode

Traditional documentation answers those questions, but the information can remain passive.

An agent-memory approach makes it possible to bring that knowledge into the next interaction when relevant.

That's the idea I wanted to explore with Incident-Memory-Copilot.

Not an AI that magically knows everything.

Not another chatbot that generates generic troubleshooting advice.

A system that can recall previous experience, reason about the current incident, and retain what was learned for the next one.

The goal is simple: the next incident shouldn't have to start from zero.

Resources

Top comments (0)