DEV Community

Sai Madhuri
Sai Madhuri

Posted on

I Built an Agent That Learns From Resolved Incidents

I Gave Incident Response a Memory With Hindsight

ai #devops #sre #hindsight #agents #incidentresponse

Production incidents rarely happen in isolation.

A Payment API returns 502s. A database connection pool is exhausted. Someone investigates, finds the fix, and closes the incident. A few months later, something similar happens and another engineer starts the investigation from the beginning.

I wanted to change that part of the workflow.

I built Incident-Memory-Copilot, an incident-response agent that uses Hindsight as its persistent memory layer. The idea is simple: when an incident happens, the agent should be able to look at what the organization learned from previous incidents, use that context during investigation, and then retain the new lesson after the incident is resolved.

The interesting part is not just recalling an old incident.

It is closing the loop.

The problem: incident knowledge gets lost

During an outage, an engineer may have access to:

Current incident details

Logs and metrics

Deployment information

Runbooks

Previous incident reports

Postmortems

Notes about what worked

Notes about what failed

The information exists, but the connection between the current incident and previous experience is usually the hard part.

A normal LLM can look at an HTTP 502 and give a reasonable list of possible causes.

But I wanted the agent to answer a more useful question:

“Have we seen something like this before, and what did we learn from it?”

That is where Hindsight became the memory layer.

What I built

The application is an operations console for incident investigation.

The overview brings together active incidents, historical memory, memory records, and the Hindsight connection.

Figure 1 — Incident Operations dashboard showing active incidents, historical memory, memory records, and Hindsight activity.

The important design choice is that memory is part of the incident lifecycle rather than a separate search feature.

The flow is:

Current Incident
↓
Recall historical context
↓
Reflect across relevant memories
↓
Recommended investigation
↓
Human review
↓
Resolution
↓
Postmortem / learning
↓
Retain new organizational memory


That gives the system a persistent learning loop instead of a one-shot answer.

The architecture

I designed the system around a simple separation of responsibilities.

The data foundation is intentionally broader than a collection of incident titles.
The system uses:

Rootly incident/log information

PagerDuty incident-response knowledge

PagerDuty postmortem knowledge

100–150 realistic synthetic incident records

50–100 realistic runbooks

50–100 realistic postmortems

The point is to give the memory layer enough operational context to answer questions about previous experience rather than just retrieve isolated documents.

The core flow: Retain, Recall, Reflect

The Hindsight integration maps naturally to the incident lifecycle.

  1. Recall before investigation

When a new incident arrives, the agent first looks for relevant historical context.

Conceptually:

memories = client.recall(
bank_id=BANK_ID,
query=current_incident
)

The query is based on the current incident rather than a generic search phrase.

That matters because the useful historical context depends on what is happening now.

For example, a Payment API connection problem should surface memories around connection pools, database exhaustion, similar payment-service incidents, and related operational lessons.

The goal is not to blindly reuse the previous fix.

The goal is to give the engineer more context before making the next decision.

  1. Reflect when several memories matter

Recall gives the agent relevant memories.

Reflection helps when the useful answer is distributed across multiple memories.

Conceptually:

reflection = client.reflect(
bank_id=BANK_ID,
query="What patterns and failed fixes should I consider?"
)

This gives the agent a way to move from individual historical incidents toward a broader operational pattern.

  1. Retain the actual lesson

After the incident is resolved, the most important information should not disappear.

The application captures:

Root cause

What worked

What failed

Lesson learned

Prevention

That information can then be taught back to organizational memory.

client.retain(
bank_id=BANK_ID,
content=incident_learning,
context="resolved production incident"
)

The exact value of retention depends on what we put into memory.

“Incident closed” is not very useful.

A much better memory is:

Symptom
→ Investigation
→ Failed action
→ Successful action
→ Root cause
→ Lesson
→ Prevention

That is the information another engineer can actually use later.

A real incident flow

The most useful screen in the application is the active-incident investigation view.
The example is a Payment API incident returning HTTP 502 errors.

The screen contains current incident information and then brings in historical context.

The investigation is separated into:

Historical matches

Hindsight reflection

Recommended investigation

Actions requiring human review

This is important because historical similarity should not automatically become the root cause.

A previous incident is evidence.

It is not proof.

If the current environment has changed, blindly repeating an old remediation can make an incident worse.

The agent therefore uses memory to improve the investigation while leaving the final operational decision with the engineer.

The part I cared about most: remembering what failed

When an incident is resolved, the application asks the engineer to capture the complete lesson.




The example shown in the application is a database connection-pool problem.

The root cause is that connection-pool limits were misconfigured in deployment v2.8.0.

Rolling back to v2.7.9 restored the connection-pool limits.

But restarting the API gateway was a failed approach because it caused a traffic spike and made the database problem worse.

That failed action is worth remembering.

If the next engineer only sees:

“Rollback to v2.7.9.”

they know what worked.

If they also see:

“Do not restart the gateway first during this type of database-exhaustion event.”

they know what to avoid.

That is a much more useful organizational memory.

Why this is different from a normal runbook

Runbooks are excellent for known procedures.

But incident response also contains experience that is difficult to encode as a fixed procedure.

A runbook can say:

Check database connection utilization.
Check application pool limits.
Check recent deployments.

An incident memory can add:

This failure previously appeared after deployment v2.8.0.
Rolling back restored the connection-pool limits.
Restarting the gateway increased traffic and worsened the condition.

The runbook tells you what to check.

The memory tells you what happened when someone checked it before.

I wanted both.

Making memory visible

I also wanted the system to make its memory activity visible instead of hiding it behind an AI response.

The dashboard exposes:

RECALL
8 memories retrieved

REFLECT
Historical pattern synthesized

RETAIN
New learning stored

This makes the lifecycle easier to understand:

Recall what happened before.

Reflect on the historical evidence.

Retain what we learned this time.

For an incident-response system, that transparency matters.

An engineer should be able to understand why the agent is recommending a particular investigation path.

Human review is part of the architecture

I intentionally kept the engineer in the loop.

The agent can retrieve historical context and produce recommended actions, but it does not silently execute a production remediation.

The workflow is:

Incident
↓
Historical memory
↓
Agent investigation
↓
Evidence-backed recommendation
↓
Human review
↓
Resolution
↓
Postmortem
↓
New memory

That boundary matters because historical information can become stale.

A service can be upgraded.

A deployment architecture can change.

A previous workaround can stop being safe.

Persistent memory should therefore support engineering judgment, not replace it.

What I learned

  1. Memory is more than storage

Putting incident documents somewhere does not automatically create useful memory.

The important part is retrieving the right experience in the context of a later decision.

  1. The quality of retained information matters

If I retain only incident titles and status values, future recall will be shallow.

Root causes, actions, failed approaches, lessons, and prevention steps are much more valuable.

  1. Failed actions deserve first-class treatment

The successful fix tells me what to do.

The failed fix tells me what not to repeat.

For incident response, both are operational knowledge.

  1. Historical similarity is evidence, not truth

A previous incident can be highly relevant without having the same root cause.

That is why current evidence and human review remain part of the workflow.

  1. The real loop is Incident → Memory → Incident

The most interesting behavior is not one successful recall.

It is the repeated cycle:

Incident
↓
Investigation
↓
Resolution
↓
Learning
↓
Hindsight Memory
↓
Next Incident

The system gets a chance to make previous operational experience useful again.

Why I chose Hindsight

I used Hindsight because its memory model fits this workflow naturally.

The Hindsight GitHub repository and Hindsight documentation describe the retain, recall, and reflect model that the system is built around.

The broader Vectorize guide to agent memory is also useful for understanding the difference between temporary conversational context and persistent agent memory.

For this project, the division is straightforward:

Retain validated incident learning.

Recall relevant historical experience.

Reflect when several memories need to be synthesized.

That gives the incident-response agent a persistent source of organizational context.

Where I would take it next

The next step is not simply adding more UI.

I would connect the system more deeply with operational systems:

Monitoring and observability platforms

Log-management systems

Incident-management platforms

Deployment and CI/CD systems

Engineering communication channels

Automated postmortem generation

Runbook repositories

I would also evaluate the memory layer itself.

For example:

How often does recall surface a genuinely relevant incident?

Which retained fields contribute most to useful recommendations?

How often does the historical context agree with the eventual root cause?

Which failed actions are successfully avoided later?

How does recall change as the memory bank grows?

Those questions would tell me whether the memory layer is actually improving incident investigation.

Closing the loop

The architecture ultimately comes down to one idea:

             ┌──────────────┐
             │   INCIDENT   │
             └──────┬───────┘
                    ↓
             ┌──────────────┐
             │ INVESTIGATE  │
             └──────┬───────┘
                    ↓
             ┌──────────────┐
             │   RESOLVE    │
             └──────┬───────┘
                    ↓
             ┌──────────────┐
             │    LEARN     │
             └──────┬───────┘
                    ↓
             ┌──────────────┐
             │    RETAIN    │
             └──────┬───────┘
                    │
                    └──────────────→ NEXT INCIDENT
Enter fullscreen mode Exit fullscreen mode

I did not want to build an assistant that gives an answer and forgets it.

I wanted an incident-response system where the resolution of one outage can become useful context for the next one.

That is the idea behind Incident-Memory-Copilot:

turn incident history into organizational memory, and organizational memory into better-informed incident response.

Project

Incident-Memory-Copilot on GitHub

Hindsight on GitHub

Hindsight Documentation

Vectorize — What Is Agent Memory?

Top comments (0)