Production incidents rarely happen in isolation.
When a Payment API starts returning 500 errors at 2 AM, the useful information is often not in the current alert. It is buried in incidents that happened weeks or months earlier: what failed, what engineers checked, what remediation worked, and what initially looked plausible but turned out to be wrong.
I wanted to build an incident-response agent that could actually use that history.
The result is an Incident Learning & Response Agent that combines an AI investigation layer with persistent organizational memory using Hindsight. The interesting part wasn't simply getting an LLM to suggest a root cause. It was making previous incident experience available at the right moment—and then feeding verified outcomes back into memory so future investigations could benefit from them.
The problem: every incident starts too cold
An LLM can be very good at reasoning about an incident description.
Give it:
"Payment API is returning intermittent HTTP 500 errors during heavy traffic. Database connection usage is near its configured limit."
It can recognize that database connection exhaustion is a plausible explanation.
But there is a problem.
The model doesn't inherently know what happened the last time our Payment API experienced the same pattern.
Maybe an earlier incident showed that increasing the connection pool and restarting the service resolved the issue. Maybe another incident looked similar but was actually caused by something else. Maybe an engineer discovered an important operational detail that isn't present in today's alert.
Without persistent memory, that context has to be manually added to every prompt.
That is exactly the problem I wanted to solve.
Instead of treating every production incident as a completely new reasoning problem, I wanted the system to reuse relevant engineering experience.
What I built
The application has a web interface where an engineer can submit an incident with:
- Incident ID
- Service
- Severity
- Description
The frontend communicates with two n8n workflows.
(https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/dnvg1l2k2a6duqhu9eqw.png)
The first workflow handles investigation. It prepares the incident, retrieves relevant historical memories from Hindsight, and passes the current incident together with that context to an LLM for analysis.
The second workflow handles the engineer's decision and outcome. Once an engineer reviews the recommendation and records what actually happened, the result is stored back in Hindsight.
The important part is that these aren't two disconnected AI calls.
They form a learning loop.
Current Incident
↓
Hindsight Recall
↓
Historical Experience
↓
AI Investigation
↓
Engineer Review
↓
Actual Outcome
↓
Hindsight Retain
↓
Future Investigations
That last step is the reason memory matters.
A successful incident isn't just closed. Its outcome becomes useful context for the next incident.
Why I used Hindsight
I didn't want to build a simple table of previous incidents and dump the entire database into every LLM prompt.
The useful question isn't:
"Show me every incident we've ever had."
It is:
"Which previous experiences are relevant to this incident?"
That's where Hindsight fits naturally.
Hindsight provides persistent memory through operations such as retain and recall. The system can store durable information and later retrieve relevant context rather than requiring the entire history to be manually supplied to the model.
I created a dedicated Hindsight memory bank for the incident-response system:
incident-response-agent
Historical incidents were stored in that memory, giving the investigation workflow a source of previous engineering experience.
For example, when a new Payment API incident arrives, the investigation workflow can recall previous incidents involving similar symptoms, such as HTTP 500 errors and database connection exhaustion.
The model then has something much more useful than generic knowledge.
It has organizational context.
Recall before reasoning
The most important design decision was putting memory retrieval before the LLM investigation.
The sequence is essentially:
Incident
↓
Prepare incident
↓
Recall relevant memories
↓
Give current incident + memories to LLM
↓
Generate investigation
(https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/a4c0rs9ob6bvto90r1ya.png)
The LLM isn't asked to solve the incident in a vacuum.
It receives the current incident along with relevant historical evidence.
For one of my test incidents, the system received:
json
{
"incident_id": "INC-032",
"service": "Payment API",
"severity": "HIGH",
"description": "Payment transactions are intermittently failing with HTTP 500 responses during heavy traffic. Database connection usage is near its limit."
}
(https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/sovhdpobxt5rrd06vb2q.png)
The recalled memories contained previous Payment API incidents with the same general failure pattern.
The resulting investigation identified database connection pool exhaustion as the likely root cause and connected that conclusion to previous incidents.
That was the behavior I was looking for.
The model wasn't simply saying, "This sounds like database connection exhaustion."
It was reasoning with evidence from previous incidents.
Retain after the engineer makes the decision
Recall is only half of the system.
The more interesting part is what happens after the incident is resolved.
An AI recommendation is not automatically a fact.
An engineer needs to review it.
So the frontend provides an approval interface where the engineer can record:
- Whether the recommendation was approved
- The action that was taken
- The actual outcome
- Engineer feedback
Only after the engineer approves the recommendation do we send the outcome to the second workflow.
The retained memory looks conceptually like this:
Incident
- Recommended action
- Engineer approval
- Actual outcome
- Engineer feedback
The actual Hindsight Retain request is built around this information:
javascript
{
items: [
{
content:
Incident ${incident_id} was investigated by the AI Incident,
Response Agent. The recommended action was: ${action}.
The action was approved by an engineer. The actual outcome
was: ${outcome}. Engineer feedback: ${engineer_feedback}.
context:
"Production incident response outcome and engineering learning",
document_id:
`${incident_id}-outcome`
}
]
}
This distinction matters.
I don't want the system to learn blindly from everything an LLM says.
I want it to learn from outcomes that have gone through an engineer review step.
That makes the memory much more useful operationally.
A concrete example
Consider a database performance incident.
The current incident says that the Order API is experiencing increased latency and request timeouts during peak traffic.
Hindsight retrieves previous incidents where similar database performance problems occurred.
The LLM uses that evidence to recommend investigating slow queries and indexing.
The engineer then checks the actual system.
In this case, slow-query logs and query analysis confirm the database bottleneck. Missing indexes are identified and the affected queries are optimized.
After the fix, the engineer records the actual outcome:
The investigation confirmed that several high-frequency database
queries were causing increased response times. Slow-query logs
identified inefficient queries, and EXPLAIN analysis showed missing
indexes on frequently filtered columns. After adding the required
indexes and optimizing the affected queries, database response times
returned to normal and Order API request timeouts stopped.
That outcome is then retained.
The next time a similar database performance incident occurs, the system has another piece of real operational experience to work with.
That's the part I find most useful about the architecture.
The system isn't just generating recommendations.
It is accumulating evidence about which recommendations actually worked.
The engineering tradeoff: memory needs a write boundary
One of the lessons I learned while building this is that persistent memory isn't automatically useful just because it is persistent.
Bad information retained forever is still bad information.
For incident response, I wanted a clear boundary between an AI suggestion and an engineering-confirmed outcome.
That's why the application has an explicit engineer review stage.
The AI can investigate.
The engineer decides whether the recommendation is appropriate.
The actual result is then captured.
This gives the memory system a much clearer signal:
AI hypothesis
↓
Human validation
↓
Observed outcome
↓
Persistent memory
That separation is important when the memory is going to influence future decisions.
Keeping the architecture simple
I intentionally kept the system modular.
The frontend is responsible for the engineer experience.
n8n handles the workflow orchestration.
The LLM handles incident analysis.
Hindsight handles persistent memory.
The frontend doesn't need to know how memory retrieval works internally. It simply receives an investigation result and provides the engineer with a way to review it.
Likewise, the frontend doesn't contain Hindsight credentials or LLM credentials. Those integrations remain behind the workflow layer.
This made it much easier to iterate.
I could improve the UI without rewriting the memory integration, and I could modify the investigation workflow without changing how the engineer interacts with the application.
What surprised me
The biggest realization was that the useful part of an incident isn't necessarily the incident description itself.
The useful part is often the relationship between:
what happened → what we thought was happening → what we did → what actually worked.
(https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/ht45avvzpmzx8vpq7ym0.png)
That is exactly the kind of information I wanted the agent to remember.
A static incident database can tell me that INC-032 existed.
A memory system can make the experience from INC-032 useful when another incident arrives.
That difference changes how I think about agent memory.
Memory isn't just about making an agent remember conversations.
For operational systems, it can be about preserving decisions, outcomes, and lessons that should influence future reasoning.
What I learned
- Retrieval is more valuable when the query has operational context
"Find similar incidents" is much more useful when the current incident contains the service, severity, symptoms, and observed conditions.
The quality of the memory retrieval depends heavily on the information available at investigation time.
- Don't let the model invent missing operational details
During testing, I noticed an important failure mode: an LLM can sometimes introduce specific configuration values or thresholds that were never present in the evidence.
For an incident-response system, that is dangerous.
The safer approach is to distinguish between what the historical evidence actually says and what the model is inferring.
If the evidence doesn't provide an exact configuration value, the recommendation should not invent one.
- Human feedback is valuable memory
The engineer's feedback isn't just UI metadata.
It can become future context.
A successful remediation plus the engineer's explanation is much more useful than storing only the original alert.
- Memory needs a feedback loop
Recall without Retain gives you historical context.
Retain without Recall gives you a growing archive.
(https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/yfiqffd2pu1gaxbkno5n.png)
Combining them creates something more interesting: accumulated experience that can influence future investigations.
- The goal isn't to replace the engineer
The system is designed to reduce repetitive investigation work while keeping the engineer responsible for the final decision.
The AI proposes.
Historical memory provides context.
The engineer validates.
The outcome becomes future knowledge.
What's next
There are several directions I would take this system next.
I would improve the memory representation of incident outcomes, add stronger evaluation around retrieval quality, and measure whether recalled incidents actually improve investigation accuracy over time.
I'd also like to make the distinction between hypothesis, evidence, and confirmed outcome even more explicit in the UI.
The long-term goal isn't to build an agent that simply gives increasingly confident answers.
It's to build one that becomes more grounded in the engineering experience accumulated around the systems it operates.
That's what made Hindsight particularly useful for this project.
It gave me a practical way to turn past incident experience into persistent context—and, more importantly, to feed confirmed outcomes back into that context so the next investigation has something new to learn from.
Resources
Incident Learning & Response Agent — GitHub
(https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/x3krqr73uew22dp826tk.png)
(https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/i04yaskc2ya7cqfk8xwe.png)
(https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/3dzuu7lq7ajlg96v4gzj.png)
(https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/1g64m6pnu2a1g1a22nrn.png)
(https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/b89lfazf52c4eizfo9by.png)
(https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/98lx4sd45eij5jaqa1q6.png)
Top comments (0)