My Incident Agent Gets Faster With Hindsight
The first time an incident happens, an AI agent has to reason about it.
The second time, I want it to remember.
That simple distinction became the central idea behind my incident-response agent. Instead of treating every production incident as an isolated question for an LLM, I built the system around persistent memory using Hindsight.
The result is a three-path incident workflow: diagnose genuinely new failures, recall strong matches from previous incidents, and adapt partial matches when the failure pattern is familiar but the service is different.
The problem with starting from zero
Incident response has a frustrating property: the same kinds of failures happen repeatedly.
A database connection pool gets exhausted. A service starts throwing out-of-memory errors. A bad environment variable breaks a deployment. A downstream dependency starts timing out and causes failures elsewhere.
An LLM can reason about each of these incidents, but reasoning from scratch every time is wasteful.
More importantly, useful incident knowledge isn't only in generic documentation. It is in the organization's own experience:
- What symptoms appeared?
- Which service was affected?
- What root cause was discovered?
- What remediation actually worked?
- Did the same pattern appear somewhere else later?
I wanted the agent to accumulate that experience.
That's where Hindsight became the memory layer.
The three paths
I designed the incident handler around three outcomes:
Incoming incident
|
v
Hindsight recall
|
+---+---+
| |
strong partial
match match
| |
v v
recall adapt
| |
+---+---+
|
no match
|
v
new diagnosis
|
v
retain memory
A strong match means the current incident is sufficiently similar to something the agent already knows.
A partial match means the agent found something useful, but it isn't safe to simply copy the old resolution.
No match means the incident needs fresh diagnosis.
That distinction matters.
I didn't want "memory" to become a fancy way of saying "copy the nearest previous answer."
The first incident teaches the system
Suppose the agent receives a database incident:
Service: payments-api
Error: DBConnectionPoolExhausted
Symptoms: requests are timing out and new database
connections cannot be acquired
If there is no useful prior memory, the agent takes the slower diagnosis path. It reasons about the symptoms, identifies a likely root cause, proposes remediation, and retains the resulting incident knowledge.
In the actual agent, the retained incident contains fields such as the service, symptom, error signature, root cause, fix, and timestamp. The Hindsight client then stores that incident for future recall.
The important part isn't the storage operation itself.
The important part is that the diagnosis becomes future context.
The second incident is different
Now the same type of problem happens again:
Service: payments-api
Error: DBConnectionPoolExhausted
Symptoms: connection acquisition failures during traffic spike
Instead of immediately asking the LLM to solve the problem from scratch, the agent recalls incident memory.
The core implementation is:
query = f"{service} {error_sig} {symptom}".strip()
raw_recall = self.hindsight.recall(query=query)
matches = self.hindsight.recall_incident_matches(
query=query,
response=raw_recall,
)
match_classification, top_match, relevant_matches = (
self.evaluate_memory_matches(matches, service)
)
The agent then classifies the retrieved memory.
If the match is strong enough and belongs to the same service, it enters the fast path. The agent first attempts to extract the root cause and proven fix directly from the recalled memory before falling back to the LLM if necessary.
That's the behavior I wanted to see:
First occurrence: reason.
Second occurrence: remember.
In the dashboard, the paths are deliberately visible:
π New diagnosisπ§ Recalled from memory𧬠Pattern adapted from memory
That makes the memory behavior observable instead of hiding it inside a prompt.
Memory doesn't mean exact duplication
The more interesting case happens when the service changes.
Imagine the original incident occurred in payments-api.
Later, orders-api experiences a similar database pool failure.
The strings aren't identical. The service is different. The surrounding symptoms may be slightly different.
A naΓ―ve retrieval system could either miss the connection entirely or blindly copy the old fix.
My agent instead treats the recalled incident as a pattern.
The current implementation first performs a service-aware recall. If that produces no usable match, it performs a second recall based on the failure family β the error signature and symptoms without the service name.
if match_classification == "no_match":
family_query = f"{error_sig} {symptom}".strip()
if family_query and family_query != query:
fb_raw_recall = self.hindsight.recall(query=family_query)
fb_matches = self.hindsight.recall_incident_matches(
query=family_query,
response=fb_raw_recall,
)
fb_classification, fb_top_match, fb_relevant = (
self.evaluate_memory_matches(fb_matches, service)
)
If that memory is relevant but not a strong same-service match, the incident enters the pattern-adaptation path.
This is where persistent memory becomes more interesting than a static knowledge base.
The agent isn't just asking:
"Have I seen this exact incident?"
It is asking:
"Have I seen something that helps explain this incident?"
I tested the learning loop
I didn't want this behavior to exist only in the code. I created a three-event verification test.
Test A submits a genuinely new incident. The agent takes the slow path, diagnoses it, and retains it.
Test B submits the same failure again for the same service. The agent recalls the previous incident and takes the fast path.
Test C submits the same failure family to a different service. The agent retrieves the earlier incident, classifies it as a partial match, and generates a service-specific adaptation.
The latest verification produced:
Test A β slow_path
no_match
retained to memory
Test B β fast_path
strong_match
recalled from memory
Test C β pattern_adapted_path
partial_match
adapted from memory
The counters were:
Resolved via Memory: 1
Patterns Adapted: 1
Novel Diagnosed: 1
Total Processed: 3/3
The cross-service test also referenced the memory created by Test A while processing the different service. That was the important result: the system wasn't just recognizing an identical incident; it was using a previous incident as a pattern for a new service.
What changed after adding memory
Before persistent memory, the workflow effectively looked like:
Incident β LLM β diagnosis
After adding Hindsight:
Incident
β
Recall
β
Strong match ββββββ reuse relevant experience
β
Partial match βββββ adapt relevant experience
β
No match ββββββββββ diagnose
β
Retain
That final arrow is the part I consider essential.
If the agent only recalled memories, it would be a retrieval system.
Because it also retains new incident experience, the system has a learning loop.
The engineering lesson
The biggest lesson for me was that adding memory isn't primarily a storage problem.
It's a behavior-design problem.
I had to decide:
- What counts as a strong match?
- When is a partial match useful?
- When should the system stop trusting memory?
- What information should become future memory?
- How do I make the path visible to the operator?
Hindsight provides the persistent memory layer. The incident agent defines how those memories affect behavior.
That separation was important to the design. Memory is useful only when it changes what the agent does with the next incident.
What I would improve next
The current prototype uses a shared organizational memory model. That works for demonstrating collective incident knowledge, but a real multi-tenant deployment would need stronger isolation boundaries.
I'd introduce organization- or workspace-scoped memory banks and preserve operator identity as metadata. That would let engineers within the same organization benefit from shared experience without allowing unrelated organizations to recall each other's incidents.
That's the next layer of the system.
The important part is already working, though:
An incident can become experience, and that experience can change how the next incident is handled.
That's what I wanted from agent memory.





Top comments (0)