How I Built an Incident Response Agent That Remembers
At 3 AM, an incident response system should not behave as if it has never seen an outage before.
That was the problem I wanted to explore: how do we make an incident-response agent use the organization's previous operational experience instead of generating another generic list of troubleshooting steps?
I built an incident-response system around three pieces: a machine-learning model for SLA breach risk, an incident agent for reasoning about the current failure, and Hindsight for retaining and recalling operational memory.
The interesting part is not simply adding an LLM to incident management. It is giving the system a way to remember what happened before.
The problem with starting from zero
A new incident usually contains a mixture of structured information and messy operational context: the affected service, symptoms, priority, category, logs, and urgency.
Incident report input
A stateless assistant can process that information, but it does not automatically have access to the organization's previous incidents and the decisions engineers made during them.
That creates a recurring problem.
An engineer may have already solved a very similar failure months ago, but the next incident starts from scratch unless that knowledge is explicitly retrieved.
I wanted the system to follow a different path:
Current Incident
↓
Understand the current symptoms
↓
Predict SLA breach risk
↓
Recall similar historical incidents
↓
Use previous resolutions as evidence
↓
Produce a response plan
↓
Store the confirmed resolution
The last step is particularly important. A successful incident should become useful context for a future incident rather than disappearing when the ticket is closed.
The architecture
The application uses a FastAPI backend with a browser-based operations dashboard.
The backend separates the major responsibilities:
model_service.py handles the machine-learning pipeline.
hindsight_service.py provides the Hindsight integration.
agent.py coordinates incident reasoning.
app.py exposes the API endpoints and serves the dashboard.
The repository also contains the historical incident data used to seed the Hindsight memory bank and automated tests for the incident flow.
At a high level, the request moves through two different kinds of information.
The first is quantitative:
Incident features
↓
RandomForest
↓
SLA breach probability
The second is experiential:
Current incident
↓
Hindsight recall
↓
Relevant historical incidents
↓
Previous causes + resolutions
The agent can then use both forms of information when constructing its response.
Why Hindsight is the interesting part
The Hindsight integration gives the application an explicit memory layer.
Historical incident information can be retained and later recalled semantically. The project uses Hindsight's recall functionality to retrieve relevant incident memories rather than relying only on exact keyword matches.
The basic idea is straightforward:
memories = await hindsight.arecall(
bank_id=BANK_ID,
query=incident_context
)
Institutional memory retrieved by Hindsight
The returned memories become context for the incident analysis.
This changes the behavior of the system.
Instead of asking only:
"What could cause these symptoms?"
the agent can effectively ask:
"Have we seen something similar before, and what happened when we did?"
That distinction matters in operational systems because previous incidents often contain information that isn't obvious from the current symptoms alone.
For example, the repository includes historical scenarios involving database connection pools, authentication certificate problems, and Kubernetes JVM memory failures. Those memories can be recalled when a new incident resembles them.
Memory is useful only if the system can learn
Retrieval alone isn't enough.
If the system can read old incidents but cannot remember newly confirmed solutions, the memory remains static.
That is why the project includes a closed-loop flow.
After an incident has been resolved, the confirmed root cause and remediation can be retained:
await hindsight.aretain(
bank_id=BANK_ID,
content=resolution
)
A later incident can then retrieve that experience.
The intended cycle is:
Incident A
↓
Diagnose
↓
Resolve
↓
Retain resolution
↓
Incident B
↓
Recall Incident A
↓
Use its resolution as historical context
This is a much more useful model of organizational memory than simply storing a conversation transcript.
Why I kept the ML model separate from memory
Another design decision was to avoid treating historical memory as a replacement for quantitative prediction.
The project uses a Scikit-learn RandomForestClassifier pipeline to estimate SLA breach risk from structured incident attributes.
The repository documents the model as being trained on ServiceNow incident data, with categorical dimensions including categories, subcategories, and assignment groups.
That produces a different kind of signal from Hindsight.
The ML model answers:
"Based on the structured characteristics of this incident, how much SLA risk is associated with it?"
Hindsight answers:
"What relevant operational experiences do we already have?"
Those are complementary questions.
A probability score by itself doesn't explain how engineers should respond. A historical memory by itself doesn't necessarily quantify the urgency of the current ticket.
Combining them gives the incident agent more context.
SLA breach risk prediction
The incident flow
A user begins with the incident intake interface.
The current service, symptoms, category, priority, and other incident information are submitted to the backend through the analysis endpoint.
The backend then performs the relevant processing.
Conceptually:
POST /api/analyze
│
├── Extract incident features
│
├── Calculate SLA risk
│
├── Recall historical memories
│
└── Generate incident analysis
Incident response dashboard
The dashboard can then present the resulting risk information, recalled historical context, and response plan.
Safe verification workflow
The repository also includes a resolution endpoint for the closed learning loop:
POST /api/resolve
↓
Confirmed resolution
↓
Hindsight retain()
Incident analysis and performance
This separation makes the workflow easier to reason about: analysis and learning are related, but they are not the same operation.
What surprised me while building it
The hardest conceptual part wasn't getting an API endpoint to return a response.
It was deciding what the agent should actually remember.
Incident memory becomes useful when it contains operationally meaningful information: what failed, what evidence pointed toward the cause, what remediation was applied, and whether that remediation was confirmed.
Simply dumping every interaction into memory would create a large collection of text without necessarily creating useful operational knowledge.
That led to a principle I kept coming back to:
Memory should capture experience, not just conversation.
The distinction is important for future incident retrieval.
If an engineer previously solved a JVM memory problem by changing a container-aware configuration, that remediation is potentially useful to another incident with similar symptoms.
The useful memory isn't the fact that someone had a conversation about the incident. It is the operational knowledge contained in the resolution.
What I learned
- Memory and reasoning solve different problems An agent can reason about a current incident, but reasoning does not automatically give it organizational experience.
Memory provides the missing historical context.
- Retrieval needs meaningful information Good recall depends on what is retained.
Incident IDs alone aren't useful memories. Root causes, symptoms, services, remediation steps, and verification results provide much more useful context.
- Quantitative and qualitative signals complement each other The ML model provides a quantitative SLA-risk signal.
Hindsight provides historical operational context.
Neither needs to replace the other.
- Closed-loop learning changes the system over time A resolved incident should not simply disappear.
Once a confirmed resolution is retained, it can become evidence for future incidents.
That is the foundation of an incident-response system that can accumulate operational knowledge.
Where this goes next
The current project is an engineering implementation rather than the end of the problem.
The next level is making the memory lifecycle increasingly reliable: deciding what should be retained, validating resolutions before they become reusable knowledge, improving retrieval quality, and measuring whether recalled incidents actually improve response outcomes.
That is the part I find most interesting.
The goal isn't to build an assistant that produces more text when an incident occurs.
The goal is to build one that can say:
"I've seen something like this before. Here's what happened, here's what we learned, and here's the evidence that makes that experience relevant now."
That is what makes memory useful in an incident-response agent.






Top comments (0)