The third time I watched an on-call engineer rediscover that a connection pool was the culprit, I stopped blaming the engineer and started blaming the tooling. The fix was written down. It was in a postmortem, in a Slack thread, and in someone's head. None of them were in front of the person paging at 3 a.m.
I built Incident Response Agent to close that gap. This is a write-up of how it works, and mostly of the one decision that mattered: treating memory as a first-class part of the system instead of a feature bolted onto an LLM prompt.
What the system does
Incident Response Agent is a dashboard and API for recording, triaging, investigating, and resolving service incidents. The backend is FastAPI with SQLAlchemy over SQLite. The frontend is React 18 with Vite and Recharts, polling every 15 seconds. An incident moves through OPEN, INVESTIGATING, CONTAINED, and RESOLVED, with severities from SEV-1 to SEV-4.
Everything an operator does is a POST to a single action endpoint: investigate, start_response, assign, contain, resolve, add_note, or escalate. Each one appends a row to an append-only timeline. Three more endpoints do the "agent" work: /investigate, /recommend, and /postmortem.
That's the boring part, and I want it to stay boring. The interesting part is what the agent knows when you hit /investigate.
The through-line: an LLM without memory is a confident stranger
My first version of /investigate sent the incident to an LLM and returned what came back. The output was plausible and generic: check dependency health, look at the last deploy, consider resource saturation. That is correct advice for roughly every incident ever filed, which makes it worthless advice for this one.
What an on-call engineer actually wants is: have we seen this before, on this service, and what fixed it? That's a memory problem, not a reasoning problem. So I split the agent into two parts with a hard boundary between them:
-
Analysis (
LLMService): stateless, replaceable, allowed to fail. -
Memory (
HindsightService): the part that makes the analysis specific to my systems.
The orchestration is small enough to read in one sitting:
async def investigate(self, incident_data: dict) -> dict:
analysis = await self.llm_service.analyze_incident(incident_data)
query = f"{incident_data.get('service', '')} {incident_data.get('symptoms', '')} {incident_data.get('error_logs', '')}"
similar = self.hindsight_service.recall_relevant_incidents(query, limit=5)
return {
"incident_summary": analysis.get("summary", "Investigation started."),
"possible_root_causes": analysis.get("possible_root_causes", []),
"historical_matches": similar,
...
}
The query is deliberately dumb: service, symptoms, and raw error logs concatenated. I did not want a clever query-rewriting step that I'd have to debug during an outage. All of the cleverness lives behind recall_relevant_incidents.
Where the first memory layer fell short
The first implementation of that method was a keyword scorer over the incidents table. I'm including it because it's a good example of something that works right up until it matters:
score = 0
if incident.service.lower() in query_lower:
score += 3
if incident.runbook_used and incident.runbook_used.lower() in query_lower:
score += 2
for term in query_lower.split():
if term and term in text:
score += 1
It's fast and transparent, and it fails in exactly the way you'd expect. An alert that says HTTP 503 and a past incident that says "upstream unavailable, pool exhausted" share almost no tokens. The scorer also has no notion of time, so a two-year-old incident on a since-rewritten service ranks the same as last month's. And splitting the query on whitespace means a pasted stack trace produces hundreds of terms that each match something.
I wanted retrieval that understood that "503s after the rollout" and "connections stuck in CLOSE_WAIT following deploy 4.12" are probably the same story. That's what pushed me to agent memory as a distinct layer instead of a fancier LIKE query.
Swapping in Hindsight
I chose Hindsight, the open-source memory system from Vectorize, largely because its interface is small. The Hindsight documentation describes three operations that mattered to me: retain to store something, recall to search, and reflect to synthesize an answer over what's stored.
Memories live in a "bank," which I use to scope everything to one team's incident history.
Because I'd kept the memory boundary in one class from day one, the swap was a contained change. The adapter now looks like this:
from hindsight_client import Hindsight
class HindsightService:
def __init__(self, db=None):
self.settings = get_settings()
self.client = Hindsight(
base_url=self.settings.hindsight_api_url,
api_key=self.settings.hindsight_api_key,
)
def retain_resolution(self, incident) -> None:
content = (
f"{incident.service} incident {incident.incident_id}: {incident.title}. "
f"Symptoms: {incident.symptoms}. Root cause: {incident.root_cause}. "
f"Resolution: {incident.resolution}. Runbook: {incident.runbook_used}. "
f"Lessons: {incident.lessons_learned}"
)
self.client.retain(
bank_id=self.settings.hindsight_bank_id,
content=content,
document_id=incident.incident_id,
context="resolved production incident",
tags=[f"service:{incident.service}", incident.severity],
timestamp=incident.resolved_at,
)
Two details here are worth calling out.
document_id is the incident ID. Retaining the same incident twice updates it rather than duplicating it. That matters because an incident's story changes: the root cause you write at 2 a.m. is often wrong by Friday, and I want the corrected version to replace the first one.
I retain on resolution, not on creation. An open incident is a hypothesis. A resolved one, with a root cause, a resolution, and a lesson, is knowledge.
Memory has to be allowed to fail
Memory is useful, but it cannot become a single point of failure for incident response.
If Hindsight is unavailable, the incident system still needs to work. The agent can fall back to the underlying incident data and clearly report that historical context was unavailable.
That distinction matters operationally. "I found no similar incidents" and "I couldn't access incident memory" are completely different answers.
I also wanted provenance in the response. When the agent finds a historical match, it returns the incident ID alongside the retrieved text. An operator can inspect the original incident instead of treating generated text as ground truth.
What changed after adding memory
The useful difference isn't that the agent suddenly became smarter.
Without memory, the response to a recurring connection-pool failure was generic: check the database, inspect recent deployments, look at resource utilization.
With memory, the agent could surface a previous incident involving the same service, similar symptoms, and a known resolution. That gives the operator a concrete place to start and, more importantly, something they can verify.
The same memory also feeds the /postmortem flow. Resolved incidents become reusable operational knowledge instead of ending as isolated records.
The principle is simple: LLMs generate analysis; memory supplies organizational context.
Lessons I would reuse
1. Put memory behind an interface you can replace.
I didn't want the rest of the application to know how retrieval worked. That made changing the implementation much smaller.
2. Decide what a "memory" is before deciding how to search it.
For this system, a resolved incident is the useful unit because it contains symptoms, root cause, resolution, and lessons.
3. Make the write path a byproduct of existing work.
Resolution already happens in the incident workflow. Using that event to retain memory avoids creating another maintenance task.
4. Surface provenance.
An agent recommendation is more useful when an operator can see which historical incident informed it.
5. Degrade loudly.
If memory is unavailable, the system should say so. A missing memory result should never look like proof that no history exists.
The behavior I care about is the difference between those two branches. With no history, the agent tells you it's guessing. With history, it tells you which past incidents it's leaning on, so you can open them and disagree.
The recommendation is never "trust me"; it's "here is the evidence, and here is what I'd check first."




Top comments (0)