At 2 AM, when a database connection pool exhausts itself for the third time this quarter, the last thing anyone wants is an AI assistant that gives the same generic advice it gave the first two times — including the fix that already failed twice.
That's the problem I set out to fix with IncidentMind, an on-call incident response agent that doesn't just analyze what's happening right now — it remembers every incident it has seen before, and critically, it remembers which fixes worked and which ones didn't.
What it actually does
IncidentMind takes an incident alert and error log as input — the kind of thing that lands in a Slack channel or PagerDuty ping — and returns a root cause analysis, a ranked list of fix steps, and an estimated urgency and resolution time. Nothing unusual so far; plenty of tools do that with a single LLM call.
What makes it different is what happens before that LLM call. Every incoming incident first gets run through Hindsight, a persistent memory layer, which searches for similar past incidents and — this is the part I care about — searches separately for records of which fixes worked and which ones failed. Both sets of results get folded into the prompt before the model ever sees the current incident.
The practical effect: if a database connection pool exhausted itself last month and the team's first instinct — restart the service — didn't fix it, the agent doesn't suggest a restart again. It says so explicitly, citing the incident it's remembering from.
The core technical story: retain, then recall, then validate
The architecture is deliberately boring in the right places. A Streamlit front end collects the incident text. A thin wrapper module talks to Hindsight's API for retain and recall calls. Groq serves the LLM inference (openai/gpt-oss-120b, with a fallback model if the primary is rate-limited).
The interesting part is the memory schema. I store two separate kinds of memory per incident, not one:
# The incident itself — what happened, and how it was ultimately resolved
client.retain(
bank_id="incidentmind-v1",
content="INCIDENT MEMORY\nService: payments-api\nRoot cause: DB pool exhaustion...",
document_id="incident-INC-001",
metadata={"service": "payments-api", "outcome": "worked"},
tags=["incident", "payments-api", "worked"],
)
# The fix outcome — tracked independently, so a failed first attempt
# is recorded even if the incident was eventually resolved differently
client.retain(
bank_id="incidentmind-v1",
content="FIX OUTCOME: 'Immediate service restart' FAILED for payments-api. Do NOT recommend again.",
document_id="outcome-INC-001-failed",
tags=["outcome", "payments-api", "failed", "fix-result"],
)
Splitting these into two memory types turned out to matter a lot. An incident's final resolution and its failed first attempts are different signals, and conflating them means the agent either forgets that a plausible-looking fix already failed, or it can't distinguish "this fix worked" from "this is just what happened during the incident." Recall queries can target either or both:
incidents = client.recall(bank_id="incidentmind-v1", query=incident_text, max_tokens=8000)
outcomes = client.recall(
bank_id="incidentmind-v1",
query=incident_text,
tags=["outcome", "payments-api"],
tags_match="all",
)
The tags_match="all" detail cost me a debugging session. My first pass used "any", which meant a query for outcome memories on payments-api also pulled back unrelated outcome memories from other services, because it matched on either tag rather than both. Precision recall matters more than I expected going in — a noisy memory context doesn't just fail to help, it actively degrades the model's confidence in citing anything specific.
Where I had to add a safety net
Language models will cite sources that sound plausible even when they weren't given them. Early in testing, I watched the agent confidently cite an incident ID that didn't exist anywhere in the recalled memories — presumably a pattern match on similar-looking IDs elsewhere in its training. That's a real problem for a tool whose entire value proposition is "trust what I'm telling you, because I remember."
The fix was a post-parse validation step: after the LLM returns its structured response, I strip any cited incident ID that isn't actually present in the set of memories that were recalled for that specific query.
def _strip_unverified_citations(parsed_response, recalled_ids):
valid_ids = set(recalled_ids)
parsed_response["fix_steps"] = [
step for step in parsed_response.get("fix_steps", [])
if not step.get("incident_refs") or
all(ref in valid_ids for ref in step["incident_refs"])
]
return parsed_response
It's a small function, but it's the difference between a demo that looks impressive and a tool I'd actually trust during a real incident. Every citation in the UI now traces back to something that was genuinely retrieved from Hindsight's memory layer, not something the model invented.
What it looks like in practice
Here's a concrete before/after. I fed the same incident — a Kafka consumer group falling 600,000 messages behind, rebalancing repeatedly — through the agent twice: once with memory disabled, once with it enabled.
Without memory: confidence capped around 35%, generic advice — check consumer lag, review max.poll.records, consider scaling consumers. Reasonable, but the kind of thing you could find in any runbook.
With memory: confidence around 85%, and the agent recalled two prior incidents on the same service with the same rebalancing pattern. It ranked its fix steps differently based on what had actually resolved those prior incidents, and it flagged that an immediate consumer restart had been tried before and hadn't helped — because a prior fix attempt on this exact failure pattern was stored as a failed outcome, and recall surfaced it.
That's the whole pitch, compressed into one interaction: the same failure mode gets progressively cheaper to diagnose the more times it happens, instead of costing the same investigation time every single time.
The part that surprised me: teaching it live
The before/after comparison above is the obvious demo. The thing I didn't expect to find compelling was how the system behaves the moment after you correct it.
If the agent suggests a fix and you mark it as failed, that outcome gets retained immediately — no re-indexing step, no batch job, no restart. Run the exact same incident through again a few seconds later, and the agent has already updated: the fix you just marked as failed shows up under a "skipped" section, citing the specific incident you just created, with an explicit note that it was avoided because it didn't work last time.
It's a small thing mechanically — one write, then one read against fresh data — but it changes how the tool feels to use. Most "AI learns from feedback" claims mean the next training run will incorporate your correction. Here, the next query, seconds later, already reflects it. That immediacy is entirely a property of using retrieval over a memory store rather than anything resembling fine-tuning, and it's the single feature that made this feel less like a chatbot with a search index bolted on and more like something that's actually paying attention.
Lessons learned
Retrieval precision beats retrieval volume. Pulling back more memories doesn't make the agent smarter — it makes the prompt noisier. Tight tag filtering on retrieval mattered more than raising the token budget.
Never trust a model's citations without checking them against ground truth. This isn't specific to memory systems, but it bites harder here, because the whole point of showing citations is to build trust that erodes instantly the first time one is wrong.
Separate "what happened" from "what worked" as distinct memory types. Treating an incident's narrative and its fix outcome as one blob makes it much harder to reason about which fixes to avoid recommending again.
A memory system needs a way to unlearn, not just accumulate. Retaining a failed fix outcome is only useful if a future recall can specifically surface "this was tried and it didn't work" as its own signal, distinct from the incident record itself.
Confidence should be earned, not assumed. Capping confidence when no relevant memory exists — rather than letting the model produce a falsely confident cold-start answer — turned out to be one of the more important design decisions, because it makes the value of memory visible rather than implicit.
If you're building anything that needs an agent to get better at a recurring task over time rather than starting from zero on every request, a memory layer that can distinguish outcomes — not just events — from each other is worth the extra design effort it takes to set up properly.
Top comments (0)