I asked an LLM why a checkout service started returning 500s after a 06:31 deploy. It told me the cause was a nil pointer dereference in order/processor.go:112, backed by log snippets it claimed showed the panic. Then it recommended rolling back the deployment immediately.
I hadn't given it logs. I hadn't given it a Helm history. It made up both, wrote a confident root cause around them, and told me to roll back immediately.
This post is about how I built an incident-response agent to stop that from happening, what changed when I gave it memory of real postmortems, and how much (or how little) I can actually prove.
What I built
The agent takes a plain-text incident description and returns a diagnosis: root cause, fix, severity, confidence, and explicit "DO NOT do X" warnings when past incidents show that a tempting fix made things worse. I call those fixes trap actions. A rollback that re-triggers the failure. A restart that wipes the state you need to recover. Scaling up a service whose real problem is a saturated dependency.
The stack is small:
- Memory: Hindsight Cloud (GitHub, docs)
-
LLM: Groq, running
openai/gpt-oss-120b - Backend: FastAPI
- UI: vanilla HTML/CSS/JS
- Data: real postmortems from the OpenSRE dataset, covering Slack, Cloudflare, GitHub, AWS, Datadog, CircleCI, and LaunchDarkly
The dataset has 114 incidents. I seeded a fresh memory bank with 104 of them and held 10 out for evaluation. Hindsight turned those 104 postmortems into 759 world facts, 5 experiences, and 182 observations, connected by 7,135 links.
I picked Hindsight over plain vector search because it doesn't just store chunks of text. It extracts facts from what you retain and links them, which lets you ask a question across incidents instead of only fetching the nearest paragraph. If agent memory is a new idea to you, Vectorize's primer is a decent starting point. I haven't benchmarked it against a plain vector store, so I'm not claiming it's better, only that it's what I used.
The pipeline
Every request goes through the same five steps: extract a signature, reflect, recall, build the prompt, generate.
def analyze_incident(incident_description: str):
sig = extract_signature(incident_description)
if sig['service'] or sig['error_type']:
parts = []
if sig['service']:
parts.append(f"service: {sig['service']}")
if sig['error_type']:
parts.append(f"type: {sig['error_type']}")
enriched_query = f"{incident_description} [{', '.join(parts)}]"
else:
enriched_query = incident_description
reflection = reflect_on_incidents(enriched_query) # Hindsight synthesizes
scored_memories = recall_with_scores(enriched_query) # raw memories, ranked
Two ways to ask memory
The interesting part is that I query memory twice, because the two calls answer different questions.
reflect asks Hindsight to reason over what it knows and return a synthesized answer. That's where cross-incident patterns show up: "this resembles a dependency-capacity failure, and in similar cases a rollback made it worse." No single postmortem says that. It comes from combining several.
recall returns the underlying memories. That gives the model concrete evidence to cite instead of only a summary to trust.
Both get pasted into the system prompt:
memory_context = f"HINDSIGHT REFLECTION:\n{reflection}\n\n"
if scored_memories:
memory_context += "RANKED MEMORIES:\n"
for item in scored_memories:
memory_context += f"- [score: {item['score']}] {item['text']}\n"
The ranking is a keyword heuristic
I re-rank recalled memories before they hit the prompt. I want to be upfront that this is crude:
score = 0.5
matches = sum(1 for term in query_terms if term in text_lower)
score += min(matches * 0.1, 0.3) # term overlap, capped
if "trap" in text_lower:
score += 0.2 # surface trap actions first
There's no embedding math here. It's term overlap (with stopwords filtered out) plus a bonus when a memory mentions a trap. The bonus exists for one reason: memories that describe trap actions are the ones that change what the agent tells you to do, so I want them near the top of the prompt. Hindsight already handles semantic relevance upstream. This layer only nudges the ordering, and it only fires if the memory text contains the literal word "trap".
The signature step is equally plain. It's a list of word-boundary regexes, like \bredis\b or \btgw\b|\btransit.?gateway\b, that tag the incident with a service and an error type. Those tags get appended to the query. I've checked it on a handful of examples, but I haven't measured whether it improves retrieval, so I'm not going to say it does.
Making the warnings mandatory
The last piece is one paragraph in the system prompt:
text
CRITICAL - TRAP ACTION AWARENESS:
If past incidents mention trap actions (fixes that made things worse),
you MUST explicitly warn against them with "DO NOT do X".
Then it's a plain Groq completion, and a regex pulls High, Medium, or Low out of the response. If the regex finds nothing, it returns "Unknown". That confidence value is the model's own claim about itself. It isn't calibrated, and you shouldn't treat it as a probability.
The worked example
The scenario: a checkout service is returning 500s on roughly 12% of requests after a deploy at 06:31, and the question is whether to roll back. I ran it through the same model twice, once with the memory block and once without.
Without memory, I got the fabrication described above: a nil pointer dereference in order/processor.go:112, fake log snippets it claimed to have read, and a confident rollback recommendation. It offered kubectl rollout undo and helm rollback as the primary fixes, and told me to open a Jira ticket. It read like a competent engineer's report. Every specific in it was invented.
The no-memory baseline invented a nil pointer dereference and fake log snippets, then recommended a rollback. None of it was provided.
That is the failure I care about. A wrong guess is normal. What's dangerous is a wrong guess dressed in fake evidence, because it looks exactly like a right answer and it's the kind of thing an on-call engineer at 6 a.m. might act on.
With memory, the agent diagnosed a likely dependency-capacity problem, such as Redis connection-pool exhaustion or a database connection-limit breach, triggered by the deployment. It pointed to similar past incidents involving pool exhaustion and Redis maxclients limits, and it told me what to look for: ERR max number of clients reached on Redis, too many connections on Postgres.
And it gave me the warning I was after, in its own section: DO NOT roll back. It explained that the memory base flags rollback as a trap for this class of failure, because rolling back reintroduces the same configuration without freeing the exhausted resource, and it named the exact command to avoid.
The memory-powered agent warns against rollback with a specific reason, citing past incidents.
I'd read this answer carefully too. It's still a hypothesis. It says "likely," and it tells the engineer to confirm in logs and metrics before changing anything. It also contained some noise: a suggestion to check for BGP or systemd-networkd changes around 06:30, which reads like bleed-through from unrelated incidents in the bank. And rollback isn't universally wrong. A model with no deployment context suggesting one is being reasonable. The problem with the baseline wasn't the rollback suggestion, it was the invented certainty behind it.
Testing on incidents it had never seen
One good example proves nothing, and my first evaluation was worse than that. I tested the agent on incidents that were also stored in its memory, and it scored a perfect 10/10. That measures lookup. It says almost nothing about how the agent behaves on an outage it hasn't encountered.
So I redid it:
- Held out 10 of the 114 incidents. The bank has only the other 104.
- Wrote symptom-only queries for each held-out incident,without root-cause or postmortem wording.
- Ran each query with memory and again as a baseline,same model, same prompt minus the memory block.
- Graded each answer against the true_category field in the dataset, and saved every output to eval_holdout_results.json.
Condition Root cause correct
With memory 9 / 10
No memory 0 / 10 fully correct (4 partial, 6 hallucinated)
A partial match means the baseline named a plausible cause in the right area — a version skew, a cache issue, a resource limit — but described the wrong mechanism or missed the true trigger.
The tenth memory-backed query hit Groq's daily rate limit and never finished. I'm counting it as a miss rather than reporting 9/9, and I have no result for it either way.
The baseline's failures had a pattern. Given no evidence, it reached for code-level explanations that sound like an engineer's first guess: connection leaks, TTL misconfigurations, missing null checks. With memory, the agent named causes in the right category and grounded them in patterns from similar past incidents. All nine completed memory-backed runs produced a specific root cause matching the true category.
What I learned
Fabrication is a grounding problem, not an intelligence problem. Same model, same question. The only thing that changed was whether it had real material to reason over. When it didn't, it filled the gap with plausible detail. When it did, it stopped inventing.
Negative advice is the valuable part. Anyone can tell you to check the logs. "This fix has backfired on other teams, and here's the mechanism" is institutional memory that gets lost between on-call rotations. Postmortems are full of it and mostly go unread.
Separate synthesis from evidence. reflect gave the model a reasoned starting point and recall gave it specifics to cite. I didn't test either alone, but the combination is what I'd build again.
Your first evaluation is probably flattering. My perfect score was measuring retrieval of documents I had just stored. The hold-out test was less impressive and much more useful.
Save raw outputs. Every number in this post traces to a JSON file. That made me more careful about what I claimed, since anyone can check.
Limitations
I'd rather list these than have someone find them in the comments:
n = 10. One miss moves the score by ten points. This is a small check, not a benchmark.
I graded my own outputs. Single grader, no second opinion, graded against the dataset's category label.
Category-level grading. "Correct" means the right kind of cause (dependency saturation, config push, network fault), not a match with the postmortem's exact sequence of events.
Held-out isn't unrelated. All 114 incidents come from one dataset and a handful of vendors. Outages cluster into recurring failure classes, so a held-out incident can resemble several retained ones. This isn't a test of truly novel failure modes.
No ablation. I don't know how much of the gain comes from reflect versus recall, from the trap boost, or from the signature enrichment.
Heuristics throughout. The ranking is keyword overlap, the signature step takes the first regex that matches, and the service list is hardcoded.
One run per query. No repeated runs, so I can't say how stable the results are.
Self-reported confidence. Not calibrated.
A stronger test would use more incidents, repeated runs, an independent grader, and a held-out set chosen to look different from what's in memory.
Where this leaves me
The claim I'll stand behind is narrow. On 10 incidents absent from its memory, the agent named the right category of root cause 9 times, where the same model without memory got none fully right, and it did so without inventing log output. Whether that holds at scale, on genuinely novel failures, or against a stronger baseline is something I haven't shown.
If you want to try the memory layer yourself, it's open source at github.com/vectorize-io/hindsight, and the docs cover retain, recall, and reflect. Point it at your own postmortems, and I'd be curious what trap actions show up.



Top comments (0)