When I started building an incident-response agent, the first decision wasn't about models or frameworks. It was about what the agent should remember. I chose real postmortems, and this post explains why, what the data looked like once it was in memory, and what I can and can't claim about it.
Why real incidents
An SRE agent is only useful if its advice matches how outages actually unfold. Real postmortems have properties that are hard to invent: vendor-specific details, tangled causal chains, and a habit of recording what the responders tried that didn't work.
That last one is the reason for the whole project. Some of the most instructive incidents are the ones where the obvious fix made things worse. A rollback re-triggered the failure. A restart wiped state needed for recovery. Scaling up added load to a saturated dependency. I call these trap actions, and a model with no memory of past outages has no reason to avoid them.
I should be clear about the limits of this argument. I didn't run a comparison between real and hand-written data, so I can't tell you real data scores better by some number. I chose it because I wanted the agent to cite incidents that actually happened, and because I didn't trust myself to invent believable failure chains.
The dataset
I used the OpenSRE incident dataset: 114 real postmortems covering Slack, Cloudflare, GitHub, AWS, Datadog, CircleCI, and LaunchDarkly. Each incident has a true_category label for its root cause, which I used later as the answer key for grading.
I retained 104 incidents into a Hindsight memory bank (GitHub, docs) and held out the other 10. The function that stores each one is deliberately boring:
python
1. - def store_incident(content: str):
2. - client.retain(bank_id=BANK_ID, content=content)
3. - The seeding log looked like this:
4. - text
5. - [15/104] Stored: LaunchDarkly - launchdarkly_legacy_routing_cold_cache
What memory did with them
I didn't index documents. Hindsight extracts facts from what you retain and derives observations and links between them. The 104 postmortems became 759 world facts, 5 experiences, and 182 observations (946 memories total), connected by 7,135 links. If you want the conceptual background, Vectorize's agent memory primer is a good start.
The practical consequence is that a query about a checkout service returning 500s can surface a pattern drawn from several unrelated vendors' incidents, not just one document that mentions the word "checkout."
Making trap actions matter
Real postmortems record failed remediations, but a model won't necessarily surface them on its own. I did two things.
First, I re-rank recalled memories with a small heuristic that boosts anything mentioning a trap:
python
score = 0.5
matches = sum(1 for term in query_terms if term in text_lower)
score += min(matches * 0.1, 0.3)
if "trap" in text_lower:
score += 0.2 # surface trap actions first
Second, the system prompt requires the model to say "DO NOT do X" whenever the retrieved context mentions a trap. The boost is crude and only fires when a memory literally contains the word "trap," so I'd treat it as a nudge, not a guarantee.
What it looked like in practice
Same model, same query: a checkout service returning 500s on roughly 12% of requests after a 06:31 deploy, and the question of whether to roll back.
Without memory, the model invented a NullPointerException, a new promoCode field, "112 occurrences" from a kubectl logs command it never ran, and a Helm revision that didn't exist. It recommended an immediate rollback.
With memory, it diagnosed a likely dependency-capacity issue (Redis or database connection-pool exhaustion), pointed to similar past pool-exhaustion incidents, and stated that the memory base flags rollback as a trap for this failure class, naming kubectl rollout undo as the command to avoid.
The memory-backed answer wasn't clean. It also suggested checking BGP and systemd-networkd changes, which read like bleed-through from unrelated incidents. That's a reminder that retrieval over a mixed dataset brings noise along with signal.
The held-out result
I held out 10 incidents, wrote symptom-only queries for them, and compared with-memory to no-memory using the same model and prompt minus the memory block. With memory: 9 of 10 root causes matched the true category (one run hit a rate limit and counts as a miss). Without memory: 0 of 10 fully correct, 4 partial, 6 hallucinated.
Limitations of this dataset choice
• It's one dataset from a handful of vendors. Outages cluster into recurring failure classes, so a held-out incident can resemble several retained ones. This doesn't test truly novel failure modes.
• n = 10, self-graded. I graded against true_category, with no second grader.
• Category-level grading. "Correct" means the right kind of cause, not a reproduction of the postmortem.
• Postmortems are written after the fact. They're curated narratives, and a live incident is messier than the symptom text I wrote for the tests.
• No comparison against other data sources.
Takeaways
- Choose data that records failed fixes . The most valuable content in a postmortem is often what the responders tried that didn't work.
- Hold some of it out before you seed. Decide the split before you retain anything, or you'll be tempted to test on what the agent has already seen.
- Keep a label you can grade against. true_category turned my evaluation from vibes into counts.
- Expect retrieval noise. Mixed data brings unrelated suggestions along, and the agent's answer should be read as a hypothesis.
- Say what you didn't compare. I have no measured comparison against other data, so I don't claim one.

Top comments (0)