Every on-call engineer has done the obvious thing during an outage and watched it make things worse. You roll back, and the rollback re-triggers the failure. You restart a pod, and the state you needed for recovery is gone. You scale up a service whose real problem is a saturated dependency, and you pile more load onto the thing that's already failing.
I call these trap actions. After reading a lot of postmortems, I think the most valuable institutional memory an on-call rotation has is knowing which tempting fixes not to try. So I built an agent around that idea.
Why "don't" is worth more than "do"
Generic advice is cheap. Any model will tell you to check the logs, look for a regression, and consider a rollback. That's the default, and it's often reasonable.
What a model without history can't tell you is that, for a particular class of failure, the default has burned other teams. That knowledge lives in postmortems, and it tends to leave the building when people rotate off call.
I built the agent to recall real postmortems from a Hindsight memory bank: 104 incidents from Slack, Cloudflare, GitHub, AWS, Datadog, CircleCI, and LaunchDarkly. When retrieved incidents describe a fix that made things worse, the agent has to say so.
Making the agent say it
There are two mechanisms, and both are blunt. First, recalled memories that mention a trap get a bonus in my re-ranking, so they land near the top of the prompt:
python
score = 0.5
matches = sum(1 for term in query_terms if term in text_lower)
score += min(matches * 0.1, 0.3)
if "trap" in text_lower:
score += 0.2 # surface trap actions first
Second, the system prompt makes the warning mandatory:
text
CRITICAL - TRAP ACTION AWARENESS:
If past incidents mention trap actions (fixes that made things worse),
you MUST explicitly warn against them with "DO NOT do X".
Neither is sophisticated. The boost only fires when a memory contains the literal word "trap," and the prompt only works if the retrieved context actually contains something to warn about.
A concrete example
The scenario: a checkout service returning 500s on roughly 12% of requests, starting after a 06:31 deploy. Should we roll back?
Without memory, the model produced a confident and entirely invented story: a NullPointerException in PaymentProcessor.validateOrder(), a new promoCode field, a kubectl logs | grep showing "112 occurrences," and a Helm history showing revision 12 at 06:31. It had no logs and no Helm history. It recommended rolling back immediately.
With memory, the agent diagnosed a likely dependency-capacity problem, such as Redis or database connection-pool exhaustion triggered by the deploy. It gave me things to check (ERR max number of clients reached on Redis, too many connections on Postgres) and then a dedicated section, "Why DO NOT Roll Back the Deployment." Its reasoning was that the memory base flags rollback as a trap for this class of failure, because rolling back reintroduces the same configuration without freeing the exhausted resource. It named the command to avoid: kubectl rollout undo.
The honest reading of this example needs some care. Rollback isn't universally wrong. A model with no deployment context suggesting it is being reasonable, and in plenty of incidents it's the right call. The baseline's real failure was fabricated certainty. And the memory-backed answer is still a hypothesis: it says "likely," reported Medium confidence, and included some noise about BGP and systemd-networkd changes that looks like bleed-through from unrelated incidents.
What institutional memory means here
A postmortem is a team's memory of an incident. The problem is that memory is written down once and rarely consulted at 3 AM. An agent that recalls it at the moment of decision changes when that knowledge gets used.
There's also a subtler point about how the memory is organized. Hindsight doesn't just store the text of each postmortem. It extracts facts and links them, which is why the 104 incidents became 759 world facts, 182 observations, and 7,135 links in my bank. A "trap" pattern can emerge from several incidents together, which is what the reflect call is for. The docs and Vectorize's primer on agent memory explain the model in more detail.
How well did it work?
I held out 10 of the 114 incidents, kept 104 in memory, and ran symptom-only queries with and without memory. With memory, 9 of 10 root causes matched the true category (one run hit a rate limit and counts as a miss). Without memory, 0 of 10 were fully correct: 4 partial, 6 hallucinated.
Note that this measures root-cause categories, not trap warnings specifically. I haven't separately measured how often the agent warns about a trap that's actually present, or how often it warns about one that isn't.
Limitations
• n = 10, graded by me against the dataset's true_category label.
• No measurement of trap-warning precision. I don't know the false-positive rate.
• The trap boost is a keyword match. If a memory describes a trap without the word "trap," it gets no bonus.
• Held-out isn't unrelated. Outages cluster into recurring failure classes, so a held-out incident can resemble retained ones.
• Warnings can mislead. A trap in past incidents isn't necessarily a trap in yours. The agent's "DO NOT" is a prompt to think, not a rule.
Takeaways
- Capture failed fixes, not just successful ones. Postmortems are unusually good at recording what didn't work.
- Put negative advice first. "Don't do X, and here's why" is harder to get from a model that has no history.
- Force the warning in the prompt. Otherwise it's easy for the model to bury it or skip it.
- Treat warnings as hypotheses. Context differs; the agent should say why, so the engineer can judge.
- Measure warnings separately. I didn't, and I'd want precision and recall on them before trusting the feature.

Top comments (0)