<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: kudikala saikeerthika</title>
    <description>The latest articles on DEV Community by kudikala saikeerthika (@kudikala_saikeerthika_672).</description>
    <link>https://dev.to/kudikala_saikeerthika_672</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149930%2F735681c5-93d1-44dc-80bf-6c2c8bf9d6de.png</url>
      <title>DEV Community: kudikala saikeerthika</title>
      <link>https://dev.to/kudikala_saikeerthika_672</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kudikala_saikeerthika_672"/>
    <language>en</language>
    <item>
      <title>What 104 Real Postmortems Taught My Agent That I Couldn't Have Written Myself</title>
      <dc:creator>kudikala saikeerthika</dc:creator>
      <pubDate>Tue, 29 Sep 2026 14:40:16 +0000</pubDate>
      <link>https://dev.to/kudikala_saikeerthika_672/what-104-real-postmortems-taught-my-agent-that-i-couldnt-have-written-myself-3n8j</link>
      <guid>https://dev.to/kudikala_saikeerthika_672/what-104-real-postmortems-taught-my-agent-that-i-couldnt-have-written-myself-3n8j</guid>
      <description>&lt;p&gt;When I started building an incident-response agent, the first decision wasn't about models or frameworks. It was about what the agent should remember. I chose real postmortems, and this post explains why, what the data looked like once it was in memory, and what I can and can't claim about it.&lt;br&gt;
&lt;strong&gt;Why real incidents&lt;/strong&gt;&lt;br&gt;
An SRE agent is only useful if its advice matches how outages actually unfold. Real postmortems have properties that are hard to invent: vendor-specific details, tangled causal chains, and a habit of recording what the responders tried that didn't work.&lt;/p&gt;

&lt;p&gt;That last one is the reason for the whole project. Some of the most instructive incidents are the ones where the obvious fix made things worse. A rollback re-triggered the failure. A restart wiped state needed for recovery. Scaling up added load to a saturated dependency. I call these trap actions, and a model with no memory of past outages has no reason to avoid them.&lt;/p&gt;

&lt;p&gt;I should be clear about the limits of this argument. I didn't run a comparison between real and hand-written data, so I can't tell you real data scores better by some number. I chose it because I wanted the agent to cite incidents that actually happened, and because I didn't trust myself to invent believable failure chains.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The dataset&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I used the OpenSRE incident dataset: 114 real postmortems covering Slack, Cloudflare, GitHub, AWS, Datadog, CircleCI, and LaunchDarkly. Each incident has a true_category label for its root cause, which I used later as the answer key for grading.&lt;/p&gt;

&lt;p&gt;I retained 104 incidents into a Hindsight memory bank (GitHub, docs) and held out the other 10. The function that stores each one is deliberately boring:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. - def store_incident(content: str):
2. -     client.retain(bank_id=BANK_ID, content=content)
3. - The seeding log looked like this:
4. - text
5. - [15/104] Stored: LaunchDarkly - launchdarkly_legacy_routing_cold_cache

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;What memory did with them&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I didn't index documents. Hindsight extracts facts from what you retain and derives observations and links between them. The 104 postmortems became 759 world facts, 5 experiences, and 182 observations (946 memories total), connected by 7,135 links. If you want the conceptual background, Vectorize's agent memory primer is a good start.&lt;/p&gt;

&lt;p&gt;The practical consequence is that a query about a checkout service returning 500s can surface a pattern drawn from several unrelated vendors' incidents, not just one document that mentions the word "checkout."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Making trap actions matter&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Real postmortems record failed remediations, but a model won't necessarily surface them on its own. I did two things.&lt;br&gt;
First, I re-rank recalled memories with a small heuristic that boosts anything mentioning a trap:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;score = 0.5
matches = sum(1 for term in query_terms if term in text_lower)
score += min(matches * 0.1, 0.3)
if "trap" in text_lower:
    score += 0.2    # surface trap actions first
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Second, the system prompt requires the model to say "DO NOT do X" whenever the retrieved context mentions a trap. The boost is crude and only fires when a memory literally contains the word "trap," so I'd treat it as a nudge, not a guarantee.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it looked like in practice&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Same model, same query: a checkout service returning 500s on roughly 12% of requests after a 06:31 deploy, and the question of whether to roll back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Without memory&lt;/strong&gt;, the model invented a NullPointerException, a new promoCode field, "112 occurrences" from a kubectl logs command it never ran, and a Helm revision that didn't exist. It recommended an immediate rollback.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With memory&lt;/strong&gt;, it diagnosed a likely dependency-capacity issue (Redis or database connection-pool exhaustion), pointed to similar past pool-exhaustion incidents, and stated that the memory base flags rollback as a trap for this failure class, naming kubectl rollout undo as the command to avoid.&lt;/p&gt;

&lt;p&gt;The memory-backed answer wasn't clean. It also suggested checking BGP and systemd-networkd changes, which read like bleed-through from unrelated incidents. That's a reminder that retrieval over a mixed dataset brings noise along with signal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The held-out result&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I held out 10 incidents, wrote symptom-only queries for them, and compared with-memory to no-memory using the same model and prompt minus the memory block. With memory: 9 of 10 root causes matched the true category (one run hit a rate limit and counts as a miss). Without memory: 0 of 10 fully correct, 4 partial, 6 hallucinated. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgzb7am84sdapqsbkwwaz.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgzb7am84sdapqsbkwwaz.jpeg" alt=" " width="800" height="382"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Limitations of this dataset choice&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;• &lt;strong&gt;It's one dataset from a handful of vendors&lt;/strong&gt;. Outages cluster into recurring failure classes, so a held-out incident can resemble several retained ones. This doesn't test truly novel failure modes.&lt;br&gt;
• &lt;strong&gt;n = 10, self-graded&lt;/strong&gt;. I graded against true_category, with no second grader.&lt;br&gt;
• &lt;strong&gt;Category-level grading&lt;/strong&gt;. "Correct" means the right kind of cause, not a reproduction of the postmortem.&lt;br&gt;
• &lt;strong&gt;Postmortems are written after the fact&lt;/strong&gt;. They're curated narratives, and a live incident is messier than the symptom text I wrote for the tests.&lt;br&gt;
• &lt;strong&gt;No comparison against other data sources.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaways&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Choose data that records failed fixes&lt;/strong&gt; . The most valuable content in a postmortem is often what the responders tried that didn't work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hold some of it out before you seed&lt;/strong&gt;. Decide the split before you retain anything, or you'll be tempted to test on what the agent has already seen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep a label you can grade against&lt;/strong&gt;. true_category turned my evaluation from vibes into counts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expect retrieval noise&lt;/strong&gt;. Mixed data brings unrelated suggestions along, and the agent's answer should be read as a hypothesis.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Say what you didn't compare&lt;/strong&gt;. I have no measured comparison against other data, so I don't claim one.&lt;/li&gt;
&lt;/ol&gt;

</description>
    </item>
  </channel>
</rss>
