<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sravya Marikokkula</title>
    <description>The latest articles on DEV Community by Sravya Marikokkula (@sravya_marikokkula_dea1b7).</description>
    <link>https://dev.to/sravya_marikokkula_dea1b7</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150028%2Fc0191cbf-7986-4791-8772-63f32248d032.png</url>
      <title>DEV Community: Sravya Marikokkula</title>
      <link>https://dev.to/sravya_marikokkula_dea1b7</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sravya_marikokkula_dea1b7"/>
    <language>en</language>
    <item>
      <title>My First 10/10 Was a Lie: How I Tested an SRE Agent Properly</title>
      <dc:creator>Sravya Marikokkula</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:02:40 +0000</pubDate>
      <link>https://dev.to/sravya_marikokkula_dea1b7/my-first-1010-was-a-lie-how-i-tested-an-sre-agent-properly-466e</link>
      <guid>https://dev.to/sravya_marikokkula_dea1b7/my-first-1010-was-a-lie-how-i-tested-an-sre-agent-properly-466e</guid>
      <description>&lt;p&gt;My first evaluation of an incident-response agent scored a perfect 10 out of 10. It was meaningless, and I want to explain why, because the mistake is easy to make and the fix is cheap.&lt;br&gt;
&lt;strong&gt;The flawed first test&lt;/strong&gt;&lt;br&gt;
The agent recalls real postmortems from a Hindsight memory bank. I picked 10 incidents, wrote queries from them, and checked whether the agent found the right root cause, warned about the trap action, and cited a real incident. It did all three, ten times.&lt;br&gt;
The problem: the 10 test incidents were also stored in memory. Retrieving a postmortem for a query derived from that postmortem is lookup. It tells you the retrieval works. It says nothing about whether the agent helps with an outage it hasn't seen.&lt;br&gt;
&lt;strong&gt;The redo&lt;/strong&gt;&lt;br&gt;
I rebuilt the test so the agent had to deal with unseen incidents:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Hold out&lt;/strong&gt;. Of the 114 OpenSRE incidents, I retained 104 in a fresh memory bank and held 10 out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Symptom-only queries&lt;/strong&gt;. For each held-out incident, I wrote a query describing the symptoms, without root-cause wording or language from the postmortem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two conditions&lt;/strong&gt;. Each query ran once with memory and once as a baseline using the same model and the same prompt minus the memory block.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grade against a label&lt;/strong&gt;. I graded each answer against the dataset's true_category field.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Save everything&lt;/strong&gt;. All outputs went into eval_holdout_results.json, so every number traces to a file.
The "same prompt minus the memory block" detail matters. The system prompt is assembled like this:
python
&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;memory_context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HINDSIGHT REFLECTION:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;reflection&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;scored_memories&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;memory_context&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;RANKED MEMORIES (by relevance):&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;scored_memories&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;memory_context&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;- [score: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;score&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;] &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The baseline receives the same instructions with memory_context removed. That isolates memory as the variable, at least at the level of "memory block present or absent."&lt;br&gt;
&lt;strong&gt;The result&lt;/strong&gt;&lt;br&gt;
Condition           Root cause correct&lt;br&gt;
With memory         9 / 10&lt;br&gt;
No memory           0 / 10 fully correct (4 partial, 6 hallucinated)&lt;br&gt;
A partial match means the baseline named a plausible cause in the right area (a version skew, a cache issue, a resource limit) but described the wrong mechanism or missed the true trigger.&lt;br&gt;
The tenth memory-backed query hit Groq's daily rate limit and never completed. I counted it as a miss instead of reporting 9/9, since I have no result for it either way.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flizj3fnifnz1miy91urd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flizj3fnifnz1miy91urd.png" alt=" " width="752" height="359"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The memory-backed diagnosis: HIGH CONFIDENCE, real incident citations &lt;br&gt;
from the held-out set, and trap-action warnings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why the baseline comparison matters&lt;/strong&gt;&lt;br&gt;
A score of 9/10 sounds impressive until you ask "compared to what?" Without a baseline, you can't tell whether the agent is good or the questions are easy. The no-memory run answers that.&lt;br&gt;
The baseline didn't just score lower; it failed in a specific way. On my demo query (a checkout service returning 500s on about 12% of requests after a 06:31 deploy), the no-memory model fabricated a NullPointerException, log counts from a kubectl command it never ran, and a Helm revision that doesn't exist. Across the held-out set it kept reaching for code-level guesses: connection leaks, TTL misconfigurations, missing null checks.&lt;br&gt;
With memory, the answer for the same query pointed to a likely dependency-capacity problem and said not to roll back, because in similar past incidents it made things worse. I'd note two caveats. The answer hedged with "likely" and a Medium confidence label, and it included some noise (a suggestion to check BGP and systemd-networkd changes). The confidence label comes from a regex over the model's own text:&lt;br&gt;
python&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;extract_confidence&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;match&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[Cc]onfidence[:\s\-\*]+(\w+)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;match&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;level&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;match&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;level&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;high&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;medium&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;low&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;level&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;capitalize&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Unknown&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It returns "Unknown" on a failed parse because my first version defaulted to "Medium," which quietly turned a parsing bug into a confidence claim. Either way, it's the model's self-report, not a calibrated probability.&lt;br&gt;
&lt;strong&gt;What this evaluation doesn't show&lt;/strong&gt;&lt;br&gt;
I'd rather list these than have someone find them:&lt;br&gt;
• &lt;strong&gt;n = 10.&lt;/strong&gt; One miss moves the score ten points. I wouldn't compute a confidence interval.&lt;br&gt;
• &lt;strong&gt;One grader, and it's me&lt;/strong&gt;. I graded my own outputs against the label. No second opinion.&lt;br&gt;
• &lt;strong&gt;Category-level grading.&lt;/strong&gt; "Correct" means the right kind of cause (dependency saturation vs. a config push vs. a network fault), not the postmortem's exact sequence of events.&lt;br&gt;
• &lt;strong&gt;"Partial" is a judgment call&lt;/strong&gt;. Four baseline runs counted as partial matches. Someone else might draw that line differently.&lt;br&gt;
• &lt;strong&gt;Held-out isn't unrelated.&lt;/strong&gt; All incidents share a dataset and a handful of vendors, and outages cluster into recurring classes. A held-out incident can resemble several retained ones.&lt;br&gt;
• &lt;strong&gt;One run per query.&lt;/strong&gt; I didn't repeat runs, so I can't speak to stability.&lt;br&gt;
• &lt;strong&gt;A rate-limited run&lt;/strong&gt;. One of ten never finished.&lt;br&gt;
• &lt;strong&gt;No ablation.&lt;/strong&gt; I can't attribute the gain to reflect, recall, the trap boost, or the signature enrichment.&lt;br&gt;
A stronger test would use more incidents, repeated runs, an independent grader, and a held-out set chosen to look different from the retained one. The memory layer is open source (GitHub, docs), and Vectorize's overview of agent memory explains the ideas behind it.&lt;br&gt;
&lt;strong&gt;Takeaways&lt;/strong&gt;&lt;br&gt;
1.&lt;strong&gt;If the test data is in memory, you're testing lookup&lt;/strong&gt;. Hold out before you seed.&lt;br&gt;
2.&lt;strong&gt;Always run a baseline.&lt;/strong&gt; A score without a comparison is hard to interpret.&lt;br&gt;
3.&lt;strong&gt;Write symptom-only queries&lt;/strong&gt;. Don't use root-cause words from the answer.&lt;br&gt;
4.&lt;strong&gt;Count failures as failures&lt;/strong&gt;. A rate-limited run is a miss, not a missing data point.&lt;br&gt;
5.&lt;strong&gt;Save raw outputs&lt;/strong&gt;. It keeps you honest about what you claim.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>devops</category>
      <category>sre</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
