DEV Community

Sravya Marikokkula
Sravya Marikokkula

Posted on

My First 10/10 Was a Lie: How I Tested an SRE Agent Properly

My first evaluation of an incident-response agent scored a perfect 10 out of 10. It was meaningless, and I want to explain why, because the mistake is easy to make and the fix is cheap.
The flawed first test
The agent recalls real postmortems from a Hindsight memory bank. I picked 10 incidents, wrote queries from them, and checked whether the agent found the right root cause, warned about the trap action, and cited a real incident. It did all three, ten times.
The problem: the 10 test incidents were also stored in memory. Retrieving a postmortem for a query derived from that postmortem is lookup. It tells you the retrieval works. It says nothing about whether the agent helps with an outage it hasn't seen.
The redo
I rebuilt the test so the agent had to deal with unseen incidents:

  1. Hold out. Of the 114 OpenSRE incidents, I retained 104 in a fresh memory bank and held 10 out.
  2. Symptom-only queries. For each held-out incident, I wrote a query describing the symptoms, without root-cause wording or language from the postmortem.
  3. Two conditions. Each query ran once with memory and once as a baseline using the same model and the same prompt minus the memory block.
  4. Grade against a label. I graded each answer against the dataset's true_category field.
  5. Save everything. All outputs went into eval_holdout_results.json, so every number traces to a file. The "same prompt minus the memory block" detail matters. The system prompt is assembled like this: python
memory_context = f"HINDSIGHT REFLECTION:\n{reflection}\n\n"
if scored_memories:
    memory_context += "RANKED MEMORIES (by relevance):\n"
    for item in scored_memories:
        memory_context += f"- [score: {item['score']}] {item['text']}\n"
Enter fullscreen mode Exit fullscreen mode

The baseline receives the same instructions with memory_context removed. That isolates memory as the variable, at least at the level of "memory block present or absent."
The result
Condition Root cause correct
With memory 9 / 10
No memory 0 / 10 fully correct (4 partial, 6 hallucinated)
A partial match means the baseline named a plausible cause in the right area (a version skew, a cache issue, a resource limit) but described the wrong mechanism or missed the true trigger.
The tenth memory-backed query hit Groq's daily rate limit and never completed. I counted it as a miss instead of reporting 9/9, since I have no result for it either way.

The memory-backed diagnosis: HIGH CONFIDENCE, real incident citations
from the held-out set, and trap-action warnings.

Why the baseline comparison matters
A score of 9/10 sounds impressive until you ask "compared to what?" Without a baseline, you can't tell whether the agent is good or the questions are easy. The no-memory run answers that.
The baseline didn't just score lower; it failed in a specific way. On my demo query (a checkout service returning 500s on about 12% of requests after a 06:31 deploy), the no-memory model fabricated a NullPointerException, log counts from a kubectl command it never ran, and a Helm revision that doesn't exist. Across the held-out set it kept reaching for code-level guesses: connection leaks, TTL misconfigurations, missing null checks.
With memory, the answer for the same query pointed to a likely dependency-capacity problem and said not to roll back, because in similar past incidents it made things worse. I'd note two caveats. The answer hedged with "likely" and a Medium confidence label, and it included some noise (a suggestion to check BGP and systemd-networkd changes). The confidence label comes from a regex over the model's own text:
python

def extract_confidence(text: str) -> str:
    match = re.search(r"[Cc]onfidence[:\s\-\*]+(\w+)", text)
    if match:
        level = match.group(1).lower()
        if level in ["high", "medium", "low"]:
            return level.capitalize()
    return "Unknown"
Enter fullscreen mode Exit fullscreen mode

It returns "Unknown" on a failed parse because my first version defaulted to "Medium," which quietly turned a parsing bug into a confidence claim. Either way, it's the model's self-report, not a calibrated probability.
What this evaluation doesn't show
I'd rather list these than have someone find them:
• n = 10. One miss moves the score ten points. I wouldn't compute a confidence interval.
• One grader, and it's me. I graded my own outputs against the label. No second opinion.
• Category-level grading. "Correct" means the right kind of cause (dependency saturation vs. a config push vs. a network fault), not the postmortem's exact sequence of events.
• "Partial" is a judgment call. Four baseline runs counted as partial matches. Someone else might draw that line differently.
• Held-out isn't unrelated. All incidents share a dataset and a handful of vendors, and outages cluster into recurring classes. A held-out incident can resemble several retained ones.
• One run per query. I didn't repeat runs, so I can't speak to stability.
• A rate-limited run. One of ten never finished.
• No ablation. I can't attribute the gain to reflect, recall, the trap boost, or the signature enrichment.
A stronger test would use more incidents, repeated runs, an independent grader, and a held-out set chosen to look different from the retained one. The memory layer is open source (GitHub, docs), and Vectorize's overview of agent memory explains the ideas behind it.
Takeaways
1.If the test data is in memory, you're testing lookup. Hold out before you seed.
2.Always run a baseline. A score without a comparison is hard to interpret.
3.Write symptom-only queries. Don't use root-cause words from the answer.
4.Count failures as failures. A rate-limited run is a miss, not a missing data point.
5.Save raw outputs. It keeps you honest about what you claim.

Top comments (1)

Collapse
 
promptalo profile image
PromptAlo •

Holding out 10 of 114 changes what a pass means: the first 10/10 passed because the test incidents were also in memory, so lookup alone could satisfy it. With the label as the pass condition, the four partial baseline answers are the soft spot. Is there a written rule for when a partial counts, or is it the grader's call each time?