<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alessandro Prandini</title>
    <description>The latest articles on DEV Community by Alessandro Prandini (@alexpran).</description>
    <link>https://dev.to/alexpran</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4117729%2F8f66778e-429e-4416-a525-eef0817b29f8.jpg</url>
      <title>DEV Community: Alessandro Prandini</title>
      <link>https://dev.to/alexpran</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/alexpran"/>
    <language>en</language>
    <item>
      <title>Bad evals, my own: five exercises from two LLM judges</title>
      <dc:creator>Alessandro Prandini</dc:creator>
      <pubDate>Mon, 14 Sep 2026 15:18:20 +0000</pubDate>
      <link>https://dev.to/alexpran/bad-evals-my-own-five-exercises-from-two-llm-judges-1f69</link>
      <guid>https://dev.to/alexpran/bad-evals-my-own-five-exercises-from-two-llm-judges-1f69</guid>
      <description>&lt;p&gt;I run two small LLM judges. One reads ~200 items a day from AI feeds and tells me which five to read (I call it brief). The other reads Reddit threads and tells me which ones are worth a comment from me (scout). Both have a regression suite, both have a promoted baseline, both go through a gate before I change a prompt. I built the gate, so I had every reason to believe the numbers.&lt;/p&gt;

&lt;p&gt;Then I reread &lt;a href="https://danluu.com/exercise-7/" rel="noopener noreferrer"&gt;Dan Luu's exercise 7&lt;/a&gt;, which we cite on the &lt;a href="https://digline.dev/why/" rel="noopener noreferrer"&gt;why page&lt;/a&gt;. His method is simple: show the benchmark as published, ask "what's wrong with this?", and only then explain. His point is that you don't need domain expertise to find these problems, just the reasoning you'd apply to any experiment. So I applied it to my own judges. Below are five exercises. Every number comes from files in the two repos, with the run id, and at the end there's a section on what this exercise found in the tool itself, which was not the plan.&lt;/p&gt;

&lt;p&gt;As in the original, the artifacts come first and the explanations later, in case you want to think before reading mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  The exercises
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Three runs
&lt;/h3&gt;

&lt;p&gt;On September 3 I ran the brief suite three times in ten minutes. 21 cases, Haiku 4.5, 5 samples per case, majority vote. The run files record the config hash and the sha256 of both prompt files; all three are identical (&lt;code&gt;98fc65b1e49e930e&lt;/code&gt;, &lt;code&gt;05c20df6…&lt;/code&gt;, &lt;code&gt;3ce6ed3b…&lt;/code&gt;), and git shows no commit touching &lt;code&gt;prompts/&lt;/code&gt; between the first and the third.&lt;/p&gt;

&lt;p&gt;Aggregates: &lt;strong&gt;16/21 → 14/21 → 16/21&lt;/strong&gt; accuracy, 10/15 → 9/15 → 10/15 precision.&lt;/p&gt;

&lt;p&gt;Here are the seven cases where at least one run was not unanimous. Each string is the five samples of the "agrees with my mark" assertion, 1 = agreed. &lt;code&gt;(p)&lt;/code&gt;/&lt;code&gt;(f)&lt;/code&gt; is the majority verdict.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;case&lt;/th&gt;
&lt;th&gt;run 06:14&lt;/th&gt;
&lt;th&gt;run 06:18&lt;/th&gt;
&lt;th&gt;run 06:24&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;controlling reasoning effort&lt;/td&gt;
&lt;td&gt;10110 (p)&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;don't classify, hallucinate&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;td&gt;11101 (p)&lt;/td&gt;
&lt;td&gt;11110 (p)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;more than just code review&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;01010 (f)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;recent developments in LLM architectures&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;td&gt;01110 (p)&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;evals skills for coding agents&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;00011 (f)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;how we built auto mode&lt;/td&gt;
&lt;td&gt;11110 (p)&lt;/td&gt;
&lt;td&gt;11011 (p)&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"it's hard to eval" is a product smell&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;td&gt;11111 (p)&lt;/td&gt;
&lt;td&gt;11110 (p)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Which of the three runs is the regression?&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  2. A reasonable clause
&lt;/h3&gt;

&lt;p&gt;Scout's judge takes a Reddit thread and returns &lt;code&gt;comment&lt;/code&gt;, &lt;code&gt;upvote&lt;/code&gt; or &lt;code&gt;skip&lt;/code&gt;. The rules live in a file called JUDGE.md. Accuracy below is over the three classes; precision and recall are on &lt;code&gt;comment&lt;/code&gt; only, since that is the verdict that costs me something. On September 9 I tried two edits, one after the other, against a baseline of 17 labelled threads. Both edits looked reasonable to me when I wrote them.&lt;/p&gt;

&lt;p&gt;Edit A adds a paragraph:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt; - an open-source maintainer of an adjacent eval/testing project asking for methodological feedback
&lt;span class="err"&gt;
&lt;/span&gt;&lt;span class="gi"&gt;+Not "comment", however much it looks like one, when the practice my experience
+would recommend is already written in the thread: the author lists it among the
+options they are weighing, or reports it as what they already do. Then the
+thread has no question left for me and the rules below decide it as if it had
+none. It stays "comment" only if I can name the failure that the author's own
+list does not cover.
+
&lt;/span&gt; verdict = "upvote" when the thread is on topic and honest but has no question for me
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Edit B adds one line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt; - an open-source maintainer of an adjacent eval/testing project asking for methodological feedback
&lt;span class="gi"&gt;+- A thread is a comment if any sentence in it asks how to detect that generated logic, prompts or outputs changed from what was previously approved, or how to know whether a fix worked — regardless of the thread's topic or the author's expertise.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Results, same judge (Haiku 4.5, the model scout was running on that morning), same 17 cases, 5 samples each:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;accuracy&lt;/th&gt;
&lt;th&gt;precision&lt;/th&gt;
&lt;th&gt;recall&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;unmodified JUDGE.md (reference)&lt;/td&gt;
&lt;td&gt;10/17&lt;/td&gt;
&lt;td&gt;9/12&lt;/td&gt;
&lt;td&gt;9/12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;edit A&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;8/17&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;6/9&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;6/12&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;edit B&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;9/17&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;8/11&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;8/12&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both edits were reverted the same morning. Later that day I moved the judge to Sonnet 5 and re-set the gate's thresholds on its run, which is why the baseline you'll see in exercises 4 and 5 has different numbers on the same 17 cases, and why the Haiku reference above would not pass that gate: the edits were tested against the reference of the moment, not against the baseline promoted afterwards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can you tell from the diff why? What would you need in order to tell?&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3. A clean precision
&lt;/h3&gt;

&lt;p&gt;Brief scores each item 1 to 5. It shows me everything scored 4 or 5, or the top 5 if fewer, and asks which ones interest me. That answer is the ground truth for the suite. The suite's 21 cases are the items I had been shown by August 25: 18 that scored 4 or 5 and 3 below-threshold fillers from the pad-to-five rule. The 10 I marked are its positives. The suite reports precision 10/15, and the gate has been green for two weeks.&lt;/p&gt;

&lt;p&gt;Here is the state of the file that holds the ground truth, 446 items, on September 6:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://digline.dev/assets/bad-evals/ex3_censored.svg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdigline.dev%2Fassets%2Fbad-evals%2Fex3_censored.svg" alt="Bar chart of the 446 items in brief's seen.json by judge score. Score 1: 111 items, 8 shown to me, none marked interesting. Score 2: 37 items, 6 shown, none marked. Score 3: 12 items, 2 shown, none marked. Score 4: 14 items, all 14 shown, 7 marked interesting. Score 5: 7 items, all 7 shown, 5 marked. Not judged: 265 items, none shown." width="576" height="273"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;judge score&lt;/th&gt;
&lt;th&gt;items&lt;/th&gt;
&lt;th&gt;shown to me&lt;/th&gt;
&lt;th&gt;marked interesting&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;111&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;not judged&lt;/td&gt;
&lt;td&gt;265&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Which metric cannot be computed from this table, no matter how many runs I do?&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Same judge, same run, two numbers
&lt;/h3&gt;

&lt;p&gt;For two weeks scout's suite had 17 cases. On September 11 I regenerated it from everything I had actually done on Reddit since the judge went live: 144 threads, each labelled with what I did (commented, upvoted, ignored). Same judge (Sonnet 5), same prompt, same config hash, one run over all 144 cases. Then I computed the metrics twice: on the 17 old cases, which are a subset of the 144, and on the full set.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://digline.dev/assets/bad-evals/ex4_17_vs_144.svg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdigline.dev%2Fassets%2Fbad-evals%2Fex4_17_vs_144.svg" alt="Two charts from the same run and the same judge. Top: on the 17-case suite accuracy is 11/17, precision 11/15 and recall 11/12; on the 144-case suite accuracy rises to 123/144 while precision falls to 14/27 and recall to 14/20. Bottom: what I did with the threads in each suite. The 17-case suite is 12 commented, 4 upvoted and 1 ignored; the 144-case suite is 20 commented, 5 upvoted and 119 ignored." width="460" height="532"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;17 cases&lt;/th&gt;
&lt;th&gt;144 cases&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;accuracy&lt;/td&gt;
&lt;td&gt;11/17 = 0.65&lt;/td&gt;
&lt;td&gt;123/144 = 0.85&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;precision&lt;/td&gt;
&lt;td&gt;11/15 = 0.73&lt;/td&gt;
&lt;td&gt;14/27 = 0.52&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;recall&lt;/td&gt;
&lt;td&gt;11/12 = 0.92&lt;/td&gt;
&lt;td&gt;14/20 = 0.70&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 17-case numbers are identical to the promoted baseline, so this is not drift between runs. The gate passes on 17 cases and fails on 144.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which of the two precisions is the true one?&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  5. My gate has the discontinuity Dan Luu complains about
&lt;/h3&gt;

&lt;p&gt;Scout's gate is three thresholds: accuracy ≥ 0.60, precision ≥ 0.70, recall ≥ 0.85. I set them from the promoted run (11/17, 11/15, 11/12) with a written rule: the threshold sits between the measured value and the value one case lower, so a single case flipping the wrong way fails the gate.&lt;/p&gt;

&lt;p&gt;Exercise 7 spends a section on exactly this pattern in Senior SWE-Bench: a continuous score, then a hard cutoff, so that one line of code more turns a "tasteful" solve into a failure. The criticism is that the cutoff is an arbitrary formula that nobody had to write down.&lt;/p&gt;

&lt;p&gt;Here is what one case is worth, in accuracy points, as a function of suite size:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://digline.dev/assets/bad-evals/ex5_one_case.svg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdigline.dev%2Fassets%2Fbad-evals%2Fex5_one_case.svg" alt="Curve of how many accuracy points a single case is worth as the suite grows from 5 to 300 cases, falling steeply and then flattening. Marked on it: my old suite at 17 cases, where one case is 5.9 points; DeepSWE, 0.9 points; my new suite at 144 cases, 0.7 points." width="576" height="259"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is my gate an instance of the same mistake? If not, what is the difference, and when does it stop holding?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The explanations
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. None of them
&lt;/h3&gt;

&lt;p&gt;Nothing regressed. The second run flipped two cases to fail because 3 out of 5 samples said so, and the third run flipped them back. The prompt, model and config were byte-identical across the three, so the only thing that moved was the judge itself.&lt;/p&gt;

&lt;p&gt;This is what I built the 5-sample floor for, and the floor did its job: the gate said "you cannot promote or reject on one run". What the three runs alone don't tell you is &lt;em&gt;how much&lt;/em&gt; the judge moves. So I pulled every run of the brief suite that has the same config hash. There are sixteen of them, from August 27 to September 10.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://digline.dev/assets/bad-evals/ex1_nonunanimity.svg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdigline.dev%2Fassets%2Fbad-evals%2Fex1_nonunanimity.svg" alt="Bar chart of 16 runs of the brief suite from August 27 to September 10, all with the same judge, prompt and config hash, showing how many of the 21 cases had five samples that did not all agree. The count ranges from 2/21 to 6/21. The three fixture runs of September 3 are highlighted at 2/21, 5/21 and 2/21." width="648" height="259"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The fraction of cases where the five samples disagree goes from 2/21 to 6/21, on runs that are supposed to be replicates. If the noise floor is what you measure once and then trust, this is a problem: the floor itself has a floor.&lt;/p&gt;

&lt;p&gt;Dan Luu re-graded Senior SWE-Bench outputs ten times and found the official result flipped 23% of the time. My number is not the same quantity, so I won't put them side by side without saying what mine is. For each non-unanimous case the pattern is 4/1 or 3/2, which means a single sample would have disagreed with the majority with probability 0.2 or 0.4. Weighted over the 21 cases, the chance that a one-sample run gives a different verdict on a given case is between 2% and 8% depending on which of the three runs you pick. The 23% and my 2-8% answer different questions (how often the published single run is wrong, vs. how often one sample would disagree with five), but they come from the same fact: an LLM grading a fixed input is not a deterministic function of the input.&lt;/p&gt;

&lt;p&gt;One caveat that matters for the honesty of this section. The run files store, per sample, whether the assertion passed, not the raw score the judge produced. So I know that sample 2 disagreed with my mark; I don't know whether it said 3 instead of 4 or 1 instead of 5. Keeping the raw output per sample became possible only this week, and only if you turn it on. I hadn't.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. You can't, and neither could I
&lt;/h3&gt;

&lt;p&gt;I looked for the bug in the two edits for a good half hour, diff in hand, and there is nothing to find there. Edit A tells the judge "if the author already lists the practice you'd recommend, don't comment". Edit B tells it "any thread asking how to detect a change from an approved state is a comment". Both are things I believe.&lt;/p&gt;

&lt;p&gt;My reading, from which cases fell rather than from any recorded reasoning (I wasn't keeping the raw output yet), is that Haiku read a sufficient condition as a necessary one. After edit B, threads that clearly matched the earlier rules but didn't contain a sentence about "detecting change" started coming back as &lt;code&gt;skip&lt;/code&gt;, because the new line read like the definition of a comment rather than one more way to earn one. Recall went from 9/12 to 8/12 on that edit and 6/12 on the other, and the cases that dropped were ones that had been unanimous &lt;code&gt;comment&lt;/code&gt; for days. The prompt edits were fine as English. They were bad as instructions to this particular model, and no amount of staring at the diff tells you that. What you need is not a sharper eye but a run against a baseline you trust.&lt;/p&gt;

&lt;p&gt;Two footnotes. First, the difference between 9/12 and 8/12 is one case, and given exercise 1 you should ask whether that's noise. It's a fair question; the reason I reverted anyway is that the cases that flipped were the stable ones, not the flapping ones, and the 6/12 of edit A is outside anything the noise floor produces. Second, the two diffs above do not exist in git. I reverted with a working-tree checkout, so the repository has exactly one blob of JUDGE.md, ever. The diffs come from the run files, which store the full text of every declared artifact at the moment of the run. September 11 was the first time that design decision paid for itself, and it paid for the whole exercise.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Recall
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;marked&lt;/code&gt; is written only for items the script showed me (lines 359 to 377 of &lt;code&gt;brief.py&lt;/code&gt;: build the list of items scoring 4 or above, pad to five if shorter, ask, write the answer on those items and only those). An item the judge scored 2 and didn't show me has no label, and never will. So the suite can measure how many of the judge's picks I liked. It cannot measure how many things I would have liked the judge didn't pick, because I never saw them.&lt;/p&gt;

&lt;p&gt;There is a second layer I hadn't noticed until I made the table: 265 of the 446 items were never scored at all. The script caps how many items it judges per run, and everything past the cap is stored as skipped. So it isn't "the judge hides the low scores from me"; it's "the judge hides the low scores from me, and something else hides 60% of the feed from the judge".&lt;/p&gt;

&lt;p&gt;The 16 low-scoring items I did see (through the pad-to-five rule) I marked as not interesting, all 16. That's mildly reassuring about the judge and says nothing about recall, since 16 of 160 low scores is not a sample of anything.&lt;/p&gt;

&lt;p&gt;One more thing hides in the gap between the two tables. 18 of the 21 suite cases scored 4 or 5 the first time; when the suite re-judges them, 15 come back at 4 or above. The three that drop are all items I hadn't marked, and it is the same three in 14 of the 16 runs. That is not exercise 1's noise showing up again at random. For two of the three my reading is that it is exercise 1's noise turned into a bias by selection: the morning pass judges each item once and shows me whatever clears 4, so an item that clears on one lucky sample gets into the suite, and the suite, sampling five times, puts it back. Across the 16 runs those two items reach 4 in 9 and 7 samples out of 80, so a lucky single sample is plausible. The third item reaches 4 in 0 samples out of 80, which makes a lucky sample very unlikely; the more probable story is that the prompt on August 24 was not the prompt the suite runs today, and I can't check, because the morning pass didn't record the prompt's hash and git starts two days later. Either way, the 10/15 precision counts three items against the judge for a decision the suite's own judge would not make.&lt;/p&gt;

&lt;p&gt;None of this makes the precision number wrong. It makes it a number about half the pipeline, reported as if it were about the pipeline.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Neither
&lt;/h3&gt;

&lt;p&gt;The 17-case suite had one negative example. One &lt;code&gt;ignored&lt;/code&gt; thread against 12 &lt;code&gt;commented&lt;/code&gt; and 4 &lt;code&gt;upvoted&lt;/code&gt;. Recall on that suite was recall on things I had already decided to engage with, and precision was computed against 15 predicted positives of which almost all were positives by construction. It's the situation Dan Luu describes for DeepSWE from the other side: I was scoring the judge on the tasks it was already good at, because those were the only tasks I had bothered to label.&lt;/p&gt;

&lt;p&gt;The 144-case suite fixes that, then breaks in a different place. The labels come from what I did, and "ignored" turns out to mean two things.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://digline.dev/assets/bad-evals/ex4_confusion.svg" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdigline.dev%2Fassets%2Fbad-evals%2Fex4_confusion.svg" alt="144 cases, my label vs the judge's majority verdict; one case had no majority verdict and is counted as a fail, not shown. Threads I commented on: 14 judged comment, 2 upvote, 4 skip. Threads I upvoted: 4 comment, 0 upvote, 1 skip. Threads I ignored: 9 comment, 0 upvote, 109 skip." width="388" height="273"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nine threads I ignored came back &lt;code&gt;comment&lt;/code&gt;, five of them unanimously across the 5 samples. With the raw outputs recorded this time, I could read the judge's reasoning on each. On at least five, the judge is right by the letter of JUDGE.md: a thread asking how to tell which of forty diffs between two runs actually matters, one asking what would make you distrust a project-level quality score, one about test-set leakage in prompt optimization. Those are questions for me. I didn't ignore them because the rules say skip. I ignored them because it was the fourth thread that morning, or because I had already commented twice on that subreddit that day. So the ground truth for "not a comment" is polluted with "not today", and the measured precision is lower than the true one by an amount I can't compute from this data.&lt;/p&gt;

&lt;p&gt;This also reaches back into exercise 2. Both edits were rejected on the 17-case suite, the one with a single negative, so was that rejection worth anything? I think so, for one reason: the verdict rested on recall (9/12 down to 6/12 and 8/12), and recall over the positives is the one quantity a suite made almost entirely of positives measures properly. Had the rejection rested on precision, it would have been a coin toss on 9 to 12 predictions.&lt;/p&gt;

&lt;p&gt;The honest answer is that the suite measures agreement with my behaviour, and my behaviour is not the rule. To measure the rule I have to change the question the script asks me at the end of each morning: not "did you comment?" but "should this have been a comment?", which is a different, slower and more annoying question. That change is next.&lt;/p&gt;

&lt;p&gt;One more thing the recorded outputs showed and no aggregate would have: across the nine false positives the judge's proposed angle is nearly always one of two anecdotes from my own experience, recycled. Precision doesn't care. The person who then writes the comment does.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. It's a different object, until it isn't
&lt;/h3&gt;

&lt;p&gt;The thresholds in Senior SWE-Bench turn a continuous quality score into a ranking of models. A discontinuity there manufactures differences between systems that the data doesn't support. My thresholds do something else: they answer one yes/no question about one system, "may this prompt change be promoted?", and the answer has to be binary because the action is binary. Somebody has to sign, and a signature is a step function.&lt;/p&gt;

&lt;p&gt;That's the defence, and I think it holds for a gate, provided the step is placed against a measured floor rather than a round number. The rule "one case worse than measured" is exactly that: the threshold is not 0.7 because 0.7 sounds right, it's 0.7 because 11/15 passed and 10/15 would not, and the 5-sample floor says a single-case move on a stable case is signal.&lt;/p&gt;

&lt;p&gt;It stops holding as the suite grows.&lt;/p&gt;

&lt;p&gt;On 17 cases one case is 5.9 accuracy points, which is more than the noise floor, so "one case worse" is a meaningful event. On 144 cases one case is 0.7 points, and the September 11 run had 13 non-unanimous cases. A threshold one case below the measured value is now inside the noise, and the rule I wrote down for 17 cases is wrong for 144. I don't have the replacement yet. The candidate is a threshold expressed in cases rather than in points, with the count of stable cases that moved as the quantity, which is roughly what the floor already knows.&lt;/p&gt;

&lt;p&gt;A smaller confession that belongs here. When I went to write "here is the history of my thresholds", there wasn't one. &lt;code&gt;promote&lt;/code&gt; overwrote a single file with no log, the baseline file recorded the run's timestamp but not the promotion's, and of the three threshold configurations that appear in run files, one was never committed and is unrecoverable. The history of my own gate lived in commit messages I happened to write. That got fixed this week, after this exercise, not before.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the exercise found in the tool
&lt;/h2&gt;

&lt;p&gt;I set out to find flaws in two prompts and found four in the tool that measures them, in order of how much they bothered me, plus one that turned out to be me.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The gate's central act leaves no trace.&lt;/strong&gt; The two rejections of September 9 exist nowhere except a commit message I wrote. &lt;code&gt;compare&lt;/code&gt; printed its verdict to a terminal and forgot it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Promote had no history.&lt;/strong&gt; See exercise 5.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A case with no majority, and what the run says about it.&lt;/strong&gt; One of the 144 cases came back 2/2/1 across the three verdicts. The suite requires 3-of-5 agreement, so I expected a suspended case; the run reports zero. It turned out the agreement rule works on the pass/fail axis, not on the verdict axis: 2 samples agreed with my mark, 3 didn't, that's a 3/5 majority for "fail", and the case fails. Correct by design, and the design wasn't written down anywhere I'd read. It is now.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-call cost budget, checked on the mean.&lt;/strong&gt; Nine individual judgments exceeded the $0.020 budget, one by 17%, and the budget assertion passed on all 144 cases because it evaluates the mean of five samples. Possibly by design. Not what the name promises.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two prompt edits, same config hash.&lt;/strong&gt; Edit A's run has the same config hash as runs with the original JUDGE.md. I filed this as a hole and it isn't one: by design the hash is the identity of the suite (assertions, thresholds, samples), a declared artifact travels with its own sha, and a prompt change shows up in the report as a separate fact, &lt;code&gt;artifacts_changed&lt;/code&gt;, next to the metrics. That's what let the edit be compared against the baseline at all. What I had actually found was that I'd never read that line of the report.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And one incident that turned into a feature. The first attempt at the 144-case run, 720 calls to Sonnet, was killed at around call 400 by a memory watchdog on my laptop. The tool had measured itself flat under 100 MB; the pressure came from an IDE. But the run file was written only at the end, so about $4 of judgments were gone with nothing on disk. The fix, a per-case journal with &lt;code&gt;--resume&lt;/code&gt;, shipped the same evening and the second attempt ran through it. $6.18 for 720 calls, against an estimate of $8.06; adaptive thinking is cheaper on the easy skips.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd take from this
&lt;/h2&gt;

&lt;p&gt;Every one of the five problems is visible in the data I already had. None required knowing anything about LLMs. Exercise 3 is a &lt;code&gt;for&lt;/code&gt; loop that writes a field on the wrong subset; exercise 4 is a class imbalance you'd catch in an intro stats course; exercise 5 is the observation that 1/n gets small. I had the noise floor, the versioned baseline, the recorded artifacts, all the machinery that's supposed to make this rigorous, and I still measured half a pipeline for two weeks and called it precision.&lt;/p&gt;

&lt;p&gt;Dan Luu's line is that evals are more about avoiding mistakes than following a process. I'd add one thing from the other side of the table: the machinery doesn't avoid the mistakes for you, but it does something almost as useful. It keeps the evidence. The diffs in exercise 2 survive only because the run stored the prompt text. The table in exercise 3 exists because every item ever judged is still in the file. When I finally asked the right question, the answer was already on disk.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;All numbers in this post come from &lt;code&gt;.digline/&lt;/code&gt; in the two repos and from &lt;code&gt;seen.json&lt;/code&gt; in each; run ids are &lt;code&gt;2026-09-03T06-14-13…&lt;/code&gt;, &lt;code&gt;06-18-43…&lt;/code&gt;, &lt;code&gt;06-24-50…&lt;/code&gt; for exercise 1, &lt;code&gt;2026-09-09T06-34-01…&lt;/code&gt; and &lt;code&gt;06-51-45…&lt;/code&gt; for exercise 2, &lt;code&gt;2026-09-11T15-09-23…&lt;/code&gt; for exercise 4. Thread titles in exercise 4 are paraphrased; the threads are public but the point is not who wrote them.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>test</category>
      <category>python</category>
      <category>ai</category>
    </item>
    <item>
      <title>My LLM eval cried wolf. Here's what I measured.</title>
      <dc:creator>Alessandro Prandini</dc:creator>
      <pubDate>Wed, 09 Sep 2026 14:41:08 +0000</pubDate>
      <link>https://dev.to/alexpran/my-llm-eval-cried-wolf-heres-what-i-measured-94n</link>
      <guid>https://dev.to/alexpran/my-llm-eval-cried-wolf-heres-what-i-measured-94n</guid>
      <description>&lt;p&gt;Disclosure first: I write digline, a small Python library for regression testing LLM applications. This post is not about the library. It is about a bug in how I was measuring my own pipeline, and about what happened once I started measuring the measurement. Skip the tool if you like; the problem is yours too.&lt;/p&gt;

&lt;h2&gt;
  
  
  The wolf
&lt;/h2&gt;

&lt;p&gt;I have a pipeline that reads for me in the morning. A few dozen RSS feeds from the AI world, a filter, and a Claude call that picks the five items worth my time. I call it the brief. Around it I had an eval suite: 21 cases, each a fixed input with a rubric, scored by an LLM judge against a mark I had given by hand. It ran on every change and it had been green for weeks. Long enough that I had stopped treating it as something that might be wrong and started treating it as the thing that told me when I was.&lt;/p&gt;

&lt;p&gt;One afternoon a case went from 5/5 to 2/5. Nothing else in the report moved. I did what you do. I diffed the prompt against the approved version: identical. The model alias, the config, the retrieval step, the trace of the call itself: identical, byte for byte. I reran the case fifteen minutes later. 5/5.&lt;/p&gt;

&lt;p&gt;I had spent an hour investigating a regression that did not exist. What had regressed was my measurement, and the measurement had no way to tell me that, because I had never asked it what "unchanged" looks like. I had one reference score per case and a threshold, and a threshold assumes the number under it is a number. It was a sample.&lt;/p&gt;

&lt;p&gt;Here is the part that matters more than the lost hour. A gate that cries wolf gets muted. Not out of negligence, out of arithmetic: if it fires once a week on nothing, by the end of the month someone (me) has learned to click through. Then the real regression arrives, the gate fires, and it gets clicked through with all the others. The failure mode of a noisy eval is not the false alarms. It is that it trains you to ignore the true one.&lt;/p&gt;

&lt;p&gt;So I stopped and asked the question I should have asked on day one: how much does this score move when I change nothing?&lt;/p&gt;

&lt;h2&gt;
  
  
  The measurement
&lt;/h2&gt;

&lt;p&gt;What I had was one run, one score per case, stored as the reference, and a comparison rule that flagged any drop larger than a declared tolerance. The tolerance was global. It treated a case whose judge is effectively deterministic and a case whose judge flips a coin on a borderline rubric line with the same seriousness, because it knew nothing about either.&lt;/p&gt;

&lt;p&gt;What I changed, now written down as an ADR:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;When a version is approved, each case is asked K times, not once.&lt;/strong&gt; In my setup K is 5. The score stays the mean of the samples; on a binary check, a mean over five votes is just the majority vote with a threshold at one half.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The reference stores, for each check, the min and max of those K samples&lt;/strong&gt; alongside the score. That range is the check's own noise floor, measured on the version I said was good.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On comparison, a drop counts as a regression only if the new score leaves the reference's band.&lt;/strong&gt; A drop that lands inside it is reported, but as "within the noise", and it is not counted. The report separates the two in the headline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The band is the reference's, never the current run's.&lt;/strong&gt; The reference is the reviewed measurement. Letting a noisy new run widen its own excuse is how a regression hides inside a model that got less stable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A flip through the threshold is never within noise.&lt;/strong&gt; Pass to fail is reported whatever the samples did. The band explains a movement; it does not excuse a result.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Why min/max per check, rather than a tighter tolerance or a standard deviation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Per check, because noise is a property of the check, not the suite.&lt;/strong&gt; One global number is a guess about all of them at once, and it is wrong for most of them in one direction or the other.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Min/max, because with five samples a standard deviation is a model I have not earned.&lt;/strong&gt; The min and the max are what was seen. They are a fact rather than a distribution, and a reader can verify them against the raw samples printed beside them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Because a band is measured, and a tolerance is declared.&lt;/strong&gt; I kept both. A tolerance says how much movement a reviewer decided is acceptable; the band says how much the system moves on its own. The report says which one spoke. Collapsing them would have deleted the ability to say "this moved more than you allowed, and less than it moves by itself".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Because a zero-width band is not a noise floor.&lt;/strong&gt; A check whose reference samples were unanimous has no interval, so every later change of mind is reported. That is right: a case decided five times out of five and now decided twice out of five has moved, and one case is a diagnosis. The floor earns its keep where the dispersion actually was.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two sentences I now keep next to the suite. Provenance in traces is observability; provenance in the reference is a contract. The trace tells you what happened. The reference tells you what you agreed counts as the same. And: the gate is a veto, not an approval. Passing means I found nothing beyond the noise. It does not mean the change is good.&lt;/p&gt;

&lt;h2&gt;
  
  
  The proof
&lt;/h2&gt;

&lt;p&gt;The fixture is public, in the brief's repository, pinned to a commit. Three consecutive runs within eleven minutes, same suite, same prompt files, same &lt;code&gt;config_hash&lt;/code&gt;, five samples per case. Nothing changed between them, and the report checked that for itself before comparing. Two cases, &lt;code&gt;evals-skills-for-coding-agents&lt;/code&gt; and &lt;code&gt;more-than-just-code-review&lt;/code&gt;, go 5/5, then 2/5, then 5/5 within the same triple. That is the wolf, reproduced on demand, twice in one morning.&lt;/p&gt;

&lt;p&gt;Then the middle run compared against the reference. Both aggregates fell. Precision comes in at 0.600000 against a reference band of 0.615385–0.666667: below the floor, reported as beyond the noise. Accuracy fell by more, 0.095238 against precision's 0.066667, and lands at 0.666667 against a band of 0.666667–0.761905: exactly on the lower edge, within the noise, not counted. The reference had already produced that accuracy once in its own five samples. It had never produced that precision.&lt;/p&gt;

&lt;p&gt;On the declared tolerance alone, 0.047619, one case out of twenty-one, both would have tripped. The same run is noisy enough to explain the larger movement and not the smaller one, and only the reference's own history can tell them apart. The two per-case flips are still reported, because a flip is never noise; the difference is that the headline now says one check moved within noise, and I know which aggregate to trust before I open a trace.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8kuz3we51qt18jf1xv23.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8kuz3we51qt18jf1xv23.png" alt="The  raw `digline compare` endraw  report for the middle run. The headline counts three checks worse and one moved within noise; the two per-case flips are listed as passing to failing, then precision is marked beyond the noise of its check and accuracy within it, each with the reference band it is being read against printed beside the scores." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The week after
&lt;/h2&gt;

&lt;p&gt;I have a second pipeline, scout, that reads Reddit feeds and asks a judge which threads are worth a comment. Seventeen cases, each with a human verdict, mine. Same suite machinery, reference approved with bands, same rule. Over the following week I made three changes to it, and the comparison handled all three differently.&lt;/p&gt;

&lt;p&gt;The first was a change to the judge's prompt. The judge kept flagging threads where the practice under discussion was one the author plainly already followed, so I added a clause: it is not a comment if the practice is already in the post. Reasonable, explained, reviewed. Recall went from 0.80 to 0.50, beyond the noise, and the like-for-like count went from 9/15 to 7/15. The per-case reasons showed what had happened. The judge had applied the clause to the author's perceived competence, "this person clearly knows this already", rather than to what the post actually said. Rejected.&lt;/p&gt;

&lt;p&gt;The second was also a prompt change. I added a sufficient condition to the list of what makes a thread worth commenting on. The judge read it as a necessary one. Accuracy went from 0.588 to 0.529 and recall from 0.75 to 0.667. Adding a sufficient condition to the comment list narrowed it. Rejected.&lt;/p&gt;

&lt;p&gt;The conclusion I drew from the two together: with Haiku 4.5 as the judge, rules with more than one clause do not compose. Each clause gets applied, but not in relation to the others, and a reasonable edit to one line changes the meaning of the ones around it in ways I could not see from the prompt.&lt;/p&gt;

&lt;p&gt;The third change was the judge itself. Same prompt, Sonnet 5 with thinking on. Recall went from 0.75 to 0.917 and precision from 0.75 to 0.733. The two comments the previous judge had missed came back, with reasoning that matched why I had labelled them in the first place. One case showed up as broken, and reading the judge's text it was the judge applying a rule I had written more faithfully than I had applied it when labelling; the label was the bug. The comparison promoted this one.&lt;/p&gt;

&lt;p&gt;The cost per judgement went up 7.5×, and that is the thinking tokens, not the price per token: roughly $10 a month instead of roughly $1.35 for this pipeline. I paid it.&lt;/p&gt;

&lt;p&gt;What I want to draw attention to is not the numbers but that the same rule produced three different verdicts, each with reasons per case. Two edits I would have defended in review turned out to break the judge in ways only the cases could show. One expensive change I would have hesitated over turned out to be worth it, and the comparison said so before I had to argue it with myself. That is what a gate with a measured floor buys: not fewer alarms, but alarms I can read.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits, and what readers taught me
&lt;/h2&gt;

&lt;p&gt;The limits, declared:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The band is min/max over K=5. It is not a confidence interval.&lt;/strong&gt; It is optimistic: at small K the observed range is narrower than the true one, so the floor will miss some noise. It will not invent any, which is the direction I care about, but a drop that lands just below the band on a check that was already wobbling deserves a look, not a rollback.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It needs a reference before it is useful.&lt;/strong&gt; Day one is run, look, approve. There is no floor until something has been approved with samples.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM-as-judge is still LLM-as-judge.&lt;/strong&gt; The band measures how much the judge moves, not whether it is right. Scout's third change is the cheerful version of that; the two rejected edits are the other one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After I posted the fixture on Reddit, three readers pointed at holes I had not seen. I am citing them as theirs.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A canary case.&lt;/strong&gt; A fixed input with one unambiguous answer, not scored for quality, only watched inside its band. If the canary moves, the model behind the alias has changed, and the right move is to re-baseline everything rather than chase whichever case happened to trip first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Store the provider fingerprint next to the scores&lt;/strong&gt;, where the provider exposes one. The alias is not the model; the fingerprint is closer to it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A stable score with changing reasoning is a judge that is right by accident&lt;/strong&gt;, and a band on the score cannot see it. Keep the judge's text, not just its number. Scout's three verdicts were readable only because scout keeps its own log of the judge's text; the suite runs did not, and that is the next thing to change.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where this lives
&lt;/h2&gt;

&lt;p&gt;The decision is written up as &lt;a href="https://digline.dev/product/adr/0006-repeated-samples-and-the-noise-floor/" rel="noopener noreferrer"&gt;ADR 0006&lt;/a&gt; on digline.dev, and the three runs, the reference and the comparison are in the &lt;a href="https://github.com/digline/brief/blob/c9d86ff22d68d3df458fa0da81348ec962d16aa7/fixtures/README.md" rel="noopener noreferrer"&gt;fixtures of the brief repository&lt;/a&gt;, pinned to the commit. If you measure your own floor and it looks different from mine, I would like to know.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>testing</category>
      <category>python</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
