<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ward Ed</title>
    <description>The latest articles on DEV Community by Ward Ed (@ward_ed_6b5e6aa8ded94a987).</description>
    <link>https://dev.to/ward_ed_6b5e6aa8ded94a987</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4163315%2F0bd91171-b421-4ec2-9ce1-8301c0077eca.jpg</url>
      <title>DEV Community: Ward Ed</title>
      <link>https://dev.to/ward_ed_6b5e6aa8ded94a987</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ward_ed_6b5e6aa8ded94a987"/>
    <language>en</language>
    <item>
      <title>Benchmark Contamination 101: How Train/Test Overlap Inflates Leaderboard Scores (and How to Catch It)</title>
      <dc:creator>Ward Ed</dc:creator>
      <pubDate>Tue, 06 Oct 2026 19:11:46 +0000</pubDate>
      <link>https://dev.to/ward_ed_6b5e6aa8ded94a987/benchmark-contamination-101-how-traintest-overlap-inflates-leaderboard-scores-and-how-to-catch-2el</link>
      <guid>https://dev.to/ward_ed_6b5e6aa8ded94a987/benchmark-contamination-101-how-traintest-overlap-inflates-leaderboard-scores-and-how-to-catch-2el</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;A leaderboard number is only as trustworthy as the gap between what a model trained on and what it was tested on. When test examples (or near-duplicates of them) leak into pretraining or fine-tuning data, the model memorizes answers instead of generalizing, and the reported score climbs for the wrong reason. This is benchmark contamination, also called train/test overlap or data leakage. It is common, often accidental, and frequently invisible in a self-reported number. This post explains why contamination inflates scores, three detection methods you can run yourself (n-gram overlap, canary strings, membership inference), a runnable Python n-gram checker, and a short checklist you can apply before you trust any benchmark claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is benchmark contamination and why does it inflate scores?
&lt;/h2&gt;

&lt;p&gt;Modern models are trained on web-scale corpora scraped from the open internet. Public benchmarks live on that same internet: on GitHub, in papers, on leaderboard repos, in blog posts that quote questions verbatim, and in Q&amp;amp;A sites where people discuss the exact items. When a benchmark's test questions and answers end up in the training mix, evaluation stops measuring reasoning and starts measuring recall.&lt;/p&gt;

&lt;p&gt;The mechanism is simple. A benchmark is supposed to be a held-out sample. If the test examples were in the training data, the model can reproduce the answer from memory. Memorization looks exactly like competence on a single scored run, but it does not transfer to genuinely new inputs. The score goes up; the underlying capability does not.&lt;/p&gt;

&lt;p&gt;Contamination comes in degrees. Verbatim contamination means the exact test string appears in training. Near-duplicate contamination means a paraphrase, a translated copy, or a reformatted version appears. Label leakage means the answer key or solution is present even if the question wording differs. Even partial exposure helps a model disproportionately on the memorized slice, which is enough to move a leaderboard when margins are a point or two.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why can self-reported leaderboard scores be inflated?
&lt;/h2&gt;

&lt;p&gt;Self-reported scores carry three structural risks that have nothing to do with bad faith.&lt;/p&gt;

&lt;p&gt;First, the submitter controls the evaluation environment: the prompt template, the decoding settings, the number of few-shot examples, and sometimes the subset of items. None of those choices are contamination, but they compound with it.&lt;/p&gt;

&lt;p&gt;Second, nobody fully audits the training corpus. For most open releases the pretraining data is described at a high level, not published item by item. A team can honestly say they did not intend to train on a benchmark while still having ingested it through a scraped mirror. Intent does not change the outcome.&lt;/p&gt;

&lt;p&gt;Third, benchmarks age. A dataset released three years ago has had three years to be copied, quoted, reformatted, and re-uploaded. The older and more popular a benchmark is, the more likely its items are somewhere in a modern crawl. This is why a model can post a record score on a classic benchmark and a mediocre score on a freshly built private one.&lt;/p&gt;

&lt;p&gt;None of this means every high score is fake. It means a score reported without a contamination check is an unverified claim, and the burden of evidence sits with whoever is pointing at the number.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you detect train/test overlap? Three methods
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. N-gram overlap
&lt;/h3&gt;

&lt;p&gt;The most direct test: take each benchmark item, slide a window of n consecutive tokens or words across it, and check whether those n-grams appear in the training corpus (or in a proxy for it). High overlap at n of 8, 13, or 50 is strong evidence that the item, or a chunk of it, was seen during training. This is the method used by several major lab technical reports, which typically flag an item as contaminated when a sufficiently long n-gram matches. The strength of n-gram overlap is that it is cheap, interpretable, and does not need model internals. Its weakness is that it misses paraphrases and translations, so it gives a lower bound on contamination, never an upper bound.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Canary strings
&lt;/h3&gt;

&lt;p&gt;A canary is a unique, random, unlikely-to-occur-naturally identifier that benchmark authors embed in their dataset specifically so they can later test whether it leaked. If a model can reproduce or recognize the canary GUID, the dataset was in its training data, full stop. BIG-bench popularized this with an embedded canary GUID. The limitation is that canaries only work if the benchmark author planted one and the dataset was consumed with the canary intact; stripped or reformatted copies defeat it. Still, when a canary test fires, it is close to conclusive.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Membership inference and behavioral tests
&lt;/h3&gt;

&lt;p&gt;When you cannot inspect the training data at all, you probe the model's behavior. Membership inference asks whether the model treats a specific example as something it has seen before. Practical variants include: comparing perplexity on benchmark items versus closely matched held-out items (memorized items often have suspiciously low loss), the guided-prompting trick (does the model complete a test item far better when primed with the dataset name?), and option-order sensitivity (a model that memorized a multiple-choice answer is unusually robust to the correct answer's position while a reasoning model is not). These tests are noisier than n-gram or canary checks and need careful controls, but they are the only options for closed training data.&lt;/p&gt;

&lt;h2&gt;
  
  
  A runnable n-gram overlap check
&lt;/h2&gt;

&lt;p&gt;Here is a compact, dependency-free n-gram overlap checker. Give it your benchmark items and any text you suspect may overlap with training (a crawl shard, a scraped page, a quoted dataset). It reports the fraction of each item's n-grams that also appear in the reference text.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;collections&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Counter&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ngrams&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;13&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Word-level n-grams, lowercased and whitespace-normalized.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;toks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;toks&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;toks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;overlap_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reference_ngrams&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;13&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Fraction of the item&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s n-grams that appear in the reference set.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;item_grams&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;ngrams&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;item_grams&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;
    &lt;span class="n"&gt;hits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;g&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;item_grams&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;g&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;reference_ngrams&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;hits&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item_grams&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;build_reference&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;corpus_texts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;13&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Precompute the n-gram set for everything you can see of training data.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;ref&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;corpus_texts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;ngrams&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;ref&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# Stand-in for a crawl shard / scraped page you suspect overlaps training.
&lt;/span&gt;    &lt;span class="n"&gt;training_proxy&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;the capital of the fictional country of zubrowka is lutz &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;and its currency is the klubeck used since the year 1932&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;benchmark_items&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The capital of the fictional country of Zubrowka is Lutz &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;and its currency is the Klubeck used since the year 1932&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# leaked
&lt;/span&gt;        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain why a positive confidence interval overlap means &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;two leaderboard scores may not be distinguishable at all&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# clean
&lt;/span&gt;    &lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="n"&gt;N&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;13&lt;/span&gt;
    &lt;span class="n"&gt;ref&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;build_reference&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;training_proxy&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;N&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;benchmark_items&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;overlap_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;N&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;flag&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CONTAMINATED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;clean&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;item &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;N&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;-gram overlap -&amp;gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;flag&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Expected output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;item 0: 100% 13-gram overlap -&amp;gt; CONTAMINATED
item 1: 0% 13-gram overlap -&amp;gt; clean
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two practical notes. First, choose n deliberately: small n (3 to 5) produces false positives on common phrasing, large n (20 or more) misses chopped-up leaks, and 8 to 13 is a reasonable default for prose. Second, this is a lower bound. A 0 percent score proves only that there was no verbatim long-span match against the text you checked; paraphrase and translation slip right past it, which is why you pair it with behavioral tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical contamination checklist
&lt;/h2&gt;

&lt;p&gt;Run this before you cite or compare any benchmark number.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Is the benchmark newer than the model's data cutoff? A test built after training data was frozen cannot have leaked. Prefer recent or private benchmarks for headline comparisons.&lt;/li&gt;
&lt;li&gt;Did the authors plant a canary, and did the reporter check it? A passed canary test is cheap evidence and should be mentioned.&lt;/li&gt;
&lt;li&gt;Was an n-gram overlap report published against the training corpus (or a representative sample)? If the training data is open, this should exist. If it is not, note that the overlap is simply unknown.&lt;/li&gt;
&lt;li&gt;Is there a gap between the model's score on the popular benchmark and on a fresh, structurally similar one? A large drop on the new test is a classic contamination signature.&lt;/li&gt;
&lt;li&gt;Are decoding settings, prompt template, and few-shot count disclosed and held constant across the models being compared? Contamination aside, undisclosed harness choices make numbers incomparable.&lt;/li&gt;
&lt;li&gt;Is the score reported with uncertainty, and is the margin over the next model larger than that uncertainty? A contaminated slice often shows up as a margin that vanishes under resampling.&lt;/li&gt;
&lt;li&gt;Who ran the evaluation? Self-reported with no third-party reproduction is an unverified claim, not a result. Independent reproduction on held-out data is the gold standard.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If a claim fails several of these, the right posture is not "the model is cheating." It is "this number is unverified, and here is specifically what would make it trustworthy."&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is benchmark contamination always intentional?&lt;/strong&gt;&lt;br&gt;
No, and usually it is not. The most common path is accidental ingestion through a web crawl, since benchmarks live on the same internet that pretraining corpora are scraped from. Accidental contamination inflates scores exactly as much as deliberate contamination, which is why detection matters more than assigning blame.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does a high n-gram overlap score prove a model will fail on real tasks?&lt;/strong&gt;&lt;br&gt;
It proves the benchmark number is unreliable for that model, not that the model is incapable. The correct response is to re-evaluate on a clean, uncontaminated test and trust that number instead. Contamination invalidates a measurement; it does not by itself measure capability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not just keep all benchmarks private?&lt;/strong&gt;&lt;br&gt;
Private benchmarks resist contamination but sacrifice reproducibility and community trust, since nobody else can inspect the items or rerun the evaluation. The common compromise is a public benchmark with a planted canary, a held-out private slice, and periodic refreshes so the test can be rebuilt once the old version has saturated the internet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What n should I use for an overlap check?&lt;/strong&gt;&lt;br&gt;
For natural-language items, 8 to 13 words is a reasonable default. Go smaller and you flag ordinary phrasing as a match; go much larger and you miss leaks that were reformatted or chopped into pieces. Report the n you used, because the overlap fraction is meaningless without it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I detect contamination without access to the training data?&lt;/strong&gt;&lt;br&gt;
Yes, but only with behavioral tests: perplexity gaps between benchmark items and matched held-out items, guided-prompting sensitivity, and answer-position robustness. These are noisier than n-gram or canary checks, so treat them as evidence that accumulates rather than a single decisive test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;When a 0.4-Point Lead Means Nothing: Reading Open-Model Leaderboard Margins Like a Statistician&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;How Hugging Face Official Benchmark Leaderboards Actually Work: .eval_results YAML, the base_model Filter, and the 30% the Default View Hides&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>machinelearning</category>
      <category>llm</category>
      <category>benchmark</category>
      <category>datascience</category>
    </item>
    <item>
      <title>When a 0.4-Point Lead Means Nothing: Reading Open-Model Leaderboard Margins Like a Statistician</title>
      <dc:creator>Ward Ed</dc:creator>
      <pubDate>Tue, 06 Oct 2026 00:44:19 +0000</pubDate>
      <link>https://dev.to/ward_ed_6b5e6aa8ded94a987/when-a-04-point-lead-means-nothing-reading-open-model-leaderboard-margins-like-a-statistician-5do7</link>
      <guid>https://dev.to/ward_ed_6b5e6aa8ded94a987/when-a-04-point-lead-means-nothing-reading-open-model-leaderboard-margins-like-a-statistician-5do7</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A leaderboard rank is a point estimate, not a fact. Every accuracy score carries a confidence interval, and on most public evals that interval is wide enough to swallow the gap between the top few models.&lt;/li&gt;
&lt;li&gt;You can estimate the error bar yourself from two numbers you already have: the accuracy and the number of eval items. For a 300-item test at 70 percent accuracy, the 95 percent confidence half-width is about 5 points. A 0.4-point lead inside that band is noise.&lt;/li&gt;
&lt;li&gt;For paired comparisons on the same questions, use McNemar's test, not two separate intervals. It is strictly more sensitive and it is what the gap actually deserves.&lt;/li&gt;
&lt;li&gt;Before trusting any ranking, ask whether the benchmark can even discriminate. If every top model scores above 95 percent, the test is saturated and the order is close to random.&lt;/li&gt;
&lt;li&gt;Popularity is not quality. The most-liked open models on the Hub (cover chart, pulled 2026-10-06) are a social signal, not an accuracy ranking. Treat the two as separate axes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the second post in &lt;strong&gt;Leaderboard Forensics&lt;/strong&gt;. The first covered how the Hub's benchmark plumbing actually works. This one is about the single most common misreading of those numbers: believing a rank order that the sample size cannot support.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why do the top open models cluster so tightly?
&lt;/h2&gt;

&lt;p&gt;Open-model evaluation has matured to the point where the frontier is crowded. On a typical reasoning or knowledge benchmark, the top ten entries often sit inside a three-point band. The cover chart above shows the eight most-liked open text-generation repositories on the Hugging Face Hub as of 2026-10-06, and likes are a popularity proxy, but the same crowding shows up in accuracy tables.&lt;/p&gt;

&lt;p&gt;Here is the uncomfortable part. When scores cluster, the rank order is mostly determined by measurement noise rather than capability. The model in position one and the model in position four may be statistically indistinguishable. The leaderboard still has to print them in some order, so it prints the noise.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D400%26c%3D%257B%2522type%2522%253A%2520%2522line%2522%252C%2520%2522data%2522%253A%2520%257B%2522labels%2522%253A%2520%255B%2522100%2522%252C%2520%2522300%2522%252C%2520%25221000%2522%252C%2520%25223000%2522%252C%2520%252210000%2522%255D%252C%2520%2522datasets%2522%253A%2520%255B%257B%2522label%2522%253A%2520%252295%2525%2520CI%2520half-width%2520%2528pp%2529%2520at%2520p%253D0.70%2522%252C%2520%2522data%2522%253A%2520%255B8.98%252C%25205.19%252C%25202.84%252C%25201.64%252C%25200.9%255D%252C%2520%2522borderColor%2522%253A%2520%2522%2523dc2626%2522%252C%2520%2522fill%2522%253A%2520false%257D%255D%257D%252C%2520%2522options%2522%253A%2520%257B%2522title%2522%253A%2520%257B%2522display%2522%253A%2520true%252C%2520%2522text%2522%253A%2520%2522How%2520many%2520eval%2520items%2520you%2520need%2520before%2520small%2520gaps%2520mean%2520anything%2520%2528illustrative%2529%2522%257D%252C%2520%2522scales%2522%253A%2520%257B%2522xAxes%2522%253A%2520%255B%257B%2522scaleLabel%2522%253A%2520%257B%2522display%2522%253A%2520true%252C%2520%2522labelString%2522%253A%2520%2522Number%2520of%2520eval%2520items%2520%2528log-ish%2529%2522%257D%257D%255D%257D%257D%257D" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D400%26c%3D%257B%2522type%2522%253A%2520%2522line%2522%252C%2520%2522data%2522%253A%2520%257B%2522labels%2522%253A%2520%255B%2522100%2522%252C%2520%2522300%2522%252C%2520%25221000%2522%252C%2520%25223000%2522%252C%2520%252210000%2522%255D%252C%2520%2522datasets%2522%253A%2520%255B%257B%2522label%2522%253A%2520%252295%2525%2520CI%2520half-width%2520%2528pp%2529%2520at%2520p%253D0.70%2522%252C%2520%2522data%2522%253A%2520%255B8.98%252C%25205.19%252C%25202.84%252C%25201.64%252C%25200.9%255D%252C%2520%2522borderColor%2522%253A%2520%2522%2523dc2626%2522%252C%2520%2522fill%2522%253A%2520false%257D%255D%257D%252C%2520%2522options%2522%253A%2520%257B%2522title%2522%253A%2520%257B%2522display%2522%253A%2520true%252C%2520%2522text%2522%253A%2520%2522How%2520many%2520eval%2520items%2520you%2520need%2520before%2520small%2520gaps%2520mean%2520anything%2520%2528illustrative%2529%2522%257D%252C%2520%2522scales%2522%253A%2520%257B%2522xAxes%2522%253A%2520%255B%257B%2522scaleLabel%2522%253A%2520%257B%2522display%2522%253A%2520true%252C%2520%2522labelString%2522%253A%2520%2522Number%2520of%2520eval%2520items%2520%2528log-ish%2529%2522%257D%257D%255D%257D%257D%257D" alt="95 percent confidence half-width versus eval size, illustrative" width="1600" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The curve above is illustrative. It plots the 95 percent confidence half-width for a model scoring 70 percent, as a function of how many questions the benchmark contains. At 100 items the half-width is about 9 points. At 1,000 items it drops to roughly 2.8 points. Only around 10,000 items does it fall under one point. Most popular public benchmarks live at the left end of that curve.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I compute the error bar on a single accuracy score?
&lt;/h2&gt;

&lt;p&gt;Accuracy on a fixed test set is a binomial proportion. If a model answers &lt;code&gt;k&lt;/code&gt; of &lt;code&gt;n&lt;/code&gt; questions correctly, the estimated accuracy is &lt;code&gt;p = k / n&lt;/code&gt;, and the standard error is &lt;code&gt;sqrt(p * (1 - p) / n)&lt;/code&gt;. A rough 95 percent interval is &lt;code&gt;p plus or minus 1.96 * SE&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The Wald interval above is fine for a quick gut check, but it misbehaves near 0 and 100 percent and for small &lt;code&gt;n&lt;/code&gt;. The Wilson interval is better behaved and still closed-form. Here is a compact, dependency-light helper.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wilson_interval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;z&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.96&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;95% Wilson score interval for a binomial proportion.
    k = correct answers, n = total eval items.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;
    &lt;span class="n"&gt;denom&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;z&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;z&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;
    &lt;span class="n"&gt;center&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;z&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;z&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;denom&lt;/span&gt;
    &lt;span class="n"&gt;half&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;z&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sqrt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;z&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;z&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;denom&lt;/span&gt;
    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;center&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;half&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;center&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;half&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hi&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;wilson_interval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.70&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;n=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;  acc=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;  95% CI=[&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;lo&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;hi&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;]  width=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;hi&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;lo&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;n=  100  acc=0.700  95% CI=[0.604, 0.782]  width=0.178
n=  300  acc=0.700  95% CI=[0.646, 0.748]  width=0.102
n= 1000  acc=0.700  95% CI=[0.671, 0.727]  width=0.056
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read that bottom line carefully. Even with 1,000 questions, the interval around a 70 percent score spans almost six points. If two models are reported as 70.3 and 69.9 on that benchmark, the difference is well inside the noise floor. The rank order is not evidence of anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  When is a leaderboard gap actually significant?
&lt;/h2&gt;

&lt;p&gt;The single-score interval is the wrong tool for comparing two models, because it ignores that they were graded on the same questions. Scores on a shared test set are correlated. The correct instrument is McNemar's test, which looks only at the items where the two models disagree.&lt;/p&gt;

&lt;p&gt;Build a 2x2 table of agreement. Let &lt;code&gt;b&lt;/code&gt; be the number of questions where model A is right and model B is wrong, and &lt;code&gt;c&lt;/code&gt; the reverse. The questions both got right or both got wrong carry no information about which is better. McNemar's statistic is &lt;code&gt;(b - c)^2 / (b + c)&lt;/code&gt;, compared against a chi-square distribution with one degree of freedom.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;scipy.stats&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;chi2&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;mcnemar&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;b: A right &amp;amp; B wrong; c: A wrong &amp;amp; B right.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt;
    &lt;span class="n"&gt;stat&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;abs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;  &lt;span class="c1"&gt;# continuity-corrected
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;chi2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cdf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;stat&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Two models 0.4 points apart on a 300-item eval:
# say 18 items where A wins, 17 where B wins, rest tied.
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;p-value = &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;mcnemar&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;18&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;17&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# p-value = 1.000
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A one-item edge out of 300 produces a p-value of essentially 1.0. There is no detectable difference. You would need the disagreement counts to be lopsided, not the headline accuracy to be marginally higher, before the gap earns a rank.&lt;/p&gt;

&lt;p&gt;This is why serious evaluation reports publish disagreement matrices, bootstrap confidence intervals, or at least the per-benchmark item counts. If a leaderboard gives you only a sorted column of three-decimal numbers, it is giving you false precision.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I tell whether a benchmark can discriminate at all?
&lt;/h2&gt;

&lt;p&gt;Before arguing about margins, check that the test has headroom. A benchmark that is saturated, where the top cluster all score above 95 percent, cannot separate frontier models no matter how many items it has. The remaining questions are either ambiguous, mislabeled, or trivially easy, and the order among the leaders is driven by which model happened to win the coin flips on the hard residue.&lt;/p&gt;

&lt;p&gt;Three quick discrimination checks:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ceiling check.&lt;/strong&gt; If the top five are all within one point of 100 percent, the benchmark is done. Move on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Baseline check.&lt;/strong&gt; Compute what a trivial strategy scores. If majority-class guessing or a tiny n-gram model gets 90 percent, the benchmark measures format compliance, not capability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spread check.&lt;/strong&gt; Look at the standard deviation across all submitted models. A benchmark where everyone scores 68 to 72 has almost no discriminative range left.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A benchmark earns your trust when the baseline is low, the ceiling is far away, and the spread across models is several times the single-score confidence width.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should I actually do with a leaderboard, then?
&lt;/h2&gt;

&lt;p&gt;Treat the leaderboard as a filter, not a verdict. Use it to shortlist the handful of models in the top band, then evaluate those candidates on your own held-out data that resembles your real traffic. The rank inside that top band is the least reliable information on the page. The membership of the band is the useful part.&lt;/p&gt;

&lt;p&gt;And keep popularity and accuracy on separate axes. The cover chart is real Hub data, and it tells you what the community has adopted, which matters for tooling, documentation, and community support. It says nothing about which model is most accurate on your task. Both signals are valuable. Conflating them is how teams end up defending a 0.4-point lead that never existed.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is the Wald interval ever good enough?&lt;/strong&gt;&lt;br&gt;
For a sanity check on a single score with a few hundred or more items, yes. For scores near 0 or 100 percent, or for small &lt;code&gt;n&lt;/code&gt;, switch to Wilson, which stays inside the valid range and does not collapse to zero width at the extremes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why McNemar instead of a two-sample proportion test?&lt;/strong&gt;&lt;br&gt;
Because both models answer the same questions, their scores are paired and correlated. A two-sample test assumes independence and will be too conservative. McNemar conditions on the disagreements, which is exactly the information that distinguishes the two models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How many eval items do I need to resolve a one-point gap?&lt;/strong&gt;&lt;br&gt;
Roughly, to get a 95 percent half-width near one point at mid-range accuracy you need on the order of 10,000 independent items. Most public benchmarks have far fewer, which is why one-point gaps are usually unresolvable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do bootstrap confidence intervals help?&lt;/strong&gt;&lt;br&gt;
Yes, especially when scoring is not a clean right-or-wrong binary, such as graded or rubric-scored tasks. Resample the items with replacement a few thousand times, recompute the metric, and read the 2.5 and 97.5 percentiles. It makes no distributional assumption and handles weird metrics gracefully.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does a larger model automatically win on a saturated benchmark?&lt;/strong&gt;&lt;br&gt;
No. On a saturated test the residual questions are dominated by noise and label quality, so scale buys you almost nothing. The honest move is to find a harder benchmark with headroom, or build one from your own data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I trust the three decimals a leaderboard prints?&lt;/strong&gt;&lt;br&gt;
Only if the item count justifies them. Three decimals on a 300-item eval is false precision. The real resolution of that test is closer to one or two points, so mentally round the column before you rank it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/ward_ed_6b5e6aa8ded94a987/how-hugging-face-official-benchmark-leaderboards-actually-work-evalresults-yaml-the-basemodel-29bo"&gt;How Hugging Face Official Benchmark Leaderboards Actually Work&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Methodology note: Hub likes in the cover chart were pulled live from the public Hugging Face models API (sort by likes, text-generation filter) on 2026-10-06. The confidence-width curve and the McNemar example use illustrative inputs and are labeled as such.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>huggingface</category>
      <category>llm</category>
      <category>benchmark</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>How Hugging Face Official Benchmark Leaderboards Actually Work: .eval_results YAML, the base_model Filter, and the 30% the Default View Hides</title>
      <dc:creator>Ward Ed</dc:creator>
      <pubDate>Mon, 05 Oct 2026 08:56:56 +0000</pubDate>
      <link>https://dev.to/ward_ed_6b5e6aa8ded94a987/how-hugging-face-official-benchmark-leaderboards-actually-work-evalresults-yaml-the-basemodel-29bo</link>
      <guid>https://dev.to/ward_ed_6b5e6aa8ded94a987/how-hugging-face-official-benchmark-leaderboards-actually-work-evalresults-yaml-the-basemodel-29bo</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Hugging Face currently tags 48 datasets as official benchmarks (&lt;code&gt;benchmark:official&lt;/code&gt;). Their leaderboards are not uploaded by the benchmark owners: they are assembled automatically from small &lt;code&gt;.eval_results/*.yaml&lt;/code&gt; files that model authors commit to their own model repos (or propose through pull requests). On October 5, 2026 I queried every board twice through the public API. The default view returned &lt;strong&gt;1,023&lt;/strong&gt; entries. With &lt;code&gt;?base_model=false&lt;/code&gt;, which keeps derivative models, it returned &lt;strong&gt;1,469&lt;/strong&gt;. So &lt;strong&gt;446 entries (30.4%) never show up in the default view&lt;/strong&gt;, and the deciding field is the &lt;code&gt;base_model&lt;/code&gt; key in each model card. Another finding: &lt;strong&gt;772 of the 1,023 visible entries (75.5%) come from open pull requests&lt;/strong&gt;, and &lt;strong&gt;none&lt;/strong&gt; were flagged &lt;code&gt;verified&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D1000%26h%3D420%26c%3D%257B%2522type%2522%253A%2522bar%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522Default%2520view%2522%252C%2522Incl.%2520derivatives%2520%2528base_model%253Dfalse%2529%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522label%2522%253A%2522Entries%2520across%252048%2520official%2520HF%2520benchmark%2520leaderboards%2522%252C%2522data%2522%253A%255B1023%252C1469%255D%252C%2522backgroundColor%2522%253A%255B%2522%25232563eb%2522%252C%2522%2523f59e0b%2522%255D%257D%255D%257D%252C%2522options%2522%253A%257B%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522446%2520of%25201%252C469%2520leaderboard%2520entries%2520%252830.4%2525%2529%2520are%2520hidden%2520by%2520default%2522%257D%257D%257D" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D1000%26h%3D420%26c%3D%257B%2522type%2522%253A%2522bar%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522Default%2520view%2522%252C%2522Incl.%2520derivatives%2520%2528base_model%253Dfalse%2529%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522label%2522%253A%2522Entries%2520across%252048%2520official%2520HF%2520benchmark%2520leaderboards%2522%252C%2522data%2522%253A%255B1023%252C1469%255D%252C%2522backgroundColor%2522%253A%255B%2522%25232563eb%2522%252C%2522%2523f59e0b%2522%255D%257D%255D%257D%252C%2522options%2522%253A%257B%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522446%2520of%25201%252C469%2520leaderboard%2520entries%2520%252830.4%2525%2529%2520are%2520hidden%2520by%2520default%2522%257D%257D%257D" alt="Visible vs hidden entries" width="2000" height="840"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What are Hugging Face official benchmark leaderboards?
&lt;/h2&gt;

&lt;p&gt;An official benchmark is just a dataset repo that carries the &lt;code&gt;benchmark:official&lt;/code&gt; tag. The Hub renders a leaderboard widget on its dataset page. Each row on that widget is one model's score on that benchmark. A dataset owner does not have to run any evaluation service for this. The Hub collects scores that model repos self-report and matches them to the dataset.&lt;/p&gt;

&lt;p&gt;You can pull the list of boards with one request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://huggingface.co/api/datasets?filter=benchmark:official&amp;amp;limit=200"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import sys,json; d=json.load(sys.stdin); print(len(d)); print('&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;'.join(x['id'] for x in d))"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When I ran it on October 5, 2026, it returned 48 datasets. Some of the better known ones: &lt;code&gt;TIGER-Lab/MMLU-Pro&lt;/code&gt;, &lt;code&gt;Idavidrein/gpqa&lt;/code&gt;, &lt;code&gt;cais/hle&lt;/code&gt;, &lt;code&gt;SWE-bench/SWE-bench_Verified&lt;/code&gt;, &lt;code&gt;hf-audio/open-asr-leaderboard&lt;/code&gt;, &lt;code&gt;openai/gsm8k&lt;/code&gt;, &lt;code&gt;MathArena/aime_2026&lt;/code&gt;, &lt;code&gt;MMMU/MMMU_Pro&lt;/code&gt; and &lt;code&gt;harborframework/terminal-bench-2.0&lt;/code&gt;. The list also includes newer, narrower boards such as &lt;code&gt;llamaindex/ParseBench&lt;/code&gt;, &lt;code&gt;llamaindex/ExtractBench&lt;/code&gt;, &lt;code&gt;LEXam-Benchmark/LEXam&lt;/code&gt;, &lt;code&gt;FutureMa/EvasionBench&lt;/code&gt; and &lt;code&gt;crosbylegal/RedlineBench&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Four of the 48 (&lt;code&gt;tiiuae/PBench&lt;/code&gt;, &lt;code&gt;mercor/ACE&lt;/code&gt;, &lt;code&gt;ChrisHayduk/nanofold-public&lt;/code&gt;, &lt;code&gt;harborframework/terminal-bench-science&lt;/code&gt;) had zero entries in both views when I checked. Tagged boards can exist before anyone has reported a score to them.&lt;/p&gt;

&lt;h2&gt;
  
  
  How are Hugging Face leaderboard entries built from .eval_results YAML?
&lt;/h2&gt;

&lt;p&gt;Every row traces back to one YAML file in a &lt;strong&gt;model&lt;/strong&gt; repo, under a directory named &lt;code&gt;.eval_results/&lt;/code&gt;. The leaderboard API includes the path in each row, so you can see it yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://huggingface.co/api/datasets/Idavidrein/gpqa/leaderboard"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import sys,json; e=json.load(sys.stdin)[1]; print({k:e[k] for k in ('rank','modelId','value','filename','verified','pullRequest')})"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each row has these fields: &lt;code&gt;rank&lt;/code&gt;, &lt;code&gt;filename&lt;/code&gt;, &lt;code&gt;value&lt;/code&gt;, &lt;code&gt;verified&lt;/code&gt;, &lt;code&gt;source&lt;/code&gt;, &lt;code&gt;pullRequest&lt;/code&gt;, &lt;code&gt;modelId&lt;/code&gt;, &lt;code&gt;author&lt;/code&gt;, &lt;code&gt;lower_is_better&lt;/code&gt; and &lt;code&gt;num_parameters&lt;/code&gt;. &lt;code&gt;filename&lt;/code&gt; is the YAML path inside the model repo, and &lt;code&gt;pullRequest&lt;/code&gt; holds a PR number when the file lives on a PR ref instead of &lt;code&gt;main&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Here is a real file. &lt;code&gt;moonshotai/Kimi-K3&lt;/code&gt; keeps five of them on &lt;code&gt;main&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://huggingface.co/api/models/moonshotai/Kimi-K3/tree/main/.eval_results"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import sys,json; [print(f['path'], f['size']) for f in json.load(sys.stdin)]"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.eval_results/apex-agents.yaml 157
.eval_results/deep-swe.yaml 156
.eval_results/gpqa.yaml 152
.eval_results/hle.yaml 139
.eval_results/moonshotai__Kimi-K3.yaml 211
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The GPQA one (&lt;code&gt;/raw/main/.eval_results/gpqa.yaml&lt;/code&gt;) is just 152 bytes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;dataset&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Idavidrein/gpqa&lt;/span&gt;
    &lt;span class="na"&gt;task_id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;diamond&lt;/span&gt;
  &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;93.5&lt;/span&gt;
  &lt;span class="na"&gt;source&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://huggingface.co/moonshotai/Kimi-K3&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Model Card&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole contract:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;dataset.id&lt;/code&gt; is the benchmark dataset repo. It has to match one of the official board IDs exactly, or the score goes nowhere.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;dataset.task_id&lt;/code&gt; picks a sub-task or split when a board has several (for GPQA, &lt;code&gt;diamond&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;value&lt;/code&gt; is the number that gets ranked. The board's &lt;code&gt;lower_is_better&lt;/code&gt; flag decides the sort order (WER on the ASR board, for example).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;source&lt;/code&gt; is an attribution link. In practice it usually points back at the model card.&lt;/li&gt;
&lt;li&gt;Optional keys show up in the wild too, like &lt;code&gt;date&lt;/code&gt; and a free-text &lt;code&gt;notes&lt;/code&gt; field for the protocol (sample count, temperature, token budget, precision). Use them. They are the only place a reader can see how the number was produced.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The file is a YAML &lt;strong&gt;list&lt;/strong&gt;, so a single file can hold several benchmark results. Many repos still use one file per benchmark, which keeps diffs and PRs small.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does my model not appear on the Hugging Face leaderboard? (the base_model filter)
&lt;/h2&gt;

&lt;p&gt;This is where most authors get stuck. The leaderboard endpoint has a parameter called &lt;code&gt;base_model&lt;/code&gt;. Without it, the API (and the default widget) &lt;strong&gt;drops models whose card declares a &lt;code&gt;base_model&lt;/code&gt;&lt;/strong&gt;, meaning fine-tunes, merges, quantizations and adapters. With &lt;code&gt;?base_model=false&lt;/code&gt;, those derivatives come back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;B&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;Idavidrein/gpqa
curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://huggingface.co/api/datasets/&lt;/span&gt;&lt;span class="nv"&gt;$B&lt;/span&gt;&lt;span class="s2"&gt;/leaderboard"&lt;/span&gt; | python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import sys,json;print(len(json.load(sys.stdin)))"&lt;/span&gt;
curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://huggingface.co/api/datasets/&lt;/span&gt;&lt;span class="nv"&gt;$B&lt;/span&gt;&lt;span class="s2"&gt;/leaderboard?base_model=false"&lt;/span&gt; | python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import sys,json;print(len(json.load(sys.stdin)))"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For GPQA that is 111 vs 175. I ran the same comparison for all 48 boards with this script:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;concurrent.futures&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ThreadPoolExecutor&lt;/span&gt;

&lt;span class="n"&gt;API&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://huggingface.co/api/datasets&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;boards&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;API&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;?filter=benchmark:official&amp;amp;limit=200&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bid&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;API&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;bid&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/leaderboard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;API&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;bid&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/leaderboard?base_model=false&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;bid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;ThreadPoolExecutor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;boards&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="n"&gt;dflt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;allr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;dflt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;allr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;allr&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;dflt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;allr&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; hidden&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output on 2026-10-05: &lt;code&gt;48 1023 1469 30.4% hidden&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D420%26c%3D%257B%2522type%2522%253A%2522horizontalBar%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522MMLU-Pro%2522%252C%2522gpqa%2522%252C%2522hle%2522%252C%2522SWE-bench_Verified%2522%252C%2522open-asr-leaderboard%2522%252C%2522SWE-bench_Pro%2522%252C%2522terminal-bench-2.0%2522%252C%2522ScreenSpot-Pro%2522%252C%2522gsm8k%2522%252C%2522terminal-bench-2.1%2522%252C%2522aime_2026%2522%252C%2522ParseBench%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522label%2522%253A%2522Default%2522%252C%2522data%2522%253A%255B141%252C111%252C89%252C67%252C62%252C44%252C38%252C23%252C18%252C34%252C25%252C30%255D%252C%2522backgroundColor%2522%253A%2522%25232563eb%2522%257D%252C%257B%2522label%2522%253A%2522Hidden%2520derivatives%2522%252C%2522data%2522%253A%255B75%252C64%252C19%252C22%252C24%252C17%252C17%252C25%252C28%252C10%252C17%252C11%255D%252C%2522backgroundColor%2522%253A%2522%2523f59e0b%2522%257D%255D%257D%252C%2522options%2522%253A%257B%2522scales%2522%253A%257B%2522xAxes%2522%253A%255B%257B%2522stacked%2522%253Atrue%257D%255D%252C%2522yAxes%2522%253A%255B%257B%2522stacked%2522%253Atrue%257D%255D%257D%252C%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522Top%252012%2520boards%253A%2520visible%2520vs%2520hidden%2520entries%2522%257D%257D%257D" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D420%26c%3D%257B%2522type%2522%253A%2522horizontalBar%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522MMLU-Pro%2522%252C%2522gpqa%2522%252C%2522hle%2522%252C%2522SWE-bench_Verified%2522%252C%2522open-asr-leaderboard%2522%252C%2522SWE-bench_Pro%2522%252C%2522terminal-bench-2.0%2522%252C%2522ScreenSpot-Pro%2522%252C%2522gsm8k%2522%252C%2522terminal-bench-2.1%2522%252C%2522aime_2026%2522%252C%2522ParseBench%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522label%2522%253A%2522Default%2522%252C%2522data%2522%253A%255B141%252C111%252C89%252C67%252C62%252C44%252C38%252C23%252C18%252C34%252C25%252C30%255D%252C%2522backgroundColor%2522%253A%2522%25232563eb%2522%257D%252C%257B%2522label%2522%253A%2522Hidden%2520derivatives%2522%252C%2522data%2522%253A%255B75%252C64%252C19%252C22%252C24%252C17%252C17%252C25%252C28%252C10%252C17%252C11%255D%252C%2522backgroundColor%2522%253A%2522%2523f59e0b%2522%257D%255D%257D%252C%2522options%2522%253A%257B%2522scales%2522%253A%257B%2522xAxes%2522%253A%255B%257B%2522stacked%2522%253Atrue%257D%255D%252C%2522yAxes%2522%253A%255B%257B%2522stacked%2522%253Atrue%257D%255D%257D%252C%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522Top%252012%2520boards%253A%2520visible%2520vs%2520hidden%2520entries%2522%257D%257D%257D" alt="Top 12 boards, visible vs hidden" width="1600" height="840"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;How much is hidden varies a lot from board to board:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Board&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;th&gt;With derivatives&lt;/th&gt;
&lt;th&gt;Hidden share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;TIGER-Lab/MMLU-Pro&lt;/td&gt;
&lt;td&gt;141&lt;/td&gt;
&lt;td&gt;216&lt;/td&gt;
&lt;td&gt;34.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Idavidrein/gpqa&lt;/td&gt;
&lt;td&gt;111&lt;/td&gt;
&lt;td&gt;175&lt;/td&gt;
&lt;td&gt;36.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cais/hle&lt;/td&gt;
&lt;td&gt;89&lt;/td&gt;
&lt;td&gt;108&lt;/td&gt;
&lt;td&gt;17.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SWE-bench/SWE-bench_Verified&lt;/td&gt;
&lt;td&gt;67&lt;/td&gt;
&lt;td&gt;89&lt;/td&gt;
&lt;td&gt;24.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;hf-audio/open-asr-leaderboard&lt;/td&gt;
&lt;td&gt;62&lt;/td&gt;
&lt;td&gt;86&lt;/td&gt;
&lt;td&gt;27.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;likaixin/ScreenSpot-Pro&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;td&gt;48&lt;/td&gt;
&lt;td&gt;52.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;openai/gsm8k&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;td&gt;46&lt;/td&gt;
&lt;td&gt;60.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMMU/MMMU_Pro&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;58.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;llamaindex/ExtractBench&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;56.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LiquidAI/ifstruct-v1.0&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;td&gt;66.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mteb/arguana&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;81.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;crosbylegal/RedlineBench&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern makes sense. Older or easier benchmarks like GSM8K attract many fine-tunes, so derivatives are the majority there and the default view shows a minority of the reported scores. On agentic and legal boards where most submissions are base models from labs (RedlineBench, skillsbench, Long-Horizon-Terminal-Bench), the two views are identical.&lt;/p&gt;

&lt;p&gt;The filter reads the model card metadata, not the weights. A model trained from scratch that lists a &lt;code&gt;base_model&lt;/code&gt; anyway (some teams do this to credit an architecture or tokenizer source) gets hidden just like a LoRA. A heavily modified derivative that leaves the field out shows up as if it were a base model. So the default view is a filter on &lt;strong&gt;declared&lt;/strong&gt; lineage, not real lineage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who verifies scores on Hugging Face official leaderboards?
&lt;/h2&gt;

&lt;p&gt;According to the API, nobody does yet. Of the 1,023 default-view rows, &lt;strong&gt;0&lt;/strong&gt; had &lt;code&gt;verified: true&lt;/code&gt;. Every number on these boards is self-reported by whoever controls the model repo or opens a PR to it.&lt;/p&gt;

&lt;p&gt;The PR path is bigger than most people expect. &lt;strong&gt;772 of 1,023 visible rows (75.5%)&lt;/strong&gt; carried a &lt;code&gt;pullRequest&lt;/code&gt; number. So the YAML sits on an open (unmerged) PR ref such as &lt;code&gt;refs/pr/1&lt;/code&gt;, not on &lt;code&gt;main&lt;/code&gt;. Reading those files directly needs the PR ref:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# file on an open PR, not on main&lt;/span&gt;
curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro/raw/refs%2Fpr%2F1/.eval_results/gsm8k.yaml"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you try &lt;code&gt;tree/main/.eval_results&lt;/code&gt; on such a model, you get &lt;code&gt;.eval_results does not exist on "main"&lt;/code&gt; even though the row shows on the leaderboard. That is expected behavior, not a bug.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D420%26c%3D%257B%2522type%2522%253A%2522doughnut%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522From%2520open%2520PRs%2520%2528772%2529%2522%252C%2522Merged%2520on%2520main%2520%2528251%2529%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522data%2522%253A%255B772%252C251%255D%252C%2522backgroundColor%2522%253A%255B%2522%2523f59e0b%2522%252C%2522%25232563eb%2522%255D%257D%255D%257D%252C%2522options%2522%253A%257B%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522Default-view%2520entries%2520by%2520source%2520ref%2520%2528verified%253A%25200%2529%2522%257D%257D%257D" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D800%26h%3D420%26c%3D%257B%2522type%2522%253A%2522doughnut%2522%252C%2522data%2522%253A%257B%2522labels%2522%253A%255B%2522From%2520open%2520PRs%2520%2528772%2529%2522%252C%2522Merged%2520on%2520main%2520%2528251%2529%2522%255D%252C%2522datasets%2522%253A%255B%257B%2522data%2522%253A%255B772%252C251%255D%252C%2522backgroundColor%2522%253A%255B%2522%2523f59e0b%2522%252C%2522%25232563eb%2522%255D%257D%255D%257D%252C%2522options%2522%253A%257B%2522title%2522%253A%257B%2522display%2522%253Atrue%252C%2522text%2522%253A%2522Default-view%2520entries%2520by%2520source%2520ref%2520%2528verified%253A%25200%2529%2522%257D%257D%257D" alt="Visible entries by source ref" width="1600" height="840"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Practical implications for anyone citing these boards:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A row is a claim, not a measurement. Follow the &lt;code&gt;source.url&lt;/code&gt; and read the &lt;code&gt;notes&lt;/code&gt;, if there are any.&lt;/li&gt;
&lt;li&gt;Protocols differ across rows on the same board. Majority voting over many samples, different thinking budgets, different task subsets: none of that is normalized. The &lt;code&gt;task_id&lt;/code&gt; field is the only structural guard, and it only separates named splits.&lt;/li&gt;
&lt;li&gt;Rows can come from someone other than the model owner, through a PR. Check the PR author before treating a number as the vendor's.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  How do I submit a model score to a Hugging Face benchmark leaderboard?
&lt;/h2&gt;

&lt;p&gt;There is no submission form. You commit a file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;huggingface_hub&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;HfApi&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;yaml&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;io&lt;/span&gt;

&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dataset&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Idavidrein/gpqa&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;diamond&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;value&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;71.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2026-10-05&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://huggingface.co/your-org/your-model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Model Card&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GPQA Diamond, 198 items, pass@1, temperature 0, bf16&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;span class="n"&gt;buf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;io&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;BytesIO&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;yaml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;safe_dump&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sort_keys&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="nc"&gt;HfApi&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;upload_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;path_or_fileobj&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;path_in_repo&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.eval_results/gpqa.yaml&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;repo_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your-org/your-model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;commit_message&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Add GPQA Diamond eval result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="c1"&gt;# create_pr=True,  # use this when you are not the repo owner
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The numbers above are placeholders for illustration. Then check whether the row appears, and &lt;strong&gt;which view&lt;/strong&gt; it appears in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;where_am_i&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;board&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://huggingface.co/api/datasets/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;board&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/leaderboard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;modelId&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?base_model=false&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;modelId&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;default_view&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;with_derivatives&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rank_default&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rank&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rank_all&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rank&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;where_am_i&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Idavidrein/gpqa&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your-org/your-model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;with_derivatives&lt;/code&gt; is true and &lt;code&gt;default_view&lt;/code&gt; is false, the &lt;code&gt;base_model&lt;/code&gt; field in your card is what hides you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklist for model authors before reporting a leaderboard rank
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Match &lt;code&gt;dataset.id&lt;/code&gt; exactly&lt;/strong&gt; to one of the 48 IDs from &lt;code&gt;filter=benchmark:official&lt;/code&gt;. Case and owner prefix matter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set &lt;code&gt;task_id&lt;/code&gt;&lt;/strong&gt; when the board has subsets. A GPQA "main" score sitting next to "diamond" scores is misleading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide on &lt;code&gt;base_model&lt;/code&gt; on purpose.&lt;/strong&gt; If your model really is a fine-tune, keeping the field is honest, but you will only appear in the derivative-inclusive view. If the field is left over from a template or only credits a tokenizer, fix the card. Don't remove it from a real derivative just to get a better spot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify with the parameterless API&lt;/strong&gt;, not the &lt;code&gt;?base_model=false&lt;/code&gt; URL and not a cached screenshot. The parameterless call is what most visitors see by default.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quote the right rank.&lt;/strong&gt; "Rank 3 including derivatives" and "rank 3" are different claims. With 30.4% of entries hidden, the gap can be large.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write the protocol in &lt;code&gt;notes&lt;/code&gt;&lt;/strong&gt;: sample count, decoding settings, token budget, precision, subset size. Nobody verifies rows centrally, so this is your credibility.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Merge your own PRs.&lt;/strong&gt; If a contributor opened the eval PR, the row already counts. Merging it puts the file on &lt;code&gt;main&lt;/code&gt; and makes the history auditable.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  How many official benchmark leaderboards are on Hugging Face?
&lt;/h3&gt;

&lt;p&gt;48 datasets carried the &lt;code&gt;benchmark:official&lt;/code&gt; tag on October 5, 2026. Four of them had no entries yet. Query &lt;code&gt;https://huggingface.co/api/datasets?filter=benchmark:official&amp;amp;limit=200&lt;/code&gt; for the current list.&lt;/p&gt;

&lt;h3&gt;
  
  
  What does &lt;code&gt;base_model=false&lt;/code&gt; do in the Hugging Face leaderboard API?
&lt;/h3&gt;

&lt;p&gt;It adds back models whose card declares a &lt;code&gt;base_model&lt;/code&gt; (fine-tunes, merges, quantizations, adapters). The default request leaves them out. Across all boards that is 1,469 entries vs 1,023, so 30.4% are hidden by default.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where does a Hugging Face leaderboard score come from?
&lt;/h3&gt;

&lt;p&gt;From a YAML file under &lt;code&gt;.eval_results/&lt;/code&gt; in the model repo, on &lt;code&gt;main&lt;/code&gt; or on an open PR ref. The leaderboard row's &lt;code&gt;filename&lt;/code&gt; field gives the path, and &lt;code&gt;pullRequest&lt;/code&gt; gives the PR number if there is one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Are Hugging Face official leaderboard scores verified?
&lt;/h3&gt;

&lt;p&gt;Not at the time of measurement. All 1,023 default-view rows had &lt;code&gt;verified: false&lt;/code&gt;. Treat each row as a self-reported claim and check its source and notes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why does my model show on the board in one browser link but not another?
&lt;/h3&gt;

&lt;p&gt;You are probably comparing the default view with the derivative-inclusive view. Run both API calls for your &lt;code&gt;modelId&lt;/code&gt;. If only the &lt;code&gt;?base_model=false&lt;/code&gt; call finds you, your model card's &lt;code&gt;base_model&lt;/code&gt; field explains it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can someone else add a score for my model?
&lt;/h3&gt;

&lt;p&gt;Yes. Anyone can open a pull request that adds &lt;code&gt;.eval_results/*.yaml&lt;/code&gt; to your repo, and 75.5% of visible rows came from PR refs when I measured. Review those PRs like any other change to your model card.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Methodology: all counts come from unauthenticated requests to the public Hugging Face API on 2026-10-05. Leaderboards change daily, so re-run the scripts above for current numbers.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>huggingface</category>
      <category>llm</category>
      <category>benchmark</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
