<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alex Fank</title>
    <description>The latest articles on DEV Community by Alex Fank (@alexfank).</description>
    <link>https://dev.to/alexfank</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4164144%2F14b83cac-de20-4e49-a70b-46b44f7aee2f.png</url>
      <title>DEV Community: Alex Fank</title>
      <link>https://dev.to/alexfank</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/alexfank"/>
    <language>en</language>
    <item>
      <title>Why averaging LLM benchmarks gives the wrong leaderboard</title>
      <dc:creator>Alex Fank</dc:creator>
      <pubDate>Mon, 05 Oct 2026 15:20:16 +0000</pubDate>
      <link>https://dev.to/alexfank/why-averaging-llm-benchmarks-gives-the-wrong-leaderboard-boc</link>
      <guid>https://dev.to/alexfank/why-averaging-llm-benchmarks-gives-the-wrong-leaderboard-boc</guid>
      <description>&lt;p&gt;Every few weeks new open-weight model is released with a table of benchmark results, and every few weeks we asked the same practical question: is it better than the one we already run? A single ranked list should answer that. Building one turned out to be harder than we expected, and our first attempt failed in an instructive way. This is the first of three posts about that work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem space
&lt;/h2&gt;

&lt;p&gt;There is no shortage of benchmarks. Epoch AI's capabilities index alone draws on &lt;a href="https://epoch.ai/data/eci-documentation" rel="noopener noreferrer"&gt;more than 50 distinct benchmarks&lt;/a&gt;. The trouble is that no model has been measured on all of them, and numbers that do exist are not always comparable.&lt;/p&gt;

&lt;p&gt;Vendors report results with their own prompts, harnesses and reasoning-effort settings. Hugging Face found that different implementations of MMLU &lt;a href="https://huggingface.co/blog/open-llm-leaderboard-mmlu" rel="noopener noreferrer"&gt;"give widely different numbers and even change the ranking order of the models"&lt;/a&gt;. Independent boards run models themselves, which helps, but every board makes its own choices, and a model appears only on the boards whose maintainers got to it.&lt;/p&gt;

&lt;p&gt;Benchmarks also age. The authors of MMLU-Pro &lt;a href="https://arxiv.org/abs/2406.01574" rel="noopener noreferrer"&gt;note&lt;/a&gt; that performance on earlier benchmarks had begun to plateau, making differences between models hard to discern. GSM8K is a familiar example: a freshly written look-alike set, GSM1k, showed &lt;a href="https://arxiv.org/abs/2405.00332" rel="noopener noreferrer"&gt;accuracy drops of up to 8%&lt;/a&gt; for some model families. New, harder benchmarks appear to separate the frontier models, and the older ones stay on the books.&lt;/p&gt;

&lt;p&gt;Some boards rotate their questions on purpose. LiveBench replaces &lt;a href="https://arxiv.org/abs/2406.19314" rel="noopener noreferrer"&gt;about one sixth of its questions in each update&lt;/a&gt;, so snapshot from one month is not strictly comparable to one from another. We treat each LiveBench snapshot as a unit and never mix scores across snapshots.&lt;/p&gt;

&lt;p&gt;The result is that every model is measured on a different, sparse subset of the available boards. Any composite has to decide what to do about the holes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we built first
&lt;/h2&gt;

&lt;p&gt;Our first version, which launched in June 2026, was deliberately plain. We normalized each result with 0 to 1 range between the board's random-guess baseline and its ceiling. Within each category (reasoning, coding, math and so on) we averaged whatever normalized scores a model had. The overall number was then an equal-weight mean of the category sub-scores. Missing entries were never filled with zeros; category a model had no data for was simply absent.&lt;/p&gt;

&lt;p&gt;To keep thin evidence out of the ranking we added a coverage gate: a model needed results in at least three categories, and at least one of them had to be reasoning or coding. There was no minimum number of benchmarks.&lt;/p&gt;

&lt;p&gt;This looked reasonable. It resembles the way several public leaderboards combine results, it is easy to explain, and no benchmark is privileged. We shipped it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it broke
&lt;/h2&gt;

&lt;p&gt;On 2026-09-20 we looked at the open-model ranking after adding several new boards and found something we could not defend. Phi 3.5 MoE Instruct was measured on three benchmarks: PIQA (0.886), GSM8K (0.887) and BoolQ (0.846). Its composite was 83.5. Kimi K3 had results on eighteen benchmarks, including some of the hardest current ones, and its composite was 66.7. The ranking put the three-benchmark model well above the eighteen-benchmark one. Overall, Kimi K3 sat at number 8, behind Llama 2 70B and Phi 3.5 MoE. The top three open entries each had exactly three boards.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5xkssn2ovoc195zlaanj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5xkssn2ovoc195zlaanj.png" alt="v1 leaderboard: three easy benchmarks beat eighteen hard ones" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The mechanism is straightforward once it is visible. Hard modern boards produce low scores, even for excellent models. A model that has been run on many of them gets lower mean for it. A model that was only ever submitted to saturated classics, where nearly everyone scores near the ceiling, is never exposed to that penalty. The average rewards avoiding difficult tests.&lt;/p&gt;

&lt;p&gt;We checked that this was not an artifact of the new boards. Recomputing without the six boards added that day still left the older models on top, with Kimi K3 at number 13.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two fixes we measured and rejected
&lt;/h2&gt;

&lt;p&gt;The first obvious fix is a minimum benchmark count. We tried it. Models with four to eleven easy boards still beat flagships, because the problem is which boards a model has, not how many.&lt;/p&gt;

&lt;p&gt;The second is to standardize each benchmark with a z-score across all models, so an easy board stops being easy. This also failed. A legacy board has a legacy population: a 2023 model compared against other 2023-era models still earns a high z-score. With z-scores over the full field, Kimi K3 only reached number 2, and Falcon 180B and StableBeluga2 were still in the top six.&lt;/p&gt;

&lt;p&gt;We kept a floor of four benchmarks anyway, partly because it matched how we already labeled confidence: three-benchmark composites swing wildly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix that worked for v2: an active panel
&lt;/h2&gt;

&lt;p&gt;What the failures had in common was that old boards were still voting. A benchmark carries information about today's models only while today's models are still being submitted to it. So we changed the rule: a benchmark counts toward the composite only if the newest result on it comes from a model released within the last 365 days.&lt;/p&gt;

&lt;p&gt;The rule is derived from the data rather than curated by hand. Nobody on our side decides that a board is obsolete; if a board goes quiet upstream, it leaves the panel by itself. With this rule, 27 boards were active, and ten saturated classics dropped out, among them GSM8K, HellaSwag, PIQA, BoolQ and MMLU. The four-benchmark floor stayed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ebhc4ug5cxj7ndc5ys1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ebhc4ug5cxj7ndc5ys1.png" alt="Active panel" width="800" height="1400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The chart shows the panel as it stands on 2026-10-05: with more boards imported since, 40 are active, and the same ten classics remain frozen.&lt;/p&gt;

&lt;p&gt;There was a cost, and we accepted it deliberately. 60 of 126 previously ranked models lost their overall rank, because they had been scored mostly on the frozen boards and fell below the floor of four active ones. Qwen2.5-7B-Instruct, with 9.7 million downloads, was among them. We chose not to hide this: a model's benchmarks page now says plainly that it no longer qualifies, instead of silently dropping the section. An absent score means there is not enough current evidence, not that the model is weak.&lt;/p&gt;

&lt;p&gt;After the change Kimi K3 ranked number 1 of 67.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why v2 was still not enough
&lt;/h2&gt;

&lt;p&gt;The active panel fixed the visible absurdity, but it is a patch on the same estimator, and the weakness remained. Averaging still penalizes models that were submitted to hard boards. We saw this directly later on: when we considered folding in two new Epoch boards, our check showed that doing so would move 107 standings, because a drop-and-average penalizes the frontier models that a hard, narrowly covered board was run on. We held those boards out at weight 0, which treats the symptom.&lt;/p&gt;

&lt;p&gt;It also cannot compare models whose benchmark sets do not overlap. If one model was tested on boards A and B, and another on boards C and D, an average of raw scores says nothing about which is stronger. The two sets of boards may differ in difficulty, and nothing in the average tells us by how much.&lt;/p&gt;

&lt;p&gt;That is not a new problem. It is the situation psychometrics has dealt with for decades: people sit different exams, and the exams have to be placed on one scale, with the difficulty of each test estimated from the responses rather than assumed. The tool for that is item response theory. In the second part we will describe how we applied it to benchmark results, and what it changed in the rankings.&lt;/p&gt;

&lt;h2&gt;
  
  
  The live leaderboard
&lt;/h2&gt;

&lt;p&gt;The current leaderboard, and the written methodology behind it, are here: &lt;a href="https://llmrun.dev/benchmark/llmrun-score" rel="noopener noreferrer"&gt;llmrun Score leaderboard&lt;/a&gt;. The full method is documented on the &lt;a href="https://llmrun.dev/benchmark/llmrun-score/methodology" rel="noopener noreferrer"&gt;methodology page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>machinelearning</category>
      <category>benchmark</category>
      <category>datascience</category>
    </item>
    <item>
      <title>Estimating tokens/s for Mixture-of-Experts models: active parameters, plus a routing term</title>
      <dc:creator>Alex Fank</dc:creator>
      <pubDate>Mon, 05 Oct 2026 14:33:15 +0000</pubDate>
      <link>https://dev.to/alexfank/estimating-tokenss-for-mixture-of-experts-models-active-parameters-plus-a-routing-term-1oni</link>
      <guid>https://dev.to/alexfank/estimating-tokenss-for-mixture-of-experts-models-active-parameters-plus-a-routing-term-1oni</guid>
      <description>&lt;p&gt;&lt;a href="https://llmrun.dev/how-it-works" rel="noopener noreferrer"&gt;llmrun.dev&lt;/a&gt; estimates whether a given LLM can be run locally and how fast it will decode. The first part of that question is arithmetic. The second required several iterations, and Mixture-of-Experts models were the case where the initial formula was least accurate. This note describes the estimator in its current form, the calibration behind it, and the cases in which it is known to fail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decode is a bandwidth problem
&lt;/h2&gt;

&lt;p&gt;During generation, each output token requires reading the weights from memory once. Batch size is 1 on a local machine, so there is almost no arithmetic over which to amortize that read, and speed is set by how fast bytes stream out of VRAM or unified memory. The dense formula used is the standard one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tok/s = (bandwidth GB/s / model size GB) x efficiency
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Efficiency is a per-platform constant covering software overhead and the gap between spec-sheet and achievable bandwidth. In the code it is 0.65 for NVIDIA (CUDA), 0.60 for AMD (ROCm 7.x), 0.50 for Intel (oneAPI) and 0.70 for Apple (Metal). These values were taken from community llama.cpp and Ollama results and are treated as calibration constants, not physical quantities. The basis of the estimate is dated in the code (last reviewed 2026-09-21, against llama.cpp / Ollama with GGUF weights), and the methodology page prints that date.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why dense math fails for MoE
&lt;/h2&gt;

&lt;p&gt;A Mixture-of-Experts model holds all its experts in memory but routes each token through only a few. VRAM therefore scales with total parameters, while bytes read per token scale with active parameters. gpt-oss-20b has 20.9B total parameters and 3.6B active. Applying the dense formula to its 12.9 GB footprint on an RTX 4090 (1008 GB/s, from the spec sheet) gives about 51 tok/s, whereas llama.cpp's &lt;code&gt;llama-bench&lt;/code&gt; measures about 222 on that card. The dense formula is off by more than 4x, in the pessimistic direction.&lt;/p&gt;

&lt;p&gt;The first correction is to scale the bytes read by the active fraction: the in-memory model size is multiplied by active/total parameters. This is an approximation, since attention and shared weights are not routed and the active count already includes them, but it allows a single ratio taken from the model's config to be used instead of modelling tensors separately. A model takes this path only if active parameters are below 90% of total. Anything denser is treated as dense, in which case the function returns exactly the dense result.&lt;/p&gt;

&lt;h2&gt;
  
  
  The routing term
&lt;/h2&gt;

&lt;p&gt;Scaling bytes alone overshoots. When only 3.6B parameters are read per token, the bandwidth term becomes very small, and a fixed per-layer cost that the formula ignored begins to dominate: evaluating the router and dispatching to expert kernels, once per layer, per token. This cost is modelled as an additive time per layer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;seconds/token = active_GB / (bandwidth x efficiency)
              + layers x per_layer_overhead_ms / 1000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The per-layer overhead constants are 0.055 ms for NVIDIA and 0.2 ms for Apple, AMD and Intel, with 0.1 ms when the brand is unknown. The NVIDIA value is much lower because CUDA graphs amortize launch cost better than the other llama.cpp backends do. If the layer count is missing, a fallback of 48 is used, the median across MoE models in the catalogue at the time of calibration. The constants were fitted against &lt;code&gt;llama-bench&lt;/code&gt; tg128 (decode) results for gpt-oss-20b and gpt-oss-120b on 8 hardware combinations. They are kept in the test suite as fixtures, so a change to any constant that breaks one of them fails the build.&lt;/p&gt;

&lt;p&gt;Two worked examples follow, computed from the constants above. The bandwidth values are the GPU/SoC spec values used in the fixtures. The measured figures are the &lt;code&gt;llama-bench&lt;/code&gt; results in the fixtures.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;RTX 4090, gpt-oss-20b&lt;/th&gt;
&lt;th&gt;M2 Ultra, gpt-oss-120b&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model in memory&lt;/td&gt;
&lt;td&gt;12.9 GB&lt;/td&gt;
&lt;td&gt;65 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Active / total params&lt;/td&gt;
&lt;td&gt;3.6B / 20.9B&lt;/td&gt;
&lt;td&gt;5.1B / 116.8B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bytes read per token&lt;/td&gt;
&lt;td&gt;2.22 GB&lt;/td&gt;
&lt;td&gt;2.84 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bandwidth (spec) x efficiency&lt;/td&gt;
&lt;td&gt;1008 x 0.65&lt;/td&gt;
&lt;td&gt;800 x 0.70&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bandwidth term&lt;/td&gt;
&lt;td&gt;3.39 ms&lt;/td&gt;
&lt;td&gt;5.07 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Layers x overhead&lt;/td&gt;
&lt;td&gt;24 x 0.055 = 1.32 ms&lt;/td&gt;
&lt;td&gt;36 x 0.2 = 7.20 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Estimate&lt;/td&gt;
&lt;td&gt;212 tok/s&lt;/td&gt;
&lt;td&gt;81.5 tok/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Measured (llama-bench)&lt;/td&gt;
&lt;td&gt;222&lt;/td&gt;
&lt;td&gt;80&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 4090 case is dominated by bandwidth. In the M2 Ultra case, routing overhead accounts for more than half of the per-token time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Accuracy
&lt;/h2&gt;

&lt;p&gt;The test suite asserts that every fixture lies within 18% of its measured value. The public page states an expected error of roughly plus or minus 20% on decode tok/s, and quoting a tighter figure is not justified. Eight fixtures on two model families constitute a small sample. The constants fit gpt-oss well. They have not been validated on every MoE shape (many fine-grained experts, large shared experts), and misses of more than 20% are expected for some of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it breaks
&lt;/h2&gt;

&lt;p&gt;CPU and RAM offload is the largest gap. Once part of the model resides in system RAM, decode speed is pulled toward the slowest tier from which weights are read, and a single bandwidth number is no longer meaningful. The methodology page lists offload as not modelled. Fit still indicates whether the model loads, but the tok/s figure assumes that everything is in fast memory.&lt;/p&gt;

&lt;p&gt;Prefill is a different regime. Reading the prompt is compute-bound, so it is estimated separately: prefill tok/s = dense FP16 TFLOPS x 1e12 x MFU / (2 x active parameters), with MFU of 0.45 for NVIDIA, 0.22 for AMD, 0.18 for Intel and 0.20 for Apple, calibrated on llama.cpp pp512 runs with Q4_K_M weights. Time to first token is the number of prompt tokens divided by that rate. MoE benefits prefill as well, since FLOPs scale with active parameters, but the error bar is wider, about 30%, and wider still where no verified compute figure is available. A fast-decoding card can therefore still impose a long wait before the first token.&lt;/p&gt;

&lt;p&gt;Context length is a further, less visible source of error. The decode formula accounts for the weights and nothing else, so it ignores the growing KV cache that attention also reads on every token. The cache itself is straightforward to size: 2 (K and V) x KV heads x head dim x layers x 2 bytes (FP16) per token. Some architectures keep a full cache on only some layers (linear attention, sliding window), and the VRAM code accounts for that. The speed estimate, however, does not decrease with context, so long-context numbers are optimistic. Quantized KV cache is also not modelled.&lt;/p&gt;

&lt;p&gt;Quantization carries its own overhead. Model size is derived from effective bits per weight, so Q4_K_M files are not exactly 4 bits per weight, and measured sizes from Ollama or llama.cpp are preferred over the computed estimate whenever they exist. Batching, speculative decoding, multi-GPU interconnects and flash-attention variants are likewise outside the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fit, briefly
&lt;/h2&gt;

&lt;p&gt;The fit side is simpler. VRAM is weights (parameters x bits per weight / 8) plus KV cache plus about 0.3 GB of framework overhead. When architecture data is unavailable, the fallback is a flat 10% overhead. "Fits" is judged with room reserved for a 16K-token context, because a model that loads with a 2K cache is not usable for a real conversation.&lt;/p&gt;

&lt;p&gt;The arithmetic for an individual model is printed on its page, for example &lt;a href="https://llmrun.dev/model/openai-gpt-oss-20b" rel="noopener noreferrer"&gt;gpt-oss-20b&lt;/a&gt;. Measurements from &lt;code&gt;llama-bench&lt;/code&gt; that disagree with an estimate by more than the stated error are welcome, since this is how the constants are corrected.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>localllm</category>
      <category>machinelearning</category>
      <category>performance</category>
    </item>
  </channel>
</rss>
