<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: sichi chen</title>
    <description>The latest articles on DEV Community by sichi chen (@sichi_chen_a4a87b20aa2dbf).</description>
    <link>https://dev.to/sichi_chen_a4a87b20aa2dbf</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4158909%2F59d63511-3618-4153-83bb-9c4410484e98.png</url>
      <title>DEV Community: sichi chen</title>
      <link>https://dev.to/sichi_chen_a4a87b20aa2dbf</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sichi_chen_a4a87b20aa2dbf"/>
    <language>en</language>
    <item>
      <title>Is the model actually getting dumber, or are we just reading tea leaves from single samples?</title>
      <dc:creator>sichi chen</dc:creator>
      <pubDate>Mon, 05 Oct 2026 13:10:12 +0000</pubDate>
      <link>https://dev.to/sichi_chen_a4a87b20aa2dbf/is-the-model-actually-getting-dumber-or-are-we-just-reading-tea-leaves-from-single-samples-45eg</link>
      <guid>https://dev.to/sichi_chen_a4a87b20aa2dbf/is-the-model-actually-getting-dumber-or-are-we-just-reading-tea-leaves-from-single-samples-45eg</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Disclosure:&lt;/strong&gt; I'm involved in building Folkbench, which I mention near the end of this post.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I've been chewing on a question that's been discussed to death but never really settled: when we say a model "got dumber" or "got nerfed," how much of that is real, and how much is just the illusion of a single sample?&lt;/p&gt;

&lt;p&gt;Lately a lot of people I know have been throwing the pelican test at models (have it draw an SVG of a pelican riding a bicycle). It's a genuinely brutal test — it's not just about writing code, it's about spatial reasoning: beak and handlebars, feet and pedals, frame and wheels. If the model's spatial understanding is even slightly off, you get something pretty abstract.&lt;/p&gt;

&lt;p&gt;But after running it a few dozen times, I think the biggest trap with the pelican test is judging by a single output.&lt;/p&gt;

&lt;p&gt;An LLM is a probabilistic sampler with randomness baked in. Same prompt: on one run the spatial awareness is maxed out — frame, cranks, foot placement all correct; run it again later and suddenly it's postmodern abstract art. If provider A gets a basically sensible pose in 16 out of 20 runs and provider B only passes 8 out of 20, &lt;em&gt;then&lt;/em&gt; the difference means something statistically. Declaring "A is the full model, B is watered down" based on one random screenshot is basically flipping a coin.&lt;/p&gt;

&lt;p&gt;To actually compare anything, you need a shared baseline — fix the model version, reasoning effort, prompt and time window, set a reference point first, and then look only at relative differences.&lt;/p&gt;

&lt;p&gt;That leads to a few really painful engineering details:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;How do you score the pelican?&lt;/strong&gt; Pure human review doesn't scale; pure LLM-as-a-judge tends to heavily favor its own output. For now it looks like you have to split it into four dimensions — Pelican (completeness), Bicycle (geometric soundness), Riding (how the bird actually connects to the bike spatially) and Animation (motion and clipping) — and cross-check with a vision judge plus static rules on the SVG DOM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The lifecycle of test prompts.&lt;/strong&gt; Once any "killer prompt" spreads, sooner or later it gets scraped into training data or targeted with fine-tuning. Long term you can't rely on a single question; you need a dynamic pool of probes covering spatial relations, instruction following and structural reasoning.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I'd been running these tests by hand for a while, and the biggest pain was the cost of record-keeping: repeated runs, saving SVGs, logging parameters, aligning timestamps — after a few dozen batches you're numb. What eats the time usually isn't the test itself but all the tedious logging and organizing around it. That said, if you just want a quick read on how different models perform, you don't have to do it all yourself — there's a ready-made leaderboard on &lt;strong&gt;Folkbench&lt;/strong&gt; (&lt;a href="https://folkbench.com/?utm_source=luntan&amp;amp;utm_campaign=dev" rel="noopener noreferrer"&gt;https://folkbench.com/?utm_source=luntan&amp;amp;utm_campaign=dev&lt;/a&gt;) that can save you the effort.&lt;/p&gt;

&lt;p&gt;If you're also playing with this test, I'd love to hear how you run it — or see your most cursed results.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
