DEV Community

sichi chen
sichi chen

Posted on

Is the model actually getting dumber, or are we just reading tea leaves from single samples?

Disclosure: I'm involved in building Folkbench, which I mention near the end of this post.

I've been chewing on a question that's been discussed to death but never really settled: when we say a model "got dumber" or "got nerfed," how much of that is real, and how much is just the illusion of a single sample?

Lately a lot of people I know have been throwing the pelican test at models (have it draw an SVG of a pelican riding a bicycle). It's a genuinely brutal test — it's not just about writing code, it's about spatial reasoning: beak and handlebars, feet and pedals, frame and wheels. If the model's spatial understanding is even slightly off, you get something pretty abstract.

But after running it a few dozen times, I think the biggest trap with the pelican test is judging by a single output.

An LLM is a probabilistic sampler with randomness baked in. Same prompt: on one run the spatial awareness is maxed out — frame, cranks, foot placement all correct; run it again later and suddenly it's postmodern abstract art. If provider A gets a basically sensible pose in 16 out of 20 runs and provider B only passes 8 out of 20, then the difference means something statistically. Declaring "A is the full model, B is watered down" based on one random screenshot is basically flipping a coin.

To actually compare anything, you need a shared baseline — fix the model version, reasoning effort, prompt and time window, set a reference point first, and then look only at relative differences.

That leads to a few really painful engineering details:

  1. How do you score the pelican? Pure human review doesn't scale; pure LLM-as-a-judge tends to heavily favor its own output. For now it looks like you have to split it into four dimensions — Pelican (completeness), Bicycle (geometric soundness), Riding (how the bird actually connects to the bike spatially) and Animation (motion and clipping) — and cross-check with a vision judge plus static rules on the SVG DOM.
  2. The lifecycle of test prompts. Once any "killer prompt" spreads, sooner or later it gets scraped into training data or targeted with fine-tuning. Long term you can't rely on a single question; you need a dynamic pool of probes covering spatial relations, instruction following and structural reasoning.

I'd been running these tests by hand for a while, and the biggest pain was the cost of record-keeping: repeated runs, saving SVGs, logging parameters, aligning timestamps — after a few dozen batches you're numb. What eats the time usually isn't the test itself but all the tedious logging and organizing around it. That said, if you just want a quick read on how different models perform, you don't have to do it all yourself — there's a ready-made leaderboard on Folkbench (https://folkbench.com/?utm_source=luntan&utm_campaign=dev) that can save you the effort.

If you're also playing with this test, I'd love to hear how you run it — or see your most cursed results.

Top comments (1)

Collapse
 
sichi_chen_a4a87b20aa2dbf profile image
sichi chen • • Edited

Following up on my own post: I wrote two longer pieces that dig deeper — one on how to actually tell whether a model "got dumber," and one on how to score the cycling pelican across those four dimensions.

Happy to be challenged on the methodology, especially the scoring rubric — curious what's worked for others.