This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
One of my test prompts said: "Read JSON files instead of CSV."
In more than half of my reruns, three models answered by also setting the delimiter and the header-row setting to unspecified. They weren't wrong. A JSON file has no delimiter and no header row. My scoring code was wrong, because it assumed the other 19 settings had to stay exactly as they were. That mistake taught me more than any leaderboard number, and it's what this post is really about: checking whether a model changes only what you asked it to change is harder than it sounds, for the model and for the person writing the test.
I keep running into the same question in my own work with prompts and code: when I change one thing, what else breaks? An assistant can make the edit you asked for and still move something you never mentioned. So I built ButterflyBench to measure that directly.
The name comes from the butterfly effect idea: a small change can produce effects elsewhere, and I wanted to measure how much of that happens in an AI's specification.
How it works. Every scenario starts from the same 20-setting specification for a command-line data tool. It covers the Python version, input and output formats, sorting, encoding, timeout, delimiter, header row, logging level and so on, and it's described to the model in plain English. Over one to ten turns, a simulated user edits it. The user might:
- set a value ("Give each file a 60-second timeout.")
- undo the last change, or a change described indirectly ("Undo the change to how dates are printed.")
- undo "the earliest change that is still in effect"
- redo a change that was just undone
- drop a requirement ("Sorting no longer matters.")
- say something that must change nothing, like a question or a remark about a teammate's setup
After every turn the model has to restate the full specification as JSON. I also stated three coupling rules up front: TSV input forces a tab delimiter, parallel processing turns the progress bar off, and HTML output forces UTF-8. That way a required side effect isn't mistaken for collateral damage.
There is no LLM judge. A small Python oracle replays the active changes after every turn and works out the exact correct specification. The Kaggle score is the share of the 40 scenarios whose final specification matches exactly. Offline I also compute change radius (how many settings moved without being asked), recovery after an undo, and the first turn where a model went wrong.
There are 40 scenarios: 6 I wrote by hand and 34 generated from templates with a fixed random seed. They cover undo by indirect description (6), undo the earliest change still in effect (6), coupled rules (6), distractors (5), undo last (4), cancel (4), long chains (4), redo (3), and one each of cancel-then-undo and indirect reference.
I'm not claiming to have invented any of this. Multi-turn instruction following is well studied, for example in Multi-IF, MultiChallenge and EvolIF. What I wanted was something small and reproducible, with an exact answer key, where the headline question is collateral change.
My first pilot was too easy. It had 10 settings and scenarios of one to three edits, and Gemini 3.7 Flash scored 5 out of 5. I rebuilt it with 20 settings, undo, redo, indirect references, distractors and coupled rules, and that version started to separate models.
Models Tested
I ran seven models from four providers: Gemini 3.7 Flash, Grok 4.20 (both the reasoning and non-reasoning versions), Claude Sonnet 5, Claude Haiku 4.5, GPT-5.4 mini and gpt-oss-20b.
I chose them by a rule I fixed before I saw any results: a spread of providers and sizes, plus a reasoning and a non-reasoning variant of the same model where one existed, which Grok 4.20 offers. Practical limits shaped the lineup too, mainly Kaggle's daily AI quota and which models its Add Models list offered. Every model got the same 40 scenarios, prompts and settings (the library defaults, temperature 0 and seed 0). Sonnet 5 ran with an 8,000-token output cap, which is far more than the few hundred tokens a reply needs.
Findings
A note on the numbers first. The leaderboard below is the public Kaggle run, one run per model. During development I also ran the same scenarios locally, some of them against an earlier version of the scenario set. Those runs varied by as many as six scenarios for the same model, and the detailed table further down comes from a separate local run, so its totals don't match the leaderboard exactly. Please read the leaderboard as a snapshot, not a ranking.
The most interesting result wasn't the leaderboard. It was that models with similar scores failed in very different ways.
| Model | Score (40 scenarios) |
|---|---|
| Grok 4.20 Reasoning | 1.00 |
| Claude Sonnet 5 | 1.00 |
| Gemini 3.7 Flash | 1.00 |
| gpt-oss-20b | 0.90 |
| Claude Haiku 4.5 | 0.78 |
| GPT-5.4 mini | 0.75 |
| Grok 4.20 Non-Reasoning | 0.75 |
Across all my runs, gpt-oss-20b scored anywhere from 30 to 36 out of 40, Haiku 29 to 31, GPT-5.4 mini 27 to 30 and Grok non-reasoning 27 to 30. So the gaps among the three lowest models are inside the noise, and the top group is one run each.
The hard part is history, not editing
For one final local run I looked at four models in detail. Three of them, Haiku, GPT-5.4 mini and Grok non-reasoning, behaved almost identically:
| Category (scenarios) | Haiku 4.5 | GPT-5.4 mini | Grok non-reasoning | gpt-oss-20b |
|---|---|---|---|---|
| Undo the earliest change still in effect (6) | 0 | 0 | 0 | 6 |
| Undo the last change (4) | 4 | 4 | 4 | 3 |
| Undo by indirect description (6) | 6 | 6 | 6 | 5 |
| Redo (3) | 3 | 3 | 3 | 2 |
| Coupled rules (6) | 6 | 5 | 3 | 5 |
| Cancel (4) | 3 | 3 | 3 | 4 |
| Distractors (5) | 4 | 5 | 5 | 4 |
| Long chains (4) | 1 | 2 | 1 | 0 |
| Total (40) | 29 | 30 | 27 | 30 |
Those three handle "undo that last change", "undo the change to X" and redo without any trouble. What they can't do is "undo the earliest change that is still in effect", where the target has to be worked out from which changes are still active. In my first full run, GPT-5.4 mini failed all ten scenarios containing that sentence. Typically it changed nothing at all, or it undid a later change instead. Haiku managed 2 of 6 in an earlier run and 0 of 6 in the final one. Most of the long-chain failures come from the same place, since three of the four long chains contain that turn.
gpt-oss-20b is the outlier. It passed all six of those scenarios and dropped a few easy ones instead. It and GPT-5.4 mini both scored 30 out of 40 in that run, with almost opposite profiles.
Reasoning or not, same model
Grok 4.20 reasoning scored 1.00 and the non-reasoning version 0.75. In my local runs the non-reasoning version got 0 of 6 on the earliest-change scenarios every time, and the reasoning version's only earlier miss was a scenario I later fixed because the wording was ambiguous. That's one model pair and one run each, and I picked the pair after I'd seen early results. So I'm treating it as an interesting lead and I'm not drawing a general conclusion about reasoning models from it.
A pass count hides how a model fails
Haiku fails 11 of 40 scenarios but moves the fewest unrelated settings (a mean change radius of 0.075). GPT-5.4 mini fails 10 and moves the most (0.325). Grok non-reasoning sits at 0.125 and gpt-oss-20b at 0.200. A single score can't tell you whether a model's mistakes are contained or spread across the whole spec.
Many of my "model failures" were my own wording
I read the transcripts before blaming any model, and four wording problems turned up. None of them changed the oracle, the scoring or the prompts, and I logged each fix.
- "Take back whatever I said about how long each file may take." Gemini 3.7 Flash set the timeout to
unspecifiedinstead of restoring 30 seconds. My opening spec also mentioned timeouts, so "everything I said about them" is a fair reading. I reworded it to "Undo the change I made to…". - "Drop the header row." My instructions told models to use
unspecifiedfor requirements the user has dropped. The same scenario passed once and failed once on identical text. That turned out to be the pattern every time: a scenario that flips between pass and fail is usually ambiguous. - "I no longer care about a time limit per file." Claude Haiku and Grok non-reasoning answered with
timeout='none'in all four fresh runs.noneis a legal value meaning no limit, so both readings are fair. - "Read JSON files instead of CSV." This is the story from the top. In 5 of 9 reruns, Sonnet 5, Grok reasoning and gpt-oss-20b all set the delimiter and header row to
unspecified, and my oracle, which only knew about three couplings, scored sensible behaviour as collateral damage. I removed the undeclared coupling from the scenario instead of adding a fourth rule.
If you build a benchmark with an exact answer key, my advice is this: when a model seems to break something, check whether a reasonable reader of your prompt would have done the same.
Final-state scoring hides blips
One gpt-oss-20b run returned unspecified for 18 of the 20 settings in turns 1 to 3, then recovered by turn 4. The final state was correct, so a final-state score would never show it. Only the first-divergence turn did.
The top is saturated
Three models scored 1.00, so ButterflyBench separates mid-range models better than strong ones.
What I'd measure next
- Paraphrases of the earliest-change instruction. All ten scenarios that contain it use the same sentence, so I can't yet tell "can't track history" from "that wording is hard".
- Repeated runs for every model, to put error bars on the leaderboard.
- Longer chains and a version that resets the state each turn, to separate failing early from failing late.
- Code as the artifact. ButterflyBench tracks a structured spec. It doesn't test whether generated code survives an edit.
Limitations
It tests tracking of a structured spec, not collateral damage inside code. It's a small public benchmark, so models could be trained on it. And the oracle encodes only the three couplings I stated.
If you're building on Kaggle Benchmarks
Three things cost me time. Nested evaluations force max_attempts=1, so retry inside your own task, otherwise one provider hiccup shows up as "Error" for a whole model. An expensive model failed with a 403 ("estimated cost exceeds your available quota") until I capped its output tokens through extra_api_params. And the main task's docstring becomes its description, which must be 255 characters or fewer.
My Benchmark
- Benchmark on Kaggle: https://www.kaggle.com/benchmarks/emmimalalexander/butterflybench
- Task and leaderboard: https://www.kaggle.com/benchmarks/tasks/emmimalalexander/butterflybench
- Notebook (oracle, scoring and all 40 scenarios): https://www.kaggle.com/code/emmimalalexander/butterflybench
Top comments (0)