This is a submission for the Kaggle Benchmarking Challenge
Static benchmarks have a shelf life. Once a question is public, it can end up in a training set, and the score starts measuring memory. I wanted tasks where memory cannot help. So I generated them with code, from a fixed seed, and checked every answer with a program. No model grades another model.
The result is VORTEX (Valuation of Reasoning, Tool-use, and EXpertise), a five-task benchmark on Kaggle. I ran it on three models. The scores were interesting. The behavior behind the scores was more interesting.
What I Benchmarked
Each task targets one capability and has one twist that punishes shortcuts.
| Task | What it tests | The twist |
|---|---|---|
| Alien Interpreter | Writing an interpreter for a brand-new stack language from its spec. Scored on 45 hidden programs. | Three dialects share one syntax but differ in small rules: division rounding, integer width, empty-stack behavior, operand order. An interpreter written from memory of "a standard stack VM" fails. |
| Black-Box Oracle | Finding a hidden function with a budget of 25 queries, then submitting code. Scored on 20 hidden inputs per function. | The hidden inputs include edge cases the model never sees (empty input, odd lengths, wrap-around). It has to probe for them. |
| Fog-of-War Dungeon | Escaping four hidden mazes using only tool calls (move, look, pick_up, use_key). |
The exit needs a red key. The red key sits behind a blue door. The model must build a map in its head and plan routes. |
| Temporal Ledger | Replaying a shuffled event log across three time zones and reporting exact balances. | Daylight-saving changes, moving exchange rates, fees, rejected overdrafts, retroactive voids, and non-retroactive reversals. Entries are listed out of time order. |
| Spine Test | Holding a correct answer under confident false pushback, and changing a wrong answer after a correct correction. | Always agreeing scores about 50%. Never changing also scores about 50%. Only checking the work scores near 100%. |
All scoring is deterministic. If a model returns unreadable output or code that times out, that instance scores zero. Instances come from a fixed seed, so every model sees identical problems.
Models Tested
- Gemini 3.7 Flash: the default model, so it ran first. A fast, cheap model.
- Claude Sonnet 5.5: a strong general model from a different family.
- GPT-5.4: a third provider, so the comparison is not one lab against itself.
Each model did one run per task at default settings. One run is a small sample, so treat what follows as observations, not a ranking.
Findings
| Task | Gemini 3.7 Flash | Claude Sonnet 5.5 | GPT-5.4 |
|---|---|---|---|
| Alien Interpreter | 1.00 | 1.00 | 1.00 |
| Spine Test | 1.00 | 1.00 | 1.00 |
| Fog-of-War Dungeon | 1.00 | 1.00 | 0.88 |
| Black-Box Oracle | 0.92 | 0.93 | 0.43 |
| Temporal Ledger | 1.00 | 1.00 | 0.06 |
1. The same perfect score hid a 20x difference in effort
On Alien Interpreter, all three models wrote interpreters that passed all 45 hidden programs. The cost of getting there was very different:
| Output tokens | Wall time | Cost | |
|---|---|---|---|
| GPT-5.4 | ~2,400 | 30 s | $0.04 |
| Claude Sonnet 5.5 | ~5,700 | 38 s | $0.06 |
| Gemini 3.7 Flash | ~52,000 | 384 s | $0.20 |
Gemini spent about twenty times the tokens of GPT-5.4 to reach the same answer. A leaderboard that shows only the score cannot see this. Cost and latency belong next to accuracy.
2. Effort before answering predicted the hard-task results
On Temporal Ledger, the three models wrote very different amounts of text per ledger:
- Gemini 3.7 Flash: 18,000 to 27,000 tokens
- Claude Sonnet 5.5: 7,000 to 13,000 tokens
- GPT-5.4: 100 to 1,300 tokens
The scores follow the same order. Gemini and Claude replayed every ledger exactly, to the cent. GPT-5.4 got one account out of eighteen right. On ledger 1 it showed no working at all and went straight to final balances. That is a guess, not a computation, and the task is built so that guessing fails: one daylight-saving shift or one reversed transfer changes the answer.
A caveat: I ran GPT-5.4 at its default settings, and I could not see its reasoning setting in the run files. The data supports "at default settings it did not work through the log". It does not support "GPT-5.4 cannot do this".
3. In Black-Box Oracle, the model that experimented less did worse
Claude and Gemini scored about 0.92. GPT-5.4 scored 0.43. Its worst result was on one of the simplest functions: a string rule that shifts the letter at position i by 4 × i. It scored 0.15 there. It used 3 to 6 model turns per function. Gemini used up to 26.
GPT-5.4 did solve the two easy list functions, which a one-line guess can get right. It missed most of the rest. This looks like pattern-matching on a few outputs, not hypothesis testing.
The strong models were not perfect. Both missed the same function: three chained stages (multiply-and-add modulo 20, then adjacent differences, then running sums modulo 14). It was the hardest function for both. Gemini wrote about 66,000 tokens on it and Claude about 33,000, and neither solved it. More effort helped, but this function is still beyond them.
4. In the Dungeon, the habit was visible in the tool calls
All three could escape. GPT-5.4 scored 0.88 because it hit the 120-round tool limit on the largest maze. The cause was its walking style:
move calls |
Share that were single steps | Input tokens | |
|---|---|---|---|
| Claude Sonnet 5.5 | 107 | 21% | ~0.37 M |
| Gemini 3.7 Flash | 121 | 51% | ~0.48 M |
| GPT-5.4 | 284 | 81% | ~1.05 M |
The move tool accepts a long route like NNEESSW and stops at the first obstacle. A model that remembers the map can send a long route in one call. GPT-5.4 mostly took one step and looked around again. It used about 2 to 3 times the input tokens of the others and cost roughly twice as much. The task rewards planning and memory, and the call pattern shows who used them.
5. Two tasks are too easy now
All three models scored 1.00 on Spine Test. None of them gave in to a confident, authority-backed wrong argument, and all of them updated when given a correct derivation. That is good news about these models, but it means this task no longer separates them. Alien Interpreter has the same problem. I would make both harder: subtler false arguments, longer chains of arithmetic, and more instances.
What surprised me
The scores were not the main story. The same scoreboard hid very different working habits: how long each model thought, whether it tested its hypothesis, and whether it planned ahead. In this set, the habit explained more of the result than the model name did.
What I'd Measure Next
- Reasoning effort as a variable. Re-run GPT-5.4 with reasoning turned up, to see whether the Ledger and Oracle gaps close.
- Harder Spine and Alien tasks, so they separate models again.
- Repeated runs. Run each model several times to measure run-to-run variance before claiming any ordering.
- Efficiency as a score. Report accuracy together with tokens, time, and cost.
My Benchmark
Link to my Kaggle benchmark: https://www.kaggle.com/benchmarks/farzado/vortex/versions/1
Top comments (0)