This is a submission for the Kaggle Benchmarking Challenge.
What I Benchmarked
FrameFlip48 measures whether a model tracks the translation frame in a short robot-motion program. Each of 24 matched pairs has identical numbers and commands. One prompt says ACTIVE_FRAME = LOCAL; its counterpart says ACTIVE_FRAME = WORLD. Cases run in separate conversations, so the model never receives the counterpart's answer.
The robot lives on an integer grid. Its turns are multiples of 90 degrees. In WORLD mode, a translation uses the fixed east/north axes. In LOCAL mode, those axes rotate with its current heading. Turns rotate around the robot's own center; they do not move the center.
For example, start at (2, -1) facing east, turn 90 degrees counterclockwise, then MOVE(3, 2):
| Translation frame | Final world position | Final heading |
|---|---|---|
| WORLD | (5, 1) |
90° |
| LOCAL | (0, 2) |
90° |
Coordinates can look plausible while being expressed in the wrong frame. Exact arithmetic gives this test a ground truth that needs no LLM judge or subjective rating.
The corpus has 48 cases, balanced across four initial headings and programs of 2, 4 and 8 commands. Each pair has distinct final LOCAL/WORLD positions. Fixed-seed generation used only oracle properties; no model outputs selected the cases. The corpus was frozen before evaluation, with SHA-256 590c916ee01d408b580f40a8f1d00262e79360cec46c18194ba8c6eb883d8e87.
Models Tested
The planned lineup used three labs' small models, selected from Kaggle's available catalog before seeing outputs. All ran on the same fixed task version 1 on October 2, 2026:
- Gemini 3.1 Flash-Lite Preview:
google/gemini-3.1-flash-lite-preview, run3955369. - GPT-5.4 nano:
openai/gpt-5.4-nano-2026-03-17, run3955370. - Claude Haiku 4.5:
anthropic/claude-haiku-4-5@20251001, run3955371.
Kaggle task creation also automatically ran Gemini 3.7 Flash (google/gemini-3.7-flash, run 3955292). That additional result is included rather than discarded.
Every case gets one response in a fresh conversation, temperature zero, with no supplied tools. Reasoning uses each model's platform default, so this is not an equal-compute comparison. The prompt requests only one JSON object with the final position and normalized heading. The frozen parser accepts plain or fenced JSON and integer-valued numbers, but rejects surrounding explanation. Formatting failures are reported separately. Transport/quota failures abort a run instead of shrinking the denominator; all four runs completed.
There was also one earlier real SDK validation with Flash-Lite: 10/48 cases and 1/24 pairs. It is separate from the server comparison below, with no prompt/scorer changes or pooling of the two runs. Mock SDK checks and deterministic baselines are not model evaluations. The full workflow used $0.2543933 of Kaggle's free inference allowance; actual paid spend was $0.
Findings
The primary score is the fraction of 24 pairs for which both variants have the exact final pose and satisfy the output contract. These are the public task results, independently re-scored from the downloaded responses:
| Model | Correct pairs / 24 | Strict correct cases / 48 | Format failures |
|---|---|---|---|
| Gemini 3.7 Flash | 24 | 48 | 0 |
| Claude Haiku 4.5 | 0 | 0 | 48 |
| Gemini 3.1 Flash-Lite Preview | 0 | 8 | 1 |
| GPT-5.4 nano | 0 | 1 | 0 |
Pose correctness and output compliance need separate diagnostics. Claude's zero is dominated by its format: all 48 responses explain the calculation before presenting a final fenced JSON object. For example, its answer to n8-h90-r0-world explains eight commands and ends with {"x":9,"y":-13,"heading_deg":270}. That pose is correct, but the complete response violates the requested output contract.
After observing this, an independent reviewer applied the same post-hoc extraction rule to all 192 saved responses: require exactly one valid pose JSON object in the entire response and require it to be terminal, optionally fenced; preceding explanation is ignored. This was analysis of existing outputs, not a model rerun, and it does not replace the frozen leaderboard score:
| Model | Post-hoc correct poses / 48 | Post-hoc correct pairs / 24 |
|---|---|---|
| Gemini 3.7 Flash | 48 | 24 |
| Claude Haiku 4.5 | 44 | 20 |
| Gemini 3.1 Flash-Lite Preview | 9 | 0 |
| GPT-5.4 nano | 1 | 0 |
Claude's four remaining pose errors are all LOCAL position errors; every heading and all 24 WORLD positions are correct. Flash-Lite's strict successes are all WORLD cases, and all are two-command programs. Thus the weak strict scores have different causes: output compliance for Claude, and substantial pose errors for the two smaller models in these runs. Gemini 3.7 Flash reaches the ceiling of this corpus; a harder extension would be needed to measure its margin.
Pair accuracy protects against a fixed-frame shortcut. A deterministic program that always uses WORLD scores 24/48 individual cases but solves 0/24 pairs. Always-LOCAL has the same pattern; the exact oracle solves 48/48 cases and 24/24 pairs. These are local code baselines, not LLM results.
The outputs do not establish that the weaker models used that shortcut. GPT-5.4 nano has just one exact opposite-frame match: for n2-h180-r0-local, it returns (0,6,90) instead of (0,4,90), matching the WORLD counterfactual. Other errors include incorrect headings and unrelated positions. An opposite-frame match is an observable signature, not proof of internal reasoning. WORLD is computationally simpler than LOCAL; the paired design controls numbers and commands, not equal reasoning difficulty.
This is a small synthetic test of 2D quarter-turn motion, not a robotics reliability certificate. Pair members are mathematically dependent, and one response per case cannot estimate run-to-run variance. Different platform defaults and the Flash-Lite validation/server discrepancy also limit broad model rankings. A future version could add mixed per-command frames, 3D rotations and a separately reported tool-enabled condition, frozen before inspecting its outputs.
My Benchmark
FrameFlip48 public Kaggle benchmark collection
The collection contains the owned task version 1 and all four completed models. It aggregates the numeric task with Average of task scores. Reproduce or inspect it using the source notebook, run outputs, and original cases, prompts and gold answers.
Original code is MIT; original synthetic data is CC0-1.0. The public collection is also under Kaggle's Apache 2.0 publication terms. No external dataset, private information or participant responses were used. The task uses the Kaggle Benchmarks SDK. Seven local tests passed; an independent integer-matrix oracle checked all 48 gold and 48 opposite-frame poses. A separate review reproduced all strict metrics and verified that all 192 scored prompts/responses match the native SDK traces.
Autonomous workflow disclosure: Codex agents designed, implemented, evaluated, reviewed and wrote this entry on my behalf. I completed the account's required personal verification. The benchmark, outputs and interpretation are public so readers can assess the work directly.
Top comments (0)