This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
My Kaggle Benchmarks task measures how well a model translates legacy COBOL, Fortran, and PHP programs into Python or Rust while preserving their observable behavior. Each task score is the number of supplied test cases passed over the total number of cases. A case passes only when the generated program exits successfully and writes the exact expected standard output.
The dataset contains 20 programs. The public split and the full split both has 20 cases per program.
Example data
{
"task_id": "fortran_real_mean_rust",
"family_id": "fortran_real_semantics",
"source_language": "fortran",
"target_language": "rust",
"difficulty": 4,
"legacy_features": ["single-precision REAL accumulation", "F14.3 fixed-point output", "loss of integer precision above 2**24"],
"prompt": "Translate this Fortran program into a complete Rust 2021 program that reads standard input and writes standard output. Reproduce the original program's observable behavior exactly for every input of the stated form, including edge cases. Assume the original runs under: gfortran -std=legacy -O0 -fwrapv (default INTEGER and REAL are 32-bit). Input: a count n (1 <= n <= 10) on the first line, then n real numbers, one per line. Output: exactly what the original prints. Return only the program source code.",
"source_code": " PROGRAM AVG\n INTEGER N, I\n REAL X, S\n READ(*,*) N\n S = 0.0\n DO 10 I = 1, N\n READ(*,*) X\n S = S + X\n 10 CONTINUE\n WRITE(*,'(F14.3)') S / N\n END", "reference_runtime": "gfortran -std=legacy -O0 -fwrapv (default INTEGER and REAL are 32-bit)",
"provenance": "original program written for this benchmark; golden outputs generated by running the original under reference_runtime",
"version": "0.2.0",
"tests": [{"stdin": "3\n16777216\n1\n1\n", "stdout": " 5592405.500\n"}, {"stdin": "3\n1\n2\n3\n", "stdout": " 2.000\n"}]
}
Models Tested
I tested my benchmark on the following models:
- Gemini 3.1 Pro Preview
- Grok 4.20 Reasoning
- Gemini 3.7 Flash
- Claude Haiku 4.5
- GPT-5.4 mini
- Claude Sonnet 4.5
- Claude Opus 4.6
- Grok 4.5
- GPT-6 Astra
Now for some reason, Grok 4.5 and GPT-6 Astra did not run and I got this error:
openai.NotFoundError: Error code: 404 - {'code': '', 'message': 'The requested model "xai/grok-4.5-0708" was not found.', 'param': '', 'type': 'not_found_error'}
It says the requested model was not found. But when I extracted the list of supported models,
import kaggle_benchmarks as kbench
print(list(kbench.llms.keys()))
The output had these models:
[... , 'openai/gpt-6-astra', 'xai/grok-4.5-0708']
Findings
In my benchmark runs, scores clustered strictly into discrete buckets—models scored 7/10, 3/10, or 0/10, with none achieving a perfect score.
The Bimodal Distribution: Because the evaluation suite executes compiled binaries and Python scripts against exact stdin/stdout assertions, models are heavily penalized for formatting drift.
The Zero-Score Trap: Models scoring 0 didn't necessarily write completely wrong code—many failed because rustc compilation failed due to strict type safety, or because the LLM included conversational intro text that broke execution.
Legacy Translation Gap: Translating legacy paradigms (like COBOL fixed-width output) into Rust or Python requires precise state management that even top-tier models struggled to execute consistently across all test cases.
And also, observing a 404 Not Found error when querying specific provider endpoints (e.g., xai/grok-4.5-0708) indicates deprecated/renamed endpoints.
My Benchmark
Check the benchmark here.

Top comments (0)