This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I maintain cli-bench, an open benchmark for coding agents that work in a terminal. Each task is a small repository with a bug to fix, a feature to build, or a function to speed up, plus a verifier script that decides pass or fail. Nothing is graded by another model. Either the tests pass and the extra checks hold, or the run fails.
For this challenge I ported six cli-bench tasks to Kaggle Benchmarks:
| Kaggle task | What the model has to do | What the verifier checks |
|---|---|---|
cb-debug-wrong-answer |
Find the bug in a small stats library | The shipped pytest suite passes |
cb-refactor-deadcode |
Delete the three unused functions and nothing else | Dead functions gone, six live ones still defined, tests pass |
cb-feature-rate-limiter |
Write a thread-safe token bucket from an interface spec | Tests for burst, refill, atomic rollback, and 100 threads at once |
cb-perf-hot-loop |
Make a pair counter at least 10x faster in pure Python | Stdlib only, same signature, randomized equivalence, 10x on uniform, clustered, and gridded points |
cb-data-log-analysis |
Answer seven exact questions about an application log it never sees | The model writes solve.py; the task runs it and compares every answer to the log |
cb-sec-patch-xss |
Close an XSS hole in a comment board | Exploit tests pass and a fresh payload renders as text |
In cli-bench, an agent gets a shell and a time budget. On Kaggle the model gets one shot: the prompt holds every file in the repo, and the model answers with whole files in FILE: path blocks. The task writes those files into a temporary directory and runs the original cli-bench verifier gates on the result. A run passes only if every gate passes. If the reply rewrites a test file or an input, that write is thrown away and the run fails, the same way cli-bench treats sabotage.
So the question this benchmark asks is: how much of a coding agent's score comes from the model reading code carefully, and how much comes from the agent loop of running tests and trying again? I already had one agent run on the full suite (Codex with gpt-5.6-luna, 23 of 36 trials passed). The Kaggle version takes the loop away.
Before running a single model, I found two bugs in my own benchmark
Porting meant reading every verifier line by line, and checking each task with a reference solution and a few wrong ones. Two tasks did not hold up.
sec/patch-xss could not be passed by a correct fix. Two shipped tests assert that the words alert and onerror appear nowhere in the rendered page. The verifier's own probe then requires the escaped payload text, including alert, to still be in the page. Escaping the comment with html.escape is the textbook fix, and it fails the tests, because <script>alert(...) still contains the word alert. Deleting the text passes the tests and fails the probe. In the Codex run, two of the three trials used html.escape and failed the tests, and the third stripped the text and failed the probe. I had written those failures up as the model's fault.
data/log-analysis described one rule and graded another. The question sheet defined error_rate as the "fraction of lines with level ERROR". The verifier divides ERROR request lines by request lines, and about one log line in seven is not a request. Both of the Codex trials that failed this task were off on error_rate and nothing else (0.0434 vs 0.05, 0.038 vs 0.0441). Those are exactly the numbers you get by following the text.
That is 5 of the 13 failed trials in my first leaderboard run that came from the benchmark, not the model. The Kaggle versions fix both: the XSS tests now check for raw markup instead of words, and the question sheet states the rule the verifier checks.
Models Tested
| Model (Kaggle slug) | Why it is in the lineup |
|---|---|
gpt-5.6-luna |
The same model my Codex agent run used, so one shot and agent loop can be compared directly. |
claude-opus-5-5-default |
Anthropic's top tier on Kaggle's model list. |
gemini-3.1-pro-preview |
Google's Pro tier on Kaggle's model list. |
qwen3-coder-480b-a35b-instruct |
An open-weights model built specifically for code. |
gpt-oss-120b |
An open-weights general model you can run on your own hardware. |
gemini-3.7-flash |
Not picked: Kaggle runs its default model when a task is pushed, so it came along for free. |
Each model ran each task once, through Kaggle's model proxy at the SDK's default temperature.
Findings
| task | Opus 5.5 | Gemini 3.1 Pro | GPT-5.6 Luna | Qwen3 Coder 480B | gpt-oss-120b | Gemini 3.7 Flash |
|---|---|---|---|---|---|---|
| cb-debug-wrong-answer | pass | pass | pass | pass | pass | pass |
| cb-refactor-deadcode | pass | pass | pass | pass | pass | pass |
| cb-feature-rate-limiter | pass | pass | pass | pass | pass | pass |
| cb-perf-hot-loop | pass | pass | pass | fail | pass | pass |
| cb-data-log-analysis | pass | pass | pass | pass | pass | pass |
| cb-sec-patch-xss | pass | pass | pass | pass | pass | pass |
| total | 6/6 | 6/6 | 6/6 | 5/6 | 6/6 | 6/6 |
1. One shot was enough for these six tasks. 35 of 36 runs passed. With the files in front of them and no way to run anything, every model fixed the stats bug, pruned exactly the three dead functions, wrote a token bucket that rolls back atomically and survives 100 threads, and escaped the comment board correctly. These tasks are too easy to separate frontier models in one shot. That is a result too: the difficulty I measured in cli-bench was not coming from these tasks.
2. The same model did better without the agent loop. GPT-5.6 Luna passed all six in one shot. As a Codex agent it passed hot-loop 1 time in 3, with a worst-case speedup of 1.2x on the gridded dataset. In one shot it wrote a spatial grid that, timed on my machine against the same verifier, ran 42.5x faster on uniform points, 26.5x on gridded, and 14.6x on clustered. The caveats are real: one run against three, one dataset seed against three, and different machines. But it is the opposite of what I expected. Having a shell did not help the agent find the fast solution, and may have pulled it toward measuring and patching a slow one.
3. Every model reached for the same idea on hot-loop, and the only failure was invisible to the unit tests. All six bucketed points into a grid. Clustered points were the worst case for every passing model (14.6x to 41x on my machine), not gridded. Qwen3 Coder's grid passed all nine shipped unit tests and still undercounted: on one random case it found 84 pairs where there are 123. Only the randomized comparison against the naive version caught it. If my verifier had stopped at the shipped tests, that bug would have scored as a pass.
4. Most of my first-round failures were my harness, again. The first batch of Kaggle runs had 9 failures. Seven were mine:
- Kaggle's task runtime does not have pytest installed. My fallback test runner did not support pytest fixtures (
tmp_path,monkeypatch), so all six models "failed" the XSS task. Every one of them had escaped the output correctly. - My reply parser expected
FILE: solve.pyand rejected gpt-oss-120b's**FILE: solve.py**. I ran its script by hand against the same log and every answer was correct.
I fixed both, pushed new versions of those two tasks, and re-ran every model on them. The table shows the re-runs. The other two first-round failures were real. One is the hot-loop bug above. In the other, Qwen3 Coder's log script called statistics.quantiles with a method name that does not exist and crashed. On the re-run it wrote a different script and passed. That is one run at the default temperature, so treat any single pass or fail as a sample, not a verdict.
5. The corrected XSS task is passable by the textbook fix. All six models used HTML escaping, and all six passed the corrected tests and the probe. In cli-bench 0.9.1, the same approach failed. That confirms the problem was the tests, not the models.
What surprised me: across both versions of this benchmark, bugs on my side (two task specs, a test runner, a parser) caused more failures than the models did.
What I would measure next:
- Harder cli-bench tasks: the flaky-test and CI tasks, and the four "houdini" probes that check whether a model games the verifier.
- Three or more runs per model, so a single crash like Qwen's is visible as variance.
- A version that gives the model a tool to run the tests, to see whether one retry helps or, as with hot-loop, hurts.
The lesson I am keeping: a benchmark that has never been run against a known-correct answer is not a benchmark yet. Every task in this port now ships with a reference solution, an empty reply, and a test-rewriting reply. A task only counts when the first passes and the other two fail, and that has to hold in the environment where the models actually run, not just on my laptop.
My Benchmark
Kaggle benchmark: https://www.kaggle.com/benchmarks/aks1321/cli-bench-one-shot
The six public tasks (each page has the full task code in its published notebook):
- https://www.kaggle.com/benchmarks/tasks/aks1321/cb-debug-wrong-answer
- https://www.kaggle.com/benchmarks/tasks/aks1321/cb-refactor-deadcode
- https://www.kaggle.com/benchmarks/tasks/aks1321/cb-feature-rate-limiter
- https://www.kaggle.com/benchmarks/tasks/aks1321/cb-perf-hot-loop
- https://www.kaggle.com/benchmarks/tasks/aks1321/cb-data-log-analysis
- https://www.kaggle.com/benchmarks/tasks/aks1321/cb-sec-patch-xss
cli-bench is Apache-2.0: https://github.com/arjunkshah12345-hash/cli-bench
Top comments (1)
tr.ee/dev-to