DEV Community

Arjun Shah
Arjun Shah

Posted on

I ported my coding-agent benchmark to Kaggle, and the first bugs I found were mine

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

I maintain cli-bench, an open benchmark for coding agents that work in a terminal. Each task is a small repository with a bug to fix, a feature to build, or a function to speed up, plus a verifier script that decides pass or fail. Nothing is graded by another model. Either the tests pass and the extra checks hold, or the run fails.

For this challenge I ported six cli-bench tasks to Kaggle Benchmarks:

Kaggle task What the model has to do What the verifier checks
cb-debug-wrong-answer Find the bug in a small stats library The shipped pytest suite passes
cb-refactor-deadcode Delete the three unused functions and nothing else Dead functions gone, six live ones still defined, tests pass
cb-feature-rate-limiter Write a thread-safe token bucket from an interface spec Tests for burst, refill, atomic rollback, and 100 threads at once
cb-perf-hot-loop Make a pair counter at least 10x faster in pure Python Stdlib only, same signature, randomized equivalence, 10x on uniform, clustered, and gridded points
cb-data-log-analysis Answer seven exact questions about an application log it never sees The model writes solve.py; the task runs it and compares every answer to the log
cb-sec-patch-xss Close an XSS hole in a comment board Exploit tests pass and a fresh payload renders as text

In cli-bench, an agent gets a shell and a time budget. On Kaggle the model gets one shot: the prompt holds every file in the repo, and the model answers with whole files in FILE: path blocks. The task writes those files into a temporary directory and runs the original cli-bench verifier gates on the result. A run passes only if every gate passes. If the reply rewrites a test file or an input, that write is thrown away and the run fails, the same way cli-bench treats sabotage.

So the question this benchmark asks is: how much of a coding agent's score comes from the model reading code carefully, and how much comes from the agent loop of running tests and trying again? I already had one agent run on the full suite (Codex with gpt-5.6-luna, 23 of 36 trials passed). The Kaggle version takes the loop away.

Before running a single model, I found two bugs in my own benchmark

Porting meant reading every verifier line by line, and checking each task with a reference solution and a few wrong ones. Two tasks did not hold up.

sec/patch-xss could not be passed by a correct fix. Two shipped tests assert that the words alert and onerror appear nowhere in the rendered page. The verifier's own probe then requires the escaped payload text, including alert, to still be in the page. Escaping the comment with html.escape is the textbook fix, and it fails the tests, because <script>alert(...) still contains the word alert. Deleting the text passes the tests and fails the probe. In the Codex run, two of the three trials used html.escape and failed the tests, and the third stripped the text and failed the probe. I had written those failures up as the model's fault.

data/log-analysis described one rule and graded another. The question sheet defined error_rate as the "fraction of lines with level ERROR". The verifier divides ERROR request lines by request lines, and about one log line in seven is not a request. Both of the Codex trials that failed this task were off on error_rate and nothing else (0.0434 vs 0.05, 0.038 vs 0.0441). Those are exactly the numbers you get by following the text.

That is 5 of the 13 failed trials in my first leaderboard run that came from the benchmark, not the model. The Kaggle versions fix both: the XSS tests now check for raw markup instead of words, and the question sheet states the rule the verifier checks.

Models Tested

Model (Kaggle slug) Why it is in the lineup
gpt-5.6-luna The same model my Codex agent run used, so one shot and agent loop can be compared directly.
claude-opus-5-5-default Anthropic's top tier on Kaggle's model list.
gemini-3.1-pro-preview Google's Pro tier on Kaggle's model list.
qwen3-coder-480b-a35b-instruct An open-weights model built specifically for code.
gpt-oss-120b An open-weights general model you can run on your own hardware.
gemini-3.7-flash Not picked: Kaggle runs its default model when a task is pushed, so it came along for free.

Each model ran each task once, through Kaggle's model proxy at the SDK's default temperature.

Findings

task Opus 5.5 Gemini 3.1 Pro GPT-5.6 Luna Qwen3 Coder 480B gpt-oss-120b Gemini 3.7 Flash
cb-debug-wrong-answer pass pass pass pass pass pass
cb-refactor-deadcode pass pass pass pass pass pass
cb-feature-rate-limiter pass pass pass pass pass pass
cb-perf-hot-loop pass pass pass fail pass pass
cb-data-log-analysis pass pass pass pass pass pass
cb-sec-patch-xss pass pass pass pass pass pass
total 6/6 6/6 6/6 5/6 6/6 6/6

1. One shot was enough for these six tasks. 35 of 36 runs passed. With the files in front of them and no way to run anything, every model fixed the stats bug, pruned exactly the three dead functions, wrote a token bucket that rolls back atomically and survives 100 threads, and escaped the comment board correctly. These tasks are too easy to separate frontier models in one shot. That is a result too: the difficulty I measured in cli-bench was not coming from these tasks.

2. The same model did better without the agent loop. GPT-5.6 Luna passed all six in one shot. As a Codex agent it passed hot-loop 1 time in 3, with a worst-case speedup of 1.2x on the gridded dataset. In one shot it wrote a spatial grid that, timed on my machine against the same verifier, ran 42.5x faster on uniform points, 26.5x on gridded, and 14.6x on clustered. The caveats are real: one run against three, one dataset seed against three, and different machines. But it is the opposite of what I expected. Having a shell did not help the agent find the fast solution, and may have pulled it toward measuring and patching a slow one.

3. Every model reached for the same idea on hot-loop, and the only failure was invisible to the unit tests. All six bucketed points into a grid. Clustered points were the worst case for every passing model (14.6x to 41x on my machine), not gridded. Qwen3 Coder's grid passed all nine shipped unit tests and still undercounted: on one random case it found 84 pairs where there are 123. Only the randomized comparison against the naive version caught it. If my verifier had stopped at the shipped tests, that bug would have scored as a pass.

4. Most of my first-round failures were my harness, again. The first batch of Kaggle runs had 9 failures. Seven were mine:

  • Kaggle's task runtime does not have pytest installed. My fallback test runner did not support pytest fixtures (tmp_path, monkeypatch), so all six models "failed" the XSS task. Every one of them had escaped the output correctly.
  • My reply parser expected FILE: solve.py and rejected gpt-oss-120b's **FILE: solve.py**. I ran its script by hand against the same log and every answer was correct.

I fixed both, pushed new versions of those two tasks, and re-ran every model on them. The table shows the re-runs. The other two first-round failures were real. One is the hot-loop bug above. In the other, Qwen3 Coder's log script called statistics.quantiles with a method name that does not exist and crashed. On the re-run it wrote a different script and passed. That is one run at the default temperature, so treat any single pass or fail as a sample, not a verdict.

5. The corrected XSS task is passable by the textbook fix. All six models used HTML escaping, and all six passed the corrected tests and the probe. In cli-bench 0.9.1, the same approach failed. That confirms the problem was the tests, not the models.

What surprised me: across both versions of this benchmark, bugs on my side (two task specs, a test runner, a parser) caused more failures than the models did.

What I would measure next:

  • Harder cli-bench tasks: the flaky-test and CI tasks, and the four "houdini" probes that check whether a model games the verifier.
  • Three or more runs per model, so a single crash like Qwen's is visible as variance.
  • A version that gives the model a tool to run the tests, to see whether one retry helps or, as with hot-loop, hurts.

The lesson I am keeping: a benchmark that has never been run against a known-correct answer is not a benchmark yet. Every task in this port now ships with a reference solution, an empty reply, and a test-rewriting reply. A task only counts when the first passes and the other two fail, and that has to hold in the environment where the models actually run, not just on my laptop.

My Benchmark

Kaggle benchmark: https://www.kaggle.com/benchmarks/aks1321/cli-bench-one-shot

The six public tasks (each page has the full task code in its published notebook):

cli-bench is Apache-2.0: https://github.com/arjunkshah12345-hash/cli-bench

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to