This is a submission for the Kaggle Benchmarking Challenge
What task(s) did you run?
In the first Agentic Football Cup — a simulated football league in which each 'player' on the teams is an LLM agent controlled using a prompt file — I conduct a manual audit of those prompt files before each match window, looking for bugs such as collisions between forwards in zones, loopholes in the handoff procedures between defenders, and scope conflicts in the anti-spam rules. This process is time-consuming.
There was a desire to know: could a large language model carry out the audit instead? And more importantly – could I place trust in it?
So I created a benchmark that acts as a junior auditor and evaluates the model against the ground truth I already have. The task provides the model with two items:
A retired team-prompt JSON (one of three Bugle Crowns builds)
The same stress-test audit prompt a human would follow
The system must provide the structured findings—including the defect ID, a GREEN/YELLOW/RED flag, and a direct quotation of the evidence—and also give one final decision, either SHIP AS-IS, FIX-THEN-SHIP, or DO NOT SHIP. Provide only the findings, with no rewrites.
The three builds are the clever part:
v2.5A.1 — a faulty version containing three known defects (a collision in the defensive-corner zone, a loophole in the intercept lane-handoff, and a conflict between the anti-spam scope). The audit has to identify all three.
v2.5A.2 is the fixed control build, so any finding by RED in this instance is a hallucinated false positive.
v2.5A.5 — a more difficult version containing eight possible findings of varying severity, designed to distinguish between competent auditors and those who are good at their job.
Scoring has three components: recalling known defects, ensuring there are no REDs on the control build, and verifying that the final decision matches the ground truth. The entire process is a Kaggle benchmark task with structured output so that a machine can check each run.
Against which models did you test it?
Three models in the final lineup, nine clean runs (every model against every build):
The tests comparing google/gemini-2.5-flash and google/gemini-2.5-pro aim to determine whether the larger model is actually the better auditor.
The small OpenAI model, openai/gpt-5.4-mini, was used to see whether a lightweight model can audit.
Two further attempts were made and had to be abandoned. anthropic/claude-sonnet-4-5 failed to get a single run: the Kaggle LLM gateway returned HTTP 429 ("heavy load") on all three attempts, twice, four(ish) hours apart — this was due to capacity ceilings on the part of the provider, not because of an invalid model ID. deepseek-ai/deepseek-r1-0528 did manage to access the model once, but the benchmark SDK's response parser was unable to process its reasoning traces. The fact that a reasoning model has reasoning that breaks the system is itself a finding, but the model cannot be used for scoring. (This is because no model from Meta is available on the Kaggle platform, which is the reason DeepSeek was assigned to that position in the first place.)
Each run had its own new chat context, with roughly 7,000 input tokens per call. Total spend was about \$0.52 of the \$10 daily quota, including failed attempts (rate-limited calls cost nothing).
What are the main insights?
1. All the models advised adopting the FIX-THEN-SHIP approach in all cases, including those relating to the clean build.
The headline is this: the v2.5A.2 control build contains no defects and should therefore be shipped as-is. Yet all three of the working models still marked it FIX-THEN-SHIP. They are cautious rather than precise. Although none produced any RED findings on the clean build (since the no-RED check passed at all locations), none would approve it. If you used one of these as your audit gate, nothing would ever get shipped.
2. Recall is inadequate since the auditors fail to detect actual bugs.
With the buggy v2.5A.1 and three known defects in the prompt, only gpt-5.4-mini picked up any of them—two out of the three. Both of the Gemini models caught none of the defects in the sweep runs. You wouldn't hire an auditor who misses two out of three known defects.
3. The hard set puts the models apart.
With respect to the eight findings on v2.5A.5, gemini-2.5-pro identified four, flash found three, and gpt-5.4-mini found none. The more powerful model did, in fact, justify its cost in this instance—though it is worth noting the irony that gemini-2.5-pro was the most effective at detecting bugs and yet still failed to approve the clean build.
4. Outputs vary run to run.
The first smoke test I ran picked up the defect related to the D1 corner collision; however, a resweep of the same model with the same build did not. When conducting benchmarking of auditors, a single run is insufficient — and if you are using an LLM auditor, you must know that its answers are inconsistent.
5. The main point: models check themselves like anxious beginners.
Flag every item, commit to nothing, and fail to identify the real bugs while refusing to mark the clear ones. This is precisely the kind of failure mode you should look for before letting an LLM be involved in any production prompt analysis—and it is the opposite of what the marketing claims these models are good at.
The next things I would look at are having a judge-LLM review the borderline YELLOW findings to see whether a second model can prioritize those that the first one has flagged; examining the latency and cost associated with each model (since which auditor would you really hire on a per-dollar basis?); and carrying out more builds, some of which involve cases where the correct decision is actually ambiguous.
Top comments (0)