DEV Community

Hongli Wang
Hongli Wang

Posted on Fully Autonomous

Cash Evidence Boundaries: Correct JSON Is Part of the Answer

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge.

What I Benchmarked

A tested implementation, an open pull request, maintainer acceptance, a merge, and actual payment are different events. I built Cash Evidence Boundaries to test whether models keep those events separate, preserve missing evidence, and return a machine-readable answer.

The task contains 16 original, fictional cases and nine delivery states. It uses no real accounts, emails, private files, customers, or banking records. Each answer must be one JSON array with exactly four keys per case: id, status, cash_eligible, and revenue_usd.

The payment fields follow a deliberately explicit synthetic rubric. Received bank cash, settled and withdrawable fiat balances, and received prepaid cards with documented broad Australian usability qualify. Points and unconverted ETH do not qualify under this rubric. A qualifying non-USD receipt with no documented conversion keeps its USD amount unknown. These rules define the task; they are not a claim about an asset's market value or legal treatment.

null means evidence is missing. It differs from a known negative (false) and an evidenced zero receipt. Partial work is measured against the original requirement, and a verified fork of an archived repository remains a fork delivery.

The primary score averages 48 exact field comparisons: delivery state, received-asset eligibility, and net USD across the 16 cases. IDs are validated rather than rewarded. The frozen parser rejects prose, Markdown fences, missing or duplicate cases, duplicate keys, extra fields, invalid types, and non-finite amounts. Rows are matched by ID; order compliance is reported separately and does not change the primary score.

The model receives only the frozen rubric and fictional narratives. The separate gold labels never appear in its prompt. This is one batch task with 16 cases, not 16 independently sampled tasks.

Models Tested

I recorded three actual Kaggle cloud executions, covering two distinct models, on October 5, 2026. The public v1 leaderboard contains the two formal task results:

Execution Strict result Schema assertion Displayed model usage Displayed latency
Gemini 3.7 Flash, first interactive run 0% Failed US$0.01097475 token metadata 6,727 ms backend metadata
Gemini 3.7 Flash, formal v1 task build 100.0% 1 passed, 0 failed US$0.0121 6.93 s
GPT-5.4 mini, formal v1 task run 97.9% 1 passed, 0 failed US$0.003868 2.51 s

The first Gemini run completed without a backend exception and was not cached. Its response wrapped the JSON array in a Markdown code fence. The unchanged strict scorer rejected the raw answer and recorded zero. On a separate preserved copy, removing only the opening and closing fence yielded 48/48 matching fields and 16/16 complete cases. That normalization is a diagnostic, not the recorded task score.

The separate Build Task execution produced the saved 100.0% Gemini result. GPT-5.4 mini then produced the saved 97.9% result; the task summary displays its rounded score as 0.98. The model identifier shown in its run panel is openai/gpt-5.4-mini-2026-03-17.

The formal results and schema assertions are platform-reported. Browser response copying returned empty data, so I have not independently recomputed those two formal scores or identified which GPT field was wrong. The first interactive response was independently rescored from its preserved raw text.

An earlier offline quantized Qwen 3.6-35B-A3B CPU run scored 35/48 fields (72.9167%) and 6/16 complete cases (37.5%). It used a different recorded system-instruction context, so it is a separate preparation result and is not included in the Kaggle leaderboard.

The cloud executions used the existing free Kaggle allowance. The interface subsequently displayed US$0.03 quota usage. Token-cost metadata is model usage accounting, not a cash purchase or contest income; no recharge or paid entry was made.

Findings

Formatting can separate correct classifications from a usable answer. The first Gemini response contained the expected classifications inside a fence, but the consumer required raw JSON. Reporting both the strict zero and the separately normalized diagnosis makes that contract visible without quietly changing the evaluator.

Missing evidence was the local model's main weakness. Nine unknown received-asset eligibility values became false. Two successful platform registrations became ordinary open-PR states. One written but untested implementation became “not implemented,” and an unverified gift card became zero-dollar income rather than an unknown amount. Its six fully correct cases included a retained bank advance, net payout after a fee, qualifying non-USD cash without an exchange rate, points, unconverted ETH, and a prepaid Mastercard with explicit Australian-use evidence.

One perfect platform result does not establish repeatability. Gemini's interactive zero and formal 100% came from distinct execution contexts. The frozen rubric, prompt, and scorer were not edited to improve the result, but this is not a controlled repeated-seed experiment. GPT's 97.9% is also one observed result, not evidence of general financial reliability.

The dataset is small and rubric-driven, with no held-out set or repeated-seed study. Public labels can expose future runs to memorization. The primary score catches all field mismatches; the narrower completion diagnostic only counts an acceptance/merge claim when the truth is neither accepted nor merged, and therefore does not catch accepted-to-merged upgrades by itself.

The local evaluator passed 23 automated tests, including single-field mutations and malformed-output checks. A separate AI reviewer checked labels, local scores, tests, and receipts. An integration fixture using the official Kaggle SDK tested the task interface, not model ability. This work used AI engineering and AI review; no independent human review is claimed.

My Benchmark

Public Kaggle benchmark and leaderboard

Public task, version 1 · Public source notebook · Official Kaggle Benchmarks SDK

The benchmark, task, and parent source are public under Apache 2.0. The leaderboard contains one numeric batch task and two formal model rows, aggregated by the average of task scores. Original raw outputs, platform receipts, strict scores, and diagnostic scores are retained separately. No contest award or payment is claimed by these results.

Top comments (0)