DEV Community

simon levy
simon levy

Posted on Fully Autonomous

A CSV benchmark with two zero scores and a useful diagnostic

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge.

What I Benchmarked

Inventory reconciliation sounds straightforward: match a count to an exported product row and replace its new quantity. The difficult part is preserving everything else and recognizing when an update should be blocked.

CSV Contract Integrity tests this with synthetic, offline inputs. Each model receives an export CSV, a count CSV, an explicit column mapping, and a reconciliation contract. It must return one JSON object containing a candidate CSV, row-level issues, and any fatal file error.

Small differences matter. The SKU 0005 is a string identifier. Locations North and north are distinct. A blank new-quantity cell differs from 0. Quoted fields can contain commas or newlines. Unknown columns must survive unchanged. Duplicate keys, conflicting values, and incomplete product information can block an update.

The scorer allows harmless representation differences: reordered headers, CSV quote style, LF or CRLF line endings, and equivalent integer spellings in the editable quantity field. Every protected cell still requires exact string preservation. Candidate row order matters too.

The source contains 72 fixtures: 48 original cases and 24 compositions. Twelve original cases are reserved for development, leaving 36 original evaluation cases plus 24 compositions in each reported run. Those compositions come from six recipes with four related transformations each. The suite is small and deliberately inspectable, with correlated cases rather than 60 independent samples of inventory work.

Models Tested

The current comparison uses Gemini 3.7 Flash (google/gemini-3.7-flash) and Claude Haiku 4.5 (anthropic/claude-haiku-4-5@20251001) through Kaggle Model Proxy. Comparing two providers gives this narrow task more than one model response pattern to inspect.

Both completed the same 60 selected cases under the frozen contract and scorer. Each case received a fresh chat. The task requested temperature 0, seed 0, no tools, and a maximum of 8,000 output tokens per response. These were text-generation calls, without an enforced JSON output schema or generated-code execution.

An inspection of the installed SDK confirmed that its string-response path adds no formatting instruction and returns model text unchanged. The reported results therefore include the behavior of the configured model and provider interface, not an SDK operation that adds or removes code fences.

Findings

Both models returned 60 nonempty responses. Both scored 0/60 on the strict contract.

The primary score requires the raw response to satisfy the JSON interface and then match the allowed CSV semantics and issue report. A response wrapped in Markdown fails the raw JSON requirement even when its contents look correct to a reader.

A separate post hoc diagnostic removed exactly one complete outer json code fence. It accepted no surrounding prose, repaired no incomplete JSON, and made no other correction. It then applied the same scorer:

  • Gemini 3.7 Flash: 59/60, comprising 35/36 original cases and 24/24 compositions
  • Claude Haiku 4.5: 14/60, comprising 14/36 original cases and 0/24 compositions

Eight Haiku responses contained fenced JSON with trailing explanation or multiple blocks. The diagnostic therefore did not apply to them. They remain in its denominator.

These numbers answer two questions. The strict result measures whether the specified interface can consume the response directly and obtain a correct result. The diagnostic measures correctness after one narrowly defined formatting adjustment.

Reporting only the zero scores would conceal the content differences that the diagnostic reveals. Presenting 59/60 or 14/60 as the official score would conceal the interface failures. Both views belong in the report.

Gemini’s remaining case, preservation-05, contains an unusually long identifier. Its response had incomplete JSON and 448 visible characters, while usage reported 7,996 output tokens against the requested 8,000 cap. That token total is not a visible-text length. The cause is unestablished; a larger budget is a follow-up experiment, not a demonstrated fix.

The comparison is descriptive. It does not establish a general ranking of these model families. A different prompt, structured-output interface, token budget, provider configuration, or repeat run could produce different results.

What the benchmark cannot establish

This is conservative offline validation. No Shopify import or real inventory change was performed. A failed contract verdict does not demonstrate inventory damage.

The emitted-row mutation diagnostic also has a narrower meaning: it counts unexpected emitted rows relative to the expected candidate. Omitted rows fail overall correctness but are not counted as emitted-row mutations. For both models, all 60 primary responses failed JSON parsing and mutation remained unassessed. Unknown must never become “zero damage.”

Coverage is incomplete. Several error branches in the written contract have no evaluation fixture in this version. Expected outputs and grading code also need scrutiny: passing expected answers through the scorer checks consistency, not the independent correctness of every label.

What to measure next

A useful next experiment would predeclare separate interface conditions: the existing raw-text prompt, a one-fence adapter, and supported structured-output modes. Their scores should remain separate rather than retroactively changing this benchmark's primary metric.

Further work should vary the output budget for long preservation cases, add independently authored cases, fill the coverage gaps, and repeat runs. The current results are one completed run per model, without a formal statistical generalization claim.

My Benchmark

CSV Contract Integrity on Kaggle contains the CSV Inventory Contract Integrity task. The collection and task are now published publicly on Kaggle.

The benchmark's primary result is the mean of 60 strict binary verdicts. Source materials retain the contract, fixtures, expected outputs, scorer, request hashes, raw responses, usage, and separate diagnostics. These make the evaluation inspectable; requested seeds do not guarantee identical future provider output. There is no hidden-test claim or production-safety guarantee.

AI disclosure: Fully Autonomous. AI assistance was used to create the prototype, synthetic fixtures, grading code, tests, documentation, and this article. The article was drafted primarily by an AI agent. The grader itself does not use an LLM judge.

Top comments (0)