This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
A JSON response can parse successfully and still change the wrong state. A response can also contain the right state but arrive wrapped in Markdown that breaks a strict consumer. I built Bilingual Patch Contracts to keep those two failure modes visible.
The task has twelve handcrafted state-update scenarios, each with an English, Chinese, and code-switched instruction body: 36 prompts. Each triplet shares an initial state and expected answer. The contract prefix stays in English, and the output keys stay canonical English. This tests changing the instruction body's language under a shared contract; it is not a fully Chinese interaction benchmark.
Every case begins with:
{"route":"email","delay_minutes":20,"active":true,"tags":["alpha","beta"],"note":"seed"}
The scenarios cover later corrections, negation, null versus empty values, ordered and case-sensitive tags, hours-to-minutes conversion, sequential conditions, instruction-like literal data, and exact copying of Unicode, backslashes, quotes, and a newline.
A pass requires the entire response to be one JSON object with exactly five keys, valid types, and every expected value. I do not strip Markdown, repair outputs, or ask another model to judge. Whitespace, object key order, and equivalent Unicode escapes are accepted. Duplicate keys, extra fields, nonfinite values, and booleans or floats in integer fields fail. Array order matters.
I use ordinary text generation, with temperature 0 and seed 0 requested through the SDK, and a fresh isolated conversation for each case. There is no constrained JSON decoding, schema enforcement, or tool use. Provider behavior may still vary between runs.
Models Tested
I ran the complete suite on Kaggle on October 1, 2026, using task version 2. I selected three lightweight models from different providers available in the platform catalog:
| Model | Exact Kaggle model identifier |
|---|---|
| Gemini 3.7 Flash | gemini-3.7-flash |
| GPT-5.4 nano | gpt-5.4-nano-2026-03-17 |
| Claude Haiku 4.5 | claude-haiku-4-5-20251001 |
I also attempted qwen3-next-80b-a3b-instruct. Both the pilot and version 2 attempt stopped with HTTP 429 and a provider message about heavy load. It has no complete score, so it is excluded from the leaderboard rather than assigned zero.
Version 2 fixes task registration so Kaggle selects the whole-suite aggregate instead of a helper function. The prompts, fixtures, and scorer were unchanged. All results below come from complete version 2 runs. I downloaded the raw responses, checked all 36 unique case IDs against the frozen prompts and answers, and independently recalculated the saved scores.
Findings
| Model | Strict exact match | Valid JSON | Valid schema | English | Chinese | Mixed |
|---|---|---|---|---|---|---|
| Gemini 3.7 Flash | 36/36 (100%) | 36/36 | 36/36 | 12/12 | 12/12 | 12/12 |
| GPT-5.4 nano | 24/36 (66.7%) | 36/36 | 36/36 | 7/12 | 8/12 | 9/12 |
| Claude Haiku 4.5 | 0/36 (0%) | 0/36 | 0/36 | 0/12 | 0/12 | 0/12 |
Valid JSON does not prove a correct update
GPT-5.4 nano produced a valid object with valid field types on every case, but twelve answers contained wrong values.
For case-sensitive-en, the instruction body is:
Tag strings are case-sensitive. Add Beta. Remove beta. Add ALPHA. Add alpha.
Starting with ["alpha","beta"], the expected tags are ["alpha","Beta","ALPHA"]. The raw model response was:
{"route":"email","delay_minutes":20,"active":true,"tags":["alpha","beta","Beta","ALPHA"],"note":"seed"}
The lowercase beta should have been removed. A JSON parser or type validator alone would accept this answer.
A presentation choice can break the contract
Claude Haiku 4.5 wrapped every answer in a Markdown code fence despite the explicit instruction “No Markdown or explanations.” For last-write-en, its full response was:
```json
{"route": "sms", "delay_minutes": 0, "active": true, "tags": ["alpha", "beta"], "note": "seed"}
```
The values in this example are correct, but the full response is not a JSON document. Its zero is a strict interface score, not evidence that it understood none of the requests.
As a separate diagnostic, removing only the complete outer code fences would make 33/36 answers pass the same value checks. That counterfactual is not the benchmark score and does not repair any leaderboard output. It shows why formatting compliance and state correctness deserve separate reporting.
Language totals need paired inspection
Nano's mixed total exceeded its English total by two cases. Within the twelve matched scenarios, seven passed in both English and mixed, three failed in both, and two passed only in mixed. For English versus Chinese, six passed in both, three failed in both, one passed only in English, and two passed only in Chinese.
These observations locate cases to inspect; they do not establish that the model is generally stronger in Chinese or code-switching. The bodies were hand-authored, and phrasing and token lengths are not perfectly controlled.
A ceiling is a prompt to extend the suite
Gemini passed every case in this run. This suite therefore cannot distinguish its reliability beyond these examples. I would next add longer instruction chains, more literal-copy edge cases, and repeated runs, while keeping new cases separate from this frozen result.
This is a small diagnostic benchmark, not a general model ranking. The three language variants are paired observations: there are twelve independent semantic scenarios, not thirty-six independent problems. A single run does not establish production reliability, and shared English contract instructions limit the multilingual claim. I did not benchmark latency, cost, or tool calling.
My Benchmark
Bilingual Patch Contracts on Kaggle
The collection uses one numeric task whose score is strict exact matches divided by 36. Its overall score is the average of task scores; with one task, it equals that task's score.
Inspect the version 2 task and model outputs and the public backing notebook for the embedded cases, expected states, scorer, and run artifacts (contract_results.json and contract_summary.json). Infrastructure errors abort the suite rather than silently shrinking the denominator.
I authored the prompts and answer fixtures for this personal project. The implementation uses the Kaggle Benchmarks SDK, following the official platform documentation.
Top comments (0)