This is a submission for the Kaggle Benchmarking Challenge.
What I Benchmarked
An AI component advising an API client has two jobs: choose the right action and return something the client can actually parse. RetryBudget-24 tests those jobs together, then separates their failure modes for analysis.
Each question describes a failed or completed request and supplies an explicit, synthetic retry policy. The model returns exactly two fields:
{"action":"RETRY","delay_ms":100}
The actions are DONE, STOP, RECONCILE, and RETRY. The policy covers status precedence, attempt limits, idempotency guarantees, capped exponential backoff, Retry-After, and a remaining-time budget. A retry is permitted only when its wait plus the next attempt's estimated duration fits the budget. Equality is allowed.
This is a text-based policy-decision task. It does not execute HTTP requests, measure inference-serving latency, or demonstrate safe production retries.
Twelve controlled pairs
The formal set contains 24 original synthetic cases arranged as 12 pairs. Within each pair, exactly one input field changes. That makes the boundary being tested easier to inspect than in a collection of unrelated puzzles.
For example, keep a 100 ms wait and a 200 ms next attempt fixed:
- 299 ms remaining means STOP.
- 300 ms remaining means RETRY with a 100 ms delay.
Other pairs change whether a server guarantees deduplication for an idempotency key, whether the attempt limit has been reached, or whether the backoff cap binds.
The cases, fixed order, prompts, oracle, and scorer were hash-frozen before any model call. Four separate development cases checked execution without tuning the prompts against model responses. Expected answers were checked against the deterministic oracle; 20 offline tests checked the original benchmark implementation.
The primary score is exact correctness on both output fields, divided by 24. JSON whitespace and key order are allowed. Code fences, explanations, extra or duplicate keys, floats, and booleans in the integer field fail the contract. No LLM judge is involved.
For context, deterministic baselines score 24/24 for the policy oracle, 6/24 for always STOP, and 7/24 for a status-only heuristic that retries transient statuses after 100 ms. The oracle is a correctness reference, not a competing language model.
Models Tested
The experiment used two models available through Kaggle's Model Proxy:
google/gemini-2.5-flashanthropic/claude-haiku-4-5@20251001
This small, cross-provider comparison kept the experiment inexpensive. It is not a survey of the strongest available models.
Both received the same frozen prompts through Kaggle Benchmarks SDK 0.6.1, with max_tokens=512, seed 0, and temperature 0 requested. Importantly, the inspected SDK path reported temperature unsupported for both and dropped that setting. Actual temperature was therefore not controlled. Reasoning settings remained at provider defaults, and requested settings do not establish identical effective inference budgets.
Each case used a separate chat. There were no retries for wrong answers. Client SDK retries were disabled; retries inside the proxy were not observable.
Findings
The planned experiment
| Model | Strict score | Valid schema | Both members of pair correct |
|---|---|---|---|
| Gemini 2.5 Flash | 16/24 | 22/24 | 4/12 |
| Claude Haiku 4.5 | 5/24 | 9/24 | 1/12 |
The strict gap looks large. Inspecting the responses changes what can reasonably be concluded from it.
Haiku wrapped 15 of its 24 answers in complete Markdown code fences. Nine of those contained the correct action and delay. Gemini produced one complete fence with correct content and one unclosed fence.
A narrowly defined, post-hoc diagnostic removed only complete outer fences, without repairing JSON, coercing types, or extracting objects from prose. Correctness then became 17/24 for Gemini and 14/24 for Haiku. These are diagnostic counts, not replacement benchmark scores.
The distinction matters for integration design. A strict JSON consumer really would reject those fenced responses. But calling every rejection a reasoning failure would hide a substantial part of what happened. A parser adapter or constrained-output interface could be worth testing in a separately frozen experiment; neither was retrofitted into these scores.
Some errors remain after formatting is separated
Both models chose RETRY on the 299 ms budget example, although 100 + 200 exceeds 299. At 300 ms, both returned the correct decision content, with Haiku's answer fenced.
On an unsafe timeout with no idempotency guarantee, Haiku returned RETRY where the policy required RECONCILE. On a capped-backoff case where min(1000, 700 * 2) should produce 1000 ms, Gemini returned 1400 and Haiku returned 700.
Those examples identify concrete policy boundaries worth testing. They do not establish how either model would behave across real API traffic.
An unintended second experiment
There was an execution-control mistake during publication preparation: Build Task was treated as a save/build operation, but it executed the original notebook again. The additional run completed before it could be stopped. That exceeded the planned 56 client calls and must be counted.
| Batch | Gemini formal score | Haiku formal score | Calls including development | Reported quota cost |
|---|---|---|---|---|
| Planned experiment | 16/24 | 5/24 | 56 | $0.0521384 |
| Unplanned Build rerun | 18/24 | 4/24 | 56 | $0.0523159 |
| Total | Separate batches | Separate batches | 112 | $0.1044543 |
All 112 recorded calls completed. The additional batch used identical prompt hashes and unchanged scoring. Its outputs are retained separately; the better Gemini result does not replace the planned result, and the batches are not pooled into a new headline score.
The amounts are reported Model Proxy quota consumption within the account's available free quota, not a cash purchase. They show that this particular experiment was inexpensive, not that future runs have a guaranteed price.
This also exposed a benchmark-workflow requirement: publication steps need their own execution budget. A call ledger inside one notebook session cannot impose a global cap across fresh platform executions.
What the results cannot tell us
This is a small synthetic set, with related pairs, one planned sample per case, and one accidental repeat. It supports failure analysis rather than a general model ranking or statistically reliable reliability estimate.
The saved artifacts do not contain finish_reason. In the planned Gemini test, 21 of 24 reported output-token counts were at least 450 against the requested 512-token limit. That warrants caution, but it does not prove truncation or explain a particular error. The models' effective generation conditions were not fully controlled.
The next useful experiment would predefine strict versus constrained-output conditions, more policy boundaries, and planned repetitions. It would also verify the platform's build behavior before enabling any model-call entry point.
My Benchmark
Kaggle benchmark collection: RetryBudget-24: Policy and JSON Contracts
Source, frozen cases, manifests, and raw results: Planned experiment snapshot and unplanned Build rerun snapshot. The self-authored benchmark code is published under Apache 2.0.
The published Kaggle task and collection contain only the unplanned batch's Gemini result, 0.75 (18/24). The two-model comparisons above come from separately preserved and audited raw experiment artifacts; they are not a claim that both models appear on the platform leaderboard. The original execution adapter also depends on its notebook context. A self-contained publication wrapper has been prepared and tested offline, but it has not been run against real models or substituted for the measured version.
AI assistance was used to design and implement the synthetic benchmark, run the workflow, audit saved outputs, and draft this article. The task uses original synthetic cases rather than third-party evaluation data. The SDK is credited below. The article explicitly retains the AI-assisted workflow's unintended rerun instead of hiding it.
Official references:
Top comments (0)