This is a submission for the Kaggle Benchmarking Challenge.
What I Benchmarked
A support model told me the correct fee, named the applicable policy in its explanation, and then put a different policy in source_ids. A person reading the explanation and software consuming the fields would receive inconsistent guidance.
Support Boundary Bench asks whether a model can select the right support decision and the evidence that justifies it. All policies, products and fees are fictional. No real customer data or actions are involved.
Responses must contain five JSON fields: decision, source_ids, missing_fields, conflict_ids, and answer_text. Decisions must be answer, clarify, or handoff. Labels stay outside the model prompt.
I prepared 10 development cases and froze 30 evaluation cases in 15 pairs. Each pair changes evidence order, a required fact, the event date, source authority, or an untrusted instruction. Some changes should alter the response; others should leave it unchanged.
A pair earns one point only when both cases pass every structural check. The score is passed pairs divided by 15. Format failures count against it; provider failures stop the suite and produce no numeric capability score. Explanation quality is a separate review dimension.
Models Tested
I selected openai/gpt-5.4-mini-2026-03-17 and google/gemini-3.7-flash from the available Kaggle models. Inputs, prompts, labels and scoring rules were identical. Temperature used the SDK default; no seed was specified; each case allowed one attempt.
After the first comparison revealed date-related failures, I recorded a plan for one complete GPT replication on all 30 unchanged cases, before scheduling it. This checks repeatability on the same cases, not generalization to a new holdout.
| Run | Valid contract | Structurally correct / assigned | Pairs passed | Request cost, USD |
|---|---|---|---|---|
| GPT baseline | 30/30 | 26/30 | 12/15 | 0.02076075 |
| GPT planned replication | 29/30 | 26/30 | 12/15 | 0.02058525 |
| Gemini baseline | 30/30 | 30/30 | 15/15 | 0.06399675 |
These runs had no missing cases or provider errors. Costs are exported request metrics, not a project invoice. Publication packaging runs are kept separate from this experiment table. On October 1, rebuilding version 4 to fix platform task selection required fresh runs: Gemini passed 30/30 cases and 15/15 pairs (USD 0.06229425); GPT passed 27/30 cases and 12/15 pairs (USD 0.02086425). Both returned 30 valid contracts. The public leaderboard displays these version 4 runs, not the historical rows above.
Findings
Correct decisions can conceal incorrect evidence
GPT selected the correct decision type in 30/30 baseline cases, but passed all structural fields in only 26/30. All four failures involved policy dates. A decision-only metric would miss them.
Here is v2-temporal-2-a from the planned replication:
| Item | Evidence or output |
|---|---|
| Event date | June 14, 2026 |
Policy te-2-a
|
58 fictional credits; ends June 15, exclusive |
Policy te-2-b
|
73 fictional credits; begins June 15, inclusive |
| Expected source | te-2-a |
| Model source_ids | ["te-2-b"] |
| Model explanation | Identifies 58 credits from te-2-a and says te-2-b does not apply yet |
The inclusive-start/exclusive-end rule was explicit in the prompt. Mechanical label derivation and AI-assisted review checked the dates; neither is being presented as human approval.
In the replication, three temporal responses again explained the right policy while returning wrong source fields. Two case IDs failed in both GPT rounds; other failures changed. This is a repeated pattern in this suite, not a universal failure rate.
The same total score can hide different failures
GPT scored 12/15 pairs twice. The baseline had four valid responses with wrong structural fields. The replication had three such responses and one contract failure: hand-off instead of the required handoff.
That response described the conflict, but its enum was invalid for the machine interface. I did not normalize it after seeing the result. This is an interface failure, not evidence that the model misunderstood the conflict. Replication decision accuracy is 29/29 among valid responses, not 30/30; assigned-case and pair denominators still include the invalid response.
Gemini passed all 15 pairs in its original evaluation run. The small sample, shared templates and unequal repetitions do not establish a general model ranking.
Verify the evaluator too
An earlier source file hard-coded GPT. A job labeled Gemini therefore still called GPT. The importer detected the identical actual model IDs and refused the comparison. That extra GPT run and its USD 0.02112975 cost remain separate; they were not relabeled or selected to replace a first result.
The corrected entry point uses platform-injected kbench.llm. Import checks the requested model against recorded evidence. For each reported run I verified 120 child-file hashes, all 30 recorded prompts, frozen input/label/scorer hashes, and agreement between the numeric parent result and independent scoring. Earlier development failures remain separate too.
What changes for a support workflow
Validate source IDs against product and event date before downstream software relies on them. A fluent explanation and valid decision enum are insufficient. This benchmark does not prove that adding such a validator improves customer outcomes; that requires another experiment.
Human label and explanation review remains pending. Development and evaluation share synthetic templates. Next I would test unseen policy templates and date boundaries, then evaluate a predeclared validation intervention. Tuning on these same cases would make them development data.
My Benchmark
Explore Support Boundary Bench on Kaggle.
Version 4 task and results · Backing notebook and evidence outputs.
Built with the Kaggle Benchmarks SDK. Frozen inputs, labels, scorer and raw per-case evidence are included in the backing notebook and its ZIP outputs. Codex assisted with implementation, synthetic data, analysis and writing; this does not replace human review.
Top comments (0)