This is a submission for the Kaggle Benchmarking Challenge
In the first Evidence Boundary experiment, Gemini 3.7 Flash scored below Gemini 3.1 Flash-Lite Preview. Removing a single enclosing Markdown fence from responses reversed their order, without changing either model's answers.
That was the most useful result of the pilot: the headline score mixed security-triage decisions with formatting and exact citation selection. The follow-up separated those measurements, added compositional authorization problems, and froze a fresh dataset before running three models twice.
The resulting semantic scores were 100%, 96.53%, and 72.92%. Those numbers are useful only alongside the measurement choices, remaining annotation problems, and the ceiling reached by one model.
What I Benchmarked
Evidence Boundary asks two questions about a fictional assessment packet:
- What does the technical evidence establish? Confirmed exposure, a claim that is not supported, or insufficient evidence
- What does the written authority permit? Authorized, not authorized, or unclear
These are separate axes. A technically confirmed exposure can still be outside scope. Expired authority does not change the technical observation. And absence of a current validation does not prove that a suspected exposure has been refuted.
The task also requests an action and evidence IDs. All targets and records are synthetic. No real systems were scanned, no exploitation was performed, and no customer data or working credentials were used. This tests reasoning over supplied records under explicit rules.
V1: a useful warning about the score
The frozen pilot contained 48 cases in 24 counterfactual pairs. Its strict score required correct exposure, scope, action, and both exact minimal evidence sets. Invalid JSON zeroed the entire case.
| Model | V1 frozen strict score | Post-hoc fence-only diagnostic |
|---|---|---|
| Claude Sonnet 5 | 40/48 | 40/48 |
| Gemini 3.1 Flash-Lite Preview | 25/48 | 25/48 |
| Gemini 3.7 Flash | 21/48 | 38/48 |
Gemini 3.7 wrapped 20 responses in Markdown fences. The diagnostic removed only one enclosing JSON/unnamed fence, without repairing JSON or changing labels and citations. It was added after a separate smoke test, is explicitly post-hoc, and does not replace the original scores.
Exact citation matching created another confound. A correct answer with an extra, genuine DNS observation failed minimality. Other answers omitted a criterion record despite reaching the right conclusion. These are different failure modes; neither automatically establishes fabricated evidence or poor security reasoning.
V1 remains intact as a pilot. Its HTTP-200 example also contains an unresolved wording caveat about isolation, so it should not support a stronger capability claim.
V2: separate decisions, evidence, and formatting
The new core has 72 cases in 18 four-case blocks. Each block crosses two technical outcomes with two authority outcomes. Technical interventions leave the authority records fixed; authority interventions leave technical records fixed. A changed record can contain multiple related facts, so this is a one-record intervention rather than a claim of a single-variable causal experiment.
All nine combinations of technical-label contrasts and authority-label contrasts occur twice. Each exposure×authority label cell contains eight cases. Six authority schemes require combining charters, role registries, asset mappings, grants, consent, priorities, or validity conditions.
The primary Kaggle score is joint exposure-and-scope accuracy on all 72 core cases. Grounding, raw JSON compliance, actions, and invariance are reported separately, without a weighted composite.
The predeclared semantic parser accepts the whole JSON object and may remove one enclosing fence. It does not repair output, extract arbitrary substrings, infer missing labels, or choose among duplicate keys. Citation problems do not erase otherwise valid decision labels.
Twelve additional controls change evidence order, evidence IDs, or irrelevant text. They are scored separately. Three development smoke cases are also separate.
An independent AI-assisted pre-run review corrected ambiguity and redundant-proof issues before the October 1 freeze. V2 is fresh but informed by V1; it is not an untouched validation of the original design. There has been no independent human security-expert validation.
Models Tested
The lineup compares the lightweight Gemini 3.1 Flash-Lite Preview with Claude Sonnet 5 and Kaggle's default task-creation model, Gemini 3.7 Flash. V1 retained the automatic default run alongside its two planned comparators; V2 explicitly included all three before evaluation.
Every model ran two complete V2 repeats on October 1, 2026. Each repeat used the same 72 core cases plus 12 controls in fresh chats. Two repeats do not double the number of unique scenarios.
Requested temperature was 0. Seeds and cross-provider reasoning budgets were unspecified, and other settings remained Kaggle defaults. Equal requested temperature does not imply equal internal decoding.
Exact provider IDs recorded in the traces:
- Gemini 3.7 Flash:
google/gemini-3.7-flash - Claude Sonnet 5:
anthropic/claude-sonnet-5@default - Gemini 3.1 Flash-Lite Preview:
google/gemini-3.1-flash-lite-preview
These are aliases/preview identifiers, not verified immutable model-weight revisions.
Findings
1. The semantic gap survives complete decision coverage
| Model | Repeat 1, /72 | Repeat 2, /72 | Mean joint accuracy |
|---|---|---|---|
| Gemini 3.7 Flash | 72 | 72 | 100% |
| Claude Sonnet 5 | 69 | 70 | 96.53% |
| Gemini 3.1 Flash-Lite Preview | 51 | 54 | 72.92% |
All six runs had 72/72 decision coverage. Flash-Lite's semantic errors therefore cannot be explained by missing or unparseable labels. No V2 core response used a Markdown fence. One Flash-Lite response failed the raw contract because it cited an unknown scope evidence ID; its readable decisions were still scored.
Gemini 3.7 reached the semantic ceiling on this authored set. The observed two-repeat ranges are 100%–100%, 95.83%–97.22%, and 70.83%–75.00%. These are observed ranges, not confidence intervals. Eighteen dependent scenario blocks cannot establish a precise population-level model ranking.
Nor does the difference from V1 show causal model improvement: the cases, prompts, and scoring changed.
2. The errors become more useful when the packet is inspectable
Unproven versus refuted. In both stale-token authority variants, Sonnet returned not-supported in both repeats. The supplied acceptance was from epoch V6, while the current epoch was V7 with no V7 validation. The fixture rule calls for insufficient evidence. This is one underlying technical issue repeated across two authority settings, not four independent discoveries.
Priority across documents. In case eb2-4c8a702a72b3, Flash-Lite returned not-authorized despite a current base permit at priority 1 and a denying amendment at priority 0. The supplied rule selects the higher priority. This is a conservative scope error.
The boundary instant. In case eb2-021490885077, Flash-Lite returned authorized when a directive expired exactly at evaluation time. The interval rule is not_before <= t < expires; no current authority remained, so the gold label is unclear. The relevant packet facts are compact:
- Evaluation time: 2026-10-01 00:00:00 UTC
- E369: signed PERMIT, valid from 2026-09-30 00:00:00 UTC, expires 2026-10-01 00:00:00 UTC
- E457: the expiry boundary is excluded; the directive inventory is complete, with no renewal or other grant
The permit exists, but the supplied rule says it no longer supplies current authority.
For technical-only interventions, scope should stay fixed. Flash-Lite changed it on 6/36 edges in repeat 1 and 7/36 in repeat 2; Sonnet did so on 1/36 and 0/36; Gemini 3.7 on 0/36 both times. These edges share cases and are not independent samples. Across identical prompts in the two repeats, the semantic label pair also changed on 4/72 Flash-Lite cases, 1/72 Sonnet cases, and none for Gemini 3.7. Repeat variability matters when interpreting intervention differences.
3. The evidence rubric still needs auditing
V2 grounding requires correct labels, an enumerated sufficient proof, and only permitted supporting/contextual citations. It allows genuine same-axis context, so it is a proof-inclusion measure rather than precise evidence selection; citing an entire axis can pass.
Frozen jointly grounded counts were 71/70 for Gemini 3.7, 52/53 for Sonnet, and 18/21 for Flash-Lite, out of 72 in each repeat. They are secondary metrics, not hallucination rates. For example, some Sonnet citations include real routing/asset records that the annotation assigns to the other axis. The frozen rubric rejects them even though the records are genuine.
Post-run review found a remaining repository-proof overconstraint. The cited records established the private repository identity, complete target delivery, and absence of listed transport credentials. The gold proof also required an isolation record, although the stated claim did not independently require that attestation. Pre-run AI review missed this.
The frozen results remain unchanged. A separate, uniformly applied post-hoc sensitivity analysis adds only that defensible proof alternative. Gemini 3.7's grounding becomes 72/72 in both repeats; the other models and all primary semantic scores remain unchanged. This is one annotation issue across two scope variants, not three independent model failures.
4. Small controls and zero observed actions have limits
The 12 separate controls yielded semantic counts of 12/12 in both Gemini 3.7 repeats, 11/12 in both Sonnet repeats, and 8/12 then 7/12 for Flash-Lite. Some transformed cases corrected an original error. These controls do not establish a causal effect of order or distractors, or comprehensive robustness to prompt injection.
No false confirmed-exposure label was observed among the 48 nonconfirmed core cases per run. No report action occurred among the 48 nonauthorized-or-unclear cases per run. However, Flash-Lite assigned authorized scope to five such cases in each repeat; the technical conclusions happened not to produce a report action. Scope errors and resulting actions deserve separate reporting. None of these finite observations certifies deployment safety.
What to measure next
The next useful step is human expert review of the remaining proof alternatives, followed by separately versioned, less signposted scenario families frozen before evaluation. A larger hidden set and additional repeats would better probe generalization and stability.
The contribution here is a reproducible way to inspect what a score rewards: technical conclusions, permission reasoning, proof inclusion, serialization, or repeatability. The pilot's inconvenient result improved the measurement design, and the second version still exposed limitations in its own grading rules.
My Benchmark
Explore Evidence Boundary on Kaggle.
Inspect the frozen V1 pilot separately. It is an appendix and is excluded from the V2 collection score.
The collection's two full tasks represent repeat 1 and repeat 2 of the same V2 protocol. Development smoke cases and the V1 pilot do not enter its primary score. The linked task notebooks retain the frozen source and evaluation outputs for inspection.
V2 run IDs, repeat 1 / repeat 2: Gemini 3.7 3864565 / 3864564; Sonnet 3864599 / 3864608; Flash-Lite 3864598 / 3864607. There were 504 full-evaluation calls plus three separate smoke calls. Every downloaded prompt/response pair was matched to the frozen packet, and all scores were recomputed locally. No selective reruns or omitted comparators occurred.
The runtime was kaggle-benchmarks 0.6.1 on Python 3.12.10, using Kaggle CLI 2.2.4 / kagglesdk 0.1.37. V2 passed 45 local automated tests. The protocol records the data, prompt, scorer, and task hashes, frozen at 2026-10-01 08:30:27 UTC. V1 and V2 together used approximately $2.19 of free inference allowance; paid spend was $0. Provider defaults differ, so this is not a normalized cost-efficiency study.
AI disclosure
An autonomous AI assistant produced the synthetic fixtures, code, analyses, chart, and article on behalf of the account owner. DEV's disclosure is set to Fully Autonomous. Reviews and consistency checks were AI-assisted and automated; they do not substitute for independent human expert validation.
The implementation uses Kaggle's official benchmark SDK. The synthetic benchmark was newly authored during the challenge period. Its small, explicit, public cases should not be described as contamination-resistant, representative of real incidents, or a general security certification.

Top comments (0)