DEV Community

Zhuoli Liang
Zhuoli Liang

Posted on

Which clause wins? I benchmarked whether LLMs follow the rules that actually govern

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge.

What I benchmarked, and why

I enter a lot of contests and bounties. The rules for a single contest are never on one page. There's a landing page, an FAQ, a contest-rules page, and the platform's general Official Rules, and they drift apart. While building Fine Print (an agent that reads contest rules for you), I kept hitting the same question: when the friendly FAQ says one thing and the Official Rules say another, does the model notice which one controls?

So I built Fine Print Bench: fictional rule packets, each with a plain question ("What is the controlling close instant in UTC?", "Is a team of five allowed?", "Which required artifacts are missing?"). The model has to return a typed answer, the minimal clause IDs that justify it, and a confidence. 160 items in two splits:

  • easy (80 items): short packets, one precedence hop, eight categories:
Category What it isolates
Explicit authority Rules vs. FAQ/landing/blog when a precedence sentence exists
Scope & amendments A newer FAQ is not an Amendment; an amendment that only names Track B
Unresolved authority Equal-authority pages with no tie-break, plus controls where the pages agree
Exact deadlines DST ending the same day, the date line, a submission 30 seconds late
Business-day windows Holidays, "counting the decision date as day one", a trigger date nobody gave you
Eligibility Missing residency, a subsidiary employer, age of majority ≠ 18
Submission artifacts Conditional requirements, a draft post, a private repo
Notification & payout "first attempted notification" vs. "after you receive our email"
  • hard (80 items): real-length rules (~800 words of boilerplate around the deciding clause), two-hop precedence (Amendment → Contest Page → Rules, withdrawn amendments, circular deference), IANA zone names around DST gaps/overlaps and the date line, business-day chains in a named zone, multi-turn corrections that flip the answer, compound packets, and aggregation over 6–40 entrant records.

About a fifth of the packets have no unique answer, and the correct output is insufficient_information. "Always abstain" can't win, because determinate and unresolved accuracy are reported separately.

The twist: every packet appears twice. The second copy is semantically identical. Only the source order, the fact order or the record order changes, or an irrelevant "About the organizer" page is added. A model that reads rules carefully should give the same answer both times.

Scoring (no LLM judge)

  • Strict accuracy: the status and value are correct (timestamps compared as instants; local-time questions must also use the requested offset), and the cited clauses equal a minimal evidence set.
  • Answer-only accuracy: the value alone, so citation strictness and reasoning errors can be told apart.
  • Presentation flip rate: the share of packet pairs where the answer (status + canonical value) changed after a harmless reorder. Set order, whitespace, offset notation and the reason label don't count as flips.
  • Confident-wrong rate: the share of answers with self-reported confidence ≥ 0.9 that were wrong.
  • 95% CIs come from a bootstrap over packet families (40 per split). Both copies of a packet are resampled together, so 80 correlated items aren't treated as 80 independent ones.

Every calendar answer is recomputed by a second, independent oracle (Python zoneinfo); every aggregation answer by a second implementation that parses the clause text itself. There are corruption tests for the scorer itself. Every model gets the same plain-text prompt and the same tolerant parser (reasoning blocks stripped, JSON pulled from code fences); rate limits and outages are retried and reported as coverage, never scored as wrong answers.

Models

Kaggle's default benchmark model (google/gemini-3.7-flash) for the leaderboard task, then the same 160 items on google/gemini-2.5-flash, google/gemma-4-31b, qwen/qwen3-235b-a22b-instruct-2507, deepseek-ai/deepseek-r1-0528 and openai/gpt-oss-120b: two generations of the same provider (does a newer model fix authority errors?), an open-weights model from the same family, and three other open-weights families including two reasoning models.

Results

Final results (run 4, 2026-10-07 PT). The default model's row is the leaderboard run (Kaggle run 3, notebook Version 1); gemini-2.5-flash is from run 2 (same items except one clarified packet, see below). Qwen3-235B, DeepSeek-R1 and Gemma-4-31B are from the run 4 exploration (results/run4/explore). Qwen and DeepSeek answered all 160 items. Gemma has 150 of 160 results: 10 hard items are missing because Kaggle's model proxy rate-limited them, so its hard score is over 70 items, not 80. gpt-oss-120b has no results at all: Kaggle's model proxy kept returning HTTP 429 (rate limit) on every attempt, so I report nothing for it rather than counting the failures as wrong answers. Numbers are strict accuracy unless noted; flip rate counts packet pairs; confident-wrong = wrong among answers with confidence ≥ 0.9. 95% CIs (brackets) were computed for the two Gemini rows only.

Model Easy strict Hard strict Hard answer-only Hard flip rate Hard confident-wrong
gemini-3.7-flash (default, leaderboard run) 0.975 [0.938, 1.0] 1.000 1.000 0.000 0.000
gemini-2.5-flash 0.975 [0.938, 1.0] 0.850 [0.775, 0.925] 0.850 0.275 0.150
gemma-4-31b (150/160; 70 hard) 0.988 0.929 0.929 0.057 (2/35) 0.071
qwen3-235b-a22b-instruct-2507 0.738 0.275 0.350 0.400 (16/40) 0.646
deepseek-r1-0528 0.900 0.650 0.688 0.200 (8/40) 0.312
gpt-oss-120b no results: every call got HTTP 429 from Kaggle's model proxy

Notes on these numbers. The default model answered all 160 items correctly (answer-only 1.000, zero flips) in the leaderboard run; its two strict misses on easy were citations written as rules:S1 instead of S1. The scorer now accepts a source-qualified ID when the prefix really is that clause's source, which makes it 160/160 — but the leaderboard number stays 0.975 + 1.000 as run. In run 2 it scored 0.988 / 1.000. For gemini-2.5-flash (run 2): one easy packet (payout_02) now also accepts citing R2 alone, and one hard packet (h_mix_03) had a judge-supplier clause that could be read as a competing permission, so I clarified its wording; on the 78 unchanged hard items it scores 0.872. The pending runs use the clarified keys.

gemini-2.5-flash per hard category: aggregation_xl 0.6, tz_dst 0.7, business_day_chains 0.8, compound 0.7; long distractors, two-hop precedence, multi-turn and 6–12-record aggregation all 1.0. Its easy split is near ceiling (0.975).

Overall strict accuracy (all items): Gemma-4-31B 0.960 (n=150), DeepSeek-R1-0528 0.775 (n=160), Qwen3-235B 0.506 (n=160).

Per-category strict accuracy, both splits combined (run 4 exploration; results/run4/fine_print_per_category.csv):

Category DeepSeek-R1 Gemma-4-31B Qwen3-235B
explicit_authority 1.0 1.0 0.7
scope_amendment 1.0 1.0 0.7
unresolved_authority 1.0 1.0 1.0
two_hop_precedence (hard) 1.0 1.0 0.6
long_distractors (hard) 1.0 1.0 0.2
multi_turn (hard) 0.7 1.0 0.8
exact_deadlines 0.3 0.9 0.6
tz_dst (hard) 0.3 0.6 0.0
business_days 1.0 1.0 0.4
business_day_chains (hard) 0.9 1.0 0.2
eligibility 0.9 1.0 0.8
submission_artifacts 1.0 1.0 1.0
notification_payout 1.0 1.0 0.7
compound (hard) 0.6 0.9 0.4
aggregation (hard) 0.5 1.0 0.0
aggregation_xl (hard) 0.2 n/a (rate-limited) 0.0

Format errors: zero for every model that returned answers. The shared parser recovered JSON from every Gemma, Qwen and DeepSeek response, so every miss above is a real wrong answer or a wrong citation, not a parsing failure. The only non-answers were gpt-oss-120b's HTTP 429s and Gemma's 10 rate-limited items, and those are counted as missing coverage, not as errors.

What surprised me

1. The 31B model beat the bigger ones. Gemma-4-31B scored 0.988 easy / 0.929 hard strict, ahead of DeepSeek-R1 (a 671B reasoning model, 0.650 hard) and Qwen3-235B (0.275 hard). It never followed an FAQ over the Rules and never broke on long distractors or two-hop precedence. Its only real weak spot was DST edge cases (0.6), which also hurt every other model. The caveat: Gemma's hard score covers 70 of 80 items, because the 10 largest aggregation items were rate-limited. Even if it had missed all 10, it would still be ahead of DeepSeek on hard.

2. Qwen is confidently wrong about which clause governs. Qwen3-235B is the only model that actually lost authority questions: 0.6 on two-hop precedence, 0.2 on long-distractor packets, 0.7 on explicit authority. On those misses it usually cited the wrong clause (for example, a Track page over the Amendment that replaced it) with confidence 0.95–1.0. On the hard split, 65% of its answers that claimed confidence ≥ 0.9 were wrong. Its self-reported confidence carries almost no information.

3. Flip rate falls as accuracy rises. On the hard split: Qwen 0.275 strict / 0.40 flips, DeepSeek-R1 0.650 / 0.20, Gemma 0.929 / 0.057, gemini-3.7-flash 1.000 / 0.000. (gemini-2.5-flash, 0.850 / 0.275, is the outlier: it flips more than its accuracy predicts.) When a model gets a hard item wrong, a harmless reorder often gives a different wrong answer. In other words, instability tells you the model doesn't know. It is a cheap and useful signal.

4. Reasoning didn't fix the calendar. DeepSeek-R1 writes long chains of thought, but it scored 0.3 on exact deadlines and DST. On several DST-gap packets it answered insufficient_information with confidence ≥ 0.95, even though the packet said exactly what to do when the local time doesn't exist. It did fine on authority (1.0 on two-hop and long distractors), so its failures are arithmetic and abstention, not reading.

The older Gemini breaks on presentation, not authority. gemini-3.7-flash, Kaggle's default model, gave the same correct answer to every packet in both orders. gemini-2.5-flash got every precedence question right, but its hard-split losses move with harmless reorders: on 11 of 40 hard packet pairs its answer changed between the two copies.

Three concrete cases (rule: the most common error type in each failing hard category, gemini-2.5-flash, run 2):

  1. Reordering a 40-record list changed the answer (aggregation_xl). Same clauses, same records, record order reversed: the base copy and the reversed copy disagreed on which IDs are disqualified in 4 of 5 packets; every one of those answers, right or wrong, had confidence ≥ 0.9.
  2. A DST gap with an explicit rule became "insufficient information" (tz_dst). "Entries close at 2:30 AM on October 4, 2026, Australia/Sydney local time. If that local time does not exist because of a daylight-saving change, the deadline is the first valid local time after it." The base copy answered insufficient_information (missing_fact, confidence 1.0); the variant, with an irrelevant page added, answered the correct 2026-10-03T16:00:00Z.
  3. Off-by-one in a Tokyo business-day chain (business_day_chains). Paperwork due three business days after first attempted notification, dates in Asia/Tokyo, with a listed holiday: the base copy got the correct 2026-11-05; the variant, with the two clauses in reverse order, said 2026-11-04 (confidence 1.0).

From run 4, using the same rule:

  1. Qwen3-235B: the wrong governing clause, with high confidence (two_hop_precedence). In h_hop_04 and h_hop_05, an Amendment overrides a Track page, which overrides the Rules. Qwen cited the Rules plus the Track clause, skipped the Amendment, and answered at confidence 0.95–1.0. On h_hop_05 the two copies even disagreed (true vs. false).
  2. DeepSeek-R1: abstaining on a DST rule it was given (tz_dst). For h_tz_01, h_tz_03 and h_tz_05, it answered insufficient_information at confidence 0.95–1.0, although the packet contained an explicit "if that local time does not exist" rule. Reasoning tokens didn't fix this. Both reasoning-heavy models I could run treated the tricky clock as missing data.
  3. Gemma-4-31B: the same failure in a smaller dose (tz_dst). It abstained on the same DST-gap packet in both copies, and on a Nepal (+05:45) packet one copy abstained while the other committed to an answer.

Gemma and DeepSeek never followed an FAQ or landing page over the Rules when a precedence sentence was present. Qwen was the only model with results that did (explicit authority 0.7). gpt-oss-120b has no results (every call got HTTP 429), so I can't say whether it fixes the arithmetic.

What it changed about how I build agents

I went in expecting authority to be the hard part, meaning whether the model knows the Rules beat the FAQ. For the good models, it isn't. Gemma and both Gemini models basically never chose the wrong source. What broke was arithmetic (DST, business days, counting records) and consistency (the same question in a different order got a different answer).

So this is how I build Fine Print now, and how I'd build any agent that reads documents with legal weight:

  • The model picks the clause. Code does the math. The LLM's job is to say "R4 governs, and here is the instant it names". Deadlines, time zones, business-day windows and record filters go through deterministic code with zoneinfo and a holiday list. Every model I tested was weakest exactly where a Python function would be perfect.
  • Run it twice, shuffled. Every extraction runs twice with the sources reordered. If the two answers disagree, it goes to a human. The flip-rate result is why: disagreement is cheap to detect, and it shows up most on the items the model actually gets wrong.
  • Don't trust self-reported confidence. Almost every answer from every model claimed ≥ 0.9. For Qwen, 65% of those were wrong on hard items. Confidence isn't the signal. Agreement and verifiable citations are.
  • Pick the model by measurement, not size. A 31B open model beat a 235B and a 671B model on this task. I would never have guessed that, and I'd have paid for the bigger one.

None of this is new advice. But I wouldn't have known which parts to move into code until I measured where the models actually fail.

What I'd measure next

  • The same packets with deterministic date tools available: does the tool fix the arithmetic and the order-sensitivity?
  • More presentation variants per packet (k > 2), so flip rate has tighter intervals
  • Real (licensed) rule pages instead of fictional ones

Where to see it

Built with AI coding agents (Codex and Grok). I designed the question and checked the gold answers. Two different agents brainstormed the categories and implemented the scorer, and a second oracle checks every calendar answer.

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

Reporting determinate and unresolved accuracy separately is the right call. "Always abstain" beats a lot of naive benchmarks.

Two statistical notes that might help when you compare models:

  1. With 80 items per split, a 95% confidence interval on accuracy is roughly ±10 percentage points around 50%. Two models 6 points apart on one split aren't really distinguishable.
  2. Because every model answers the same items, a paired test (McNemar on the items where they disagree) is much more powerful than comparing two accuracy numbers. Often only 15–20 items differ, and that's where the signal is.

Are you scoring the confidence field as well? Brier score or a calibration curve would show whether a model that's wrong on precedence at least knows it's unsure, which matters more for an agent like Fine Print than raw accuracy.