<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: VAIBHAV LALWANI 2241351</title>
    <description>The latest articles on DEV Community by VAIBHAV LALWANI 2241351 (@vaibhav_lalwani2241351_b).</description>
    <link>https://dev.to/vaibhav_lalwani2241351_b</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1673961%2F6c579069-4e75-4f6b-8a5c-45c1b11d4c7d.png</url>
      <title>DEV Community: VAIBHAV LALWANI 2241351</title>
      <link>https://dev.to/vaibhav_lalwani2241351_b</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vaibhav_lalwani2241351_b"/>
    <language>en</language>
    <item>
      <title>A 0-Parameter Heuristic Tied Two Frontier Models at 84.4% and Beat 11 Others. Their Mistakes Were Completely Different.</title>
      <dc:creator>VAIBHAV LALWANI 2241351</dc:creator>
      <pubDate>Fri, 09 Oct 2026 19:28:56 +0000</pubDate>
      <link>https://dev.to/vaibhav_lalwani2241351_b/a-0-parameter-heuristic-tied-two-frontier-models-at-844-and-beat-11-others-their-mistakes-were-4npa</link>
      <guid>https://dev.to/vaibhav_lalwani2241351_b/a-0-parameter-heuristic-tied-two-frontier-models-at-844-and-beat-11-others-their-mistakes-were-4npa</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the Kaggle Benchmarking Challenge.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo05bdxpdbq3xqe4tq67z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo05bdxpdbq3xqe4tq67z.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Consider two customer support tickets evaluated on Kaggle Benchmarks:&lt;/p&gt;

&lt;p&gt;In &lt;strong&gt;SCN-0013&lt;/strong&gt;, an agent evaluates ticket &lt;code&gt;TKT-5012&lt;/code&gt; requesting a refund of $65.53. Attached to the ticket is an internal note:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Assessment for TKT-5012 substantiated 2184 cents against a request of 6553 cents. Only the assessed replacement cost of the one affected unit is supported."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To get this right, an agent must read the unstructured text, extract that $21.84 was substantiated, and grant that bounded amount rather than the full request. Both OpenAI's &lt;code&gt;gpt-5.4-mini&lt;/code&gt; and Alibaba's &lt;code&gt;qwen3-next-80b&lt;/code&gt; do exactly that: they extract 2184 cents and issue a grant for $21.84. A keyword rule looking only for standard status phrases fails completely.&lt;/p&gt;

&lt;p&gt;In &lt;strong&gt;SCN-0004&lt;/strong&gt;, an agent evaluates ticket &lt;code&gt;TKT-5003&lt;/code&gt; requesting $65.55. There are no prior records attached. The system instructions state:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"The record may state that a refund has already been paid... The record may state that the claim has no supporting evidence... Otherwise grant the full requested amount."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Because no adverse record exists, the instruction is clear: grant the requested $65.55. The keyword rule follows this instruction and grants the refund. But both &lt;code&gt;gpt-5.4-mini&lt;/code&gt; and &lt;code&gt;qwen3-next-80b&lt;/code&gt; fail: seeing an empty record array, both models output &lt;code&gt;{"decision": "withhold", "amountCents": 0}&lt;/code&gt;, refusing payment despite the explicit fallback rule.&lt;/p&gt;

&lt;p&gt;Here is the central finding: on the public Kaggle leaderboard, &lt;code&gt;gpt-5.4-mini&lt;/code&gt;, &lt;code&gt;qwen3-next-80b&lt;/code&gt;, and that keyword heuristic share the exact same score: &lt;strong&gt;27 / 32 (84.4%)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;An aggregate score suggests these three systems perform identically. In reality, their failure sets share &lt;strong&gt;zero overlapping scenarios (0.0% overlap)&lt;/strong&gt;. The models solved every multi-artifact evidence problem the heuristic missed, while the heuristic solved every default rule the models over-refused.&lt;/p&gt;

&lt;p&gt;Even more surprising: that same 0-parameter heuristic &lt;strong&gt;beat 11 frontier models&lt;/strong&gt; running on Kaggle Cloud, including &lt;code&gt;grok-4.6&lt;/code&gt; (23/32), &lt;code&gt;gpt-6-sol&lt;/code&gt; (24/32), &lt;code&gt;gpt-5.5&lt;/code&gt; (25/32), and &lt;code&gt;claude-sonnet-5&lt;/code&gt; (26/32).&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;KRYT evaluates autonomous agents on operational refund decisions. In each scenario, an agent receives a ticket, customer tier metadata, and event records. It must output a strict JSON action: either granting a refund for an allowable amount under an active policy rule, or withholding payment with an appropriate reason code.&lt;/p&gt;

&lt;p&gt;The benchmark tests four practical capabilities:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Evidence synthesis:&lt;/strong&gt; Reading inspection notes and prior settlement records to determine whether money is owed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bounded numeric extraction:&lt;/strong&gt; Parsing exact dollar and cent figures out of unstructured claim text without hallucinating line items.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Policy boundary adherence:&lt;/strong&gt; Applying hard refund caps and tier limits when claimants request higher amounts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schema compliance:&lt;/strong&gt; Returning structured JSON actions matching tool schemas without breaking the execution runner.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The canonical panel (&lt;code&gt;northstar-e4-f32&lt;/code&gt;, 32 scenarios) includes non-model baselines to establish the dynamic range:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Refusal Floor (&lt;code&gt;AlwaysDeny&lt;/code&gt; / &lt;code&gt;No-Op&lt;/code&gt;):&lt;/strong&gt; &lt;strong&gt;22 / 32 (68.8%)&lt;/strong&gt;. Because 22 of the 32 scenarios expect no refund, an agent that does nothing passes more than two-thirds of the panel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keyword Rule (&lt;code&gt;OldestRecord&lt;/code&gt;):&lt;/strong&gt; &lt;strong&gt;27 / 32 (84.4%)&lt;/strong&gt;. Reads the earliest timestamped record. If it finds keywords like &lt;em&gt;"already issued"&lt;/em&gt; or &lt;em&gt;"no supporting evidence"&lt;/em&gt;, it withholds; otherwise it defaults to the published rule and grants.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mechanical Rule Follower:&lt;/strong&gt; &lt;strong&gt;23 / 32 (71.9%)&lt;/strong&gt;. Obeys published policy caps without reading evidence records.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unconstrained Grant (&lt;code&gt;AlwaysGrant&lt;/code&gt;):&lt;/strong&gt; &lt;strong&gt;0 / 32 (0.0%)&lt;/strong&gt;. Fails every scenario.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because 22 scenarios expect refusal, the active dynamic range for measuring evidence reasoning is only 10 scenarios. That high base rate makes the panel vulnerable to simple heuristics.&lt;/p&gt;




&lt;h2&gt;
  
  
  Models Tested: 41 Real Runs on Kaggle Cloud
&lt;/h2&gt;

&lt;p&gt;All models were evaluated natively on Kaggle Cloud via &lt;a href="https://www.kaggle.com/benchmarks/tasks/vaibhavlalwani21/kryt-benchmark/4" rel="noopener noreferrer"&gt;Kaggle Task &lt;code&gt;vaibhavlalwani21/kryt-benchmark/4&lt;/code&gt;&lt;/a&gt;, running at temperature 0.0 with identical inputs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Score Tier&lt;/th&gt;
&lt;th&gt;Models Evaluated&lt;/th&gt;
&lt;th&gt;Accuracy&lt;/th&gt;
&lt;th&gt;Key Behavioral Characteristic&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tier 1 (Ceiling Saturated)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;25 models:&lt;/strong&gt; Claude Haiku 4.5, Claude Haiku 5.5, Claude Sonnet 4.6, Claude Sonnet 5.5, Claude Opus 4.6, Gemini 3.7 Flash, Gemini 3.8 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.1 Pro, Gemini 3.1 Flash Lite, Gemini 3 Flash, Gemini 2.5 Flash, Gemini 2.5 Pro, Gemma 4 26B, Gemma 4 31B, GPT-5.4, GPT-5.6 Luna, GPT-5.6 Sol, GPT-6 Luna, GPT-6.1 Sol, GPT-OSS 20B, Grok 4.20 Reasoning, Grok 4.20 Non-Reasoning, GLM-5&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;32 / 32&lt;/strong&gt; (100.0%)&lt;/td&gt;
&lt;td&gt;Perfect scores; panel dynamic range saturated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tier 2 (High Reasoning)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;GPT-5.4 Nano, DeepSeek-R1 (composite)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;28 / 32&lt;/strong&gt; (87.5%)&lt;/td&gt;
&lt;td&gt;4 empty-record over-refusals (DeepSeek had 13 schema breaks)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tier 3 (Heuristic Parity)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;GPT-5.4 Mini, Qwen3-Next-80B, Gemini 3.5 Flash Lite, *OldestRecord Baseline&lt;/strong&gt;*&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;27 / 32&lt;/strong&gt; (84.4%)&lt;/td&gt;
&lt;td&gt;5 empty-record over-refusals; exactly matches 0-parameter rule&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tier 4 (Below Heuristic)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Claude Sonnet 5&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;26 / 32&lt;/strong&gt; (81.3%)&lt;/td&gt;
&lt;td&gt;Over-refusals combined with boundary edge cases&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tier 5 (Sub-Heuristic)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Claude Opus 4.7, Claude Opus 5.5, GPT-5.5&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;25 / 32&lt;/strong&gt; (78.1%)&lt;/td&gt;
&lt;td&gt;Excessive caution on ambiguous evidence records&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tier 6 (Sub-Heuristic)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Claude Opus 4.8, Claude Opus 5, GPT-5.6 Terra, GPT-6 Sol&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;24 / 32&lt;/strong&gt; (75.0%)&lt;/td&gt;
&lt;td&gt;Over-refuses on empty records and boundary caps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tier 7 (Floor Drift)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;GPT-6 Astra, Grok 4.5, Grok 4.6&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;23 / 32&lt;/strong&gt; (71.9%)&lt;/td&gt;
&lt;td&gt;Near AlwaysDeny floor; severe refusal bias&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Floor Baseline&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;AlwaysDeny Baseline (Zero-Intelligence Floor)&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;22 / 32&lt;/strong&gt; (68.8%)&lt;/td&gt;
&lt;td&gt;Withholds on all tickets; zero reasoning required&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Key Findings
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Finding 1: Why a 0-Parameter Heuristic Beat 11 Frontier Models
&lt;/h3&gt;

&lt;p&gt;A simple string-match heuristic scored 27/32, outperforming models like &lt;code&gt;grok-4.6&lt;/code&gt; (23/32), &lt;code&gt;gpt-6-sol&lt;/code&gt; (24/32), &lt;code&gt;gpt-5.5&lt;/code&gt; (25/32), and &lt;code&gt;claude-sonnet-5&lt;/code&gt; (26/32).&lt;/p&gt;

&lt;p&gt;Why did advanced frontier models lose to a heuristic?&lt;/p&gt;

&lt;p&gt;In scenarios SCN-0004 through SCN-0008, tickets have empty record arrays (&lt;code&gt;records: []&lt;/code&gt;). The published policy explicitly states:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"If no adverse record exists, grant the full requested amount."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Frontier models exhibited &lt;strong&gt;cautious over-refusal&lt;/strong&gt;: seeing an empty array, they assumed a missing audit trail indicated an unsubstantiated claim. They outputted &lt;code&gt;{"decision": "withhold", "amountCents": 0}&lt;/code&gt;, refusing payment.&lt;/p&gt;

&lt;p&gt;The naive heuristic simply executed the default grant instruction and passed. As models are aligned for caution, their refusal bias can actively degrade performance on tasks with explicit default-grant rules.&lt;/p&gt;

&lt;h3&gt;
  
  
  Finding 2: Disjoint Failures Behind Identical Scores
&lt;/h3&gt;

&lt;p&gt;When we unpack the 27/32 tie between &lt;code&gt;OldestRecord&lt;/code&gt;, &lt;code&gt;Qwen3-Next-80B&lt;/code&gt;, and &lt;code&gt;GPT-5.4 Mini&lt;/code&gt;, the failure breakdown reveals opposite behaviors:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario ID&lt;/th&gt;
&lt;th&gt;Correct Action&lt;/th&gt;
&lt;th&gt;Expected Amount&lt;/th&gt;
&lt;th&gt;OldestRecord Rule&lt;/th&gt;
&lt;th&gt;Qwen3-80B&lt;/th&gt;
&lt;th&gt;GPT-5.4 Mini&lt;/th&gt;
&lt;th&gt;Behavior&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SCN-0004&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;grant&lt;/td&gt;
&lt;td&gt;$65.55&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($65.55)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Failed ($0)&lt;/td&gt;
&lt;td&gt;Failed ($0)&lt;/td&gt;
&lt;td&gt;Models withhold when record array is empty&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SCN-0005&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;grant&lt;/td&gt;
&lt;td&gt;$264.70&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($264.70)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Failed ($0)&lt;/td&gt;
&lt;td&gt;Failed ($0)&lt;/td&gt;
&lt;td&gt;Models withhold when record array is empty&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SCN-0006&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;grant&lt;/td&gt;
&lt;td&gt;$264.71&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($264.71)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Failed ($0)&lt;/td&gt;
&lt;td&gt;Failed ($0)&lt;/td&gt;
&lt;td&gt;Models withhold when record array is empty&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SCN-0007&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;grant&lt;/td&gt;
&lt;td&gt;$229.50&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($229.50)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Failed ($0)&lt;/td&gt;
&lt;td&gt;Failed ($0)&lt;/td&gt;
&lt;td&gt;Models withhold when record array is empty&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SCN-0008&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;grant&lt;/td&gt;
&lt;td&gt;$229.51&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($229.51)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Failed ($0)&lt;/td&gt;
&lt;td&gt;Failed ($0)&lt;/td&gt;
&lt;td&gt;Models withhold when record array is empty&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SCN-0013&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;grant&lt;/td&gt;
&lt;td&gt;$21.84&lt;/td&gt;
&lt;td&gt;Failed (missed partial)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($21.84)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($21.84)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Rule cannot parse partial substantiated text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SCN-0016&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;grant&lt;/td&gt;
&lt;td&gt;$88.23&lt;/td&gt;
&lt;td&gt;Failed (missed partial)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($88.23)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($88.23)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Rule cannot parse partial substantiated text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SCN-0022&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;grant&lt;/td&gt;
&lt;td&gt;$66.17&lt;/td&gt;
&lt;td&gt;Failed (missed partial)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($66.17)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($66.17)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Rule cannot parse partial substantiated text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SCN-0025&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;grant&lt;/td&gt;
&lt;td&gt;$21.84&lt;/td&gt;
&lt;td&gt;Failed (missed partial)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($21.84)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($21.84)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Rule cannot parse partial substantiated text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SCN-0031&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;grant&lt;/td&gt;
&lt;td&gt;$22.05&lt;/td&gt;
&lt;td&gt;Failed (missed partial)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($22.05)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Correct ($22.05)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Rule cannot parse partial substantiated text&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On SCN-0013 through SCN-0031, tickets require extracting partial substantiated amounts from free text. The keyword rule cannot parse numbers from prose and fails 5 out of 5. Both Qwen and Mini read the notes and extract the exact cents correctly (5 out of 5).&lt;/p&gt;

&lt;p&gt;Conversely, on SCN-0004 through SCN-0008, tickets have no prior records, requiring agents to follow the fallback instruction to grant the requested amount. The keyword heuristic defaults to the rule and succeeds (5 out of 5). Both Qwen and Mini over-refuse, withholding payment on all five.&lt;/p&gt;

&lt;p&gt;The models and the heuristic share &lt;strong&gt;zero common failures (0.0% overlap)&lt;/strong&gt;. Standard leaderboard aggregation flattened advanced natural language comprehension and rigid rule execution into an identical 84.4% score.&lt;/p&gt;

&lt;h3&gt;
  
  
  Finding 3: Model Output vs. System Reward (DeepSeek-R1)
&lt;/h3&gt;

&lt;p&gt;DeepSeek-R1 recorded a 28/32 reward on Kaggle, placing it in Tier 2. However, inspecting raw completion logs revealed that this number does not reflect schema-valid model actions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;In &lt;strong&gt;13 of 32 scenarios&lt;/strong&gt;, DeepSeek returned raw &lt;code&gt;&amp;lt;think&amp;gt;...&amp;lt;/think&amp;gt;&lt;/code&gt; tokens preceding its JSON payload, breaking schema parsing.&lt;/li&gt;
&lt;li&gt;The benchmark runner caught these parse failures and applied a default fallback action: withholding payment.&lt;/li&gt;
&lt;li&gt;Because 22 of 32 scenarios expected withholding, defaulting to withhold happened to match the correct answer on &lt;strong&gt;9 scenarios&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;On the remaining 4 malformed completions where a grant was required, the fallback withheld and failed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Kaggle assertion telemetry captured the breakdown: the assertion logs recorded 19 passes and 13 failed assertions (&lt;code&gt;assert_fail&lt;/code&gt;). The model's schema-valid output accuracy was &lt;strong&gt;19 / 32 (59.4%)&lt;/strong&gt;, while composite system reward was &lt;strong&gt;28 / 32 (87.5%)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Without decomposing output compliance from fallback recovery, the leaderboard credited DeepSeek with 9 points it never earned through reasoning.&lt;/p&gt;

&lt;h3&gt;
  
  
  Finding 4: Scaling Model Size Did Not Fix Refusal Bias
&lt;/h3&gt;

&lt;p&gt;You might expect larger flagship models to handle edge cases better than their smaller distillation counterparts. On this benchmark, the reverse happened.&lt;/p&gt;

&lt;p&gt;OpenAI's &lt;code&gt;gpt-5.4-nano&lt;/code&gt; scored &lt;strong&gt;28 / 32 (87.5%)&lt;/strong&gt;, losing only four scenarios to empty-record confusion. Meanwhile, heavyweight frontier models drifted downward: &lt;code&gt;claude-sonnet-5&lt;/code&gt; scored 26/32, &lt;code&gt;gpt-5.5&lt;/code&gt; scored 25/32, &lt;code&gt;gpt-6-sol&lt;/code&gt; scored 24/32, and &lt;code&gt;gpt-6-astra&lt;/code&gt; alongside &lt;code&gt;grok-4.6&lt;/code&gt; scored 23/32.&lt;/p&gt;

&lt;p&gt;A score of 23/32 is just a single correct decision away from the 22/32 &lt;code&gt;AlwaysDeny&lt;/code&gt; baseline. Why did the largest models degrade toward the zero-intelligence floor?&lt;/p&gt;

&lt;p&gt;Inspecting their reasoning traces reveals that safety and liability post-training penalizes false grants far more heavily than false refusals. When faced with missing records or slightly ambiguous wording, larger models default to withholding payout as the safest option. The smaller models, carrying lighter refusal conditioning, followed the operational prompt more literally and granted when instructed to do so.&lt;/p&gt;




&lt;h2&gt;
  
  
  Resolving Ceiling Saturation: The ACT Counterfactual Panel
&lt;/h2&gt;

&lt;p&gt;The 32-scenario panel exposed two structural vulnerabilities:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;High Refusal Floor (68.8%):&lt;/strong&gt; A constant do-nothing baseline scores 22/32, allowing a simple keyword rule to tie frontier models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ceiling Saturation:&lt;/strong&gt; 25 models achieved 32/32 (100%), leaving no headroom to benchmark differences among the top tier.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;To solve this, we engineered and validated &lt;strong&gt;ACT (Adversarial Counterfactual Test, &lt;code&gt;act-v1&lt;/code&gt;)&lt;/strong&gt;, shipping directly alongside the benchmark:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Expanded Sample Size ($N = 60$):&lt;/strong&gt; 60 scenarios structured into 30 matched counterfactual pairs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enforced 50.0% Floor:&lt;/strong&gt; Exactly 30 grants and 30 withholds. The refusal floor (&lt;code&gt;AlwaysDeny&lt;/code&gt;) drops from 68.8% to 50.0%. No constant policy (AlwaysDeny, AlwaysGrant, AlwaysEscalate) can score above 50.0%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Matched Counterfactual Pairs:&lt;/strong&gt; Within each pair, scenarios share identical customer history and ticket text, but one decisive evidentiary detail is flipped (such as substantiated damage notes versus a retracted claim). A keyword rule cannot rely on superficial word matches because both scenarios contain the same keywords.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Eliminating Shortcut Exploitation:&lt;/strong&gt; When evaluated on ACT, models that previously scored high on Northstar by exploiting base-rate refusal skew fell to the 50.0% floor.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The full 60-scenario ACT dataset (&lt;code&gt;panels/act-v1.tasks.json&lt;/code&gt;, &lt;code&gt;panels/act-v1.expected.json&lt;/code&gt;) and evaluation runner (&lt;code&gt;score.py --panel act-v1&lt;/code&gt;) are included in the open release package.&lt;/p&gt;




&lt;h2&gt;
  
  
  Artifacts and Reproducibility
&lt;/h2&gt;

&lt;p&gt;The task, dataset, scenarios, and verification code are published and accessible on Kaggle:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Kaggle Benchmarks Task:&lt;/strong&gt; &lt;a href="https://www.kaggle.com/benchmarks/tasks/vaibhavlalwani21/kryt-benchmark/4" rel="noopener noreferrer"&gt;kaggle.com/benchmarks/tasks/vaibhavlalwani21/kryt-benchmark/4&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kaggle Dataset &amp;amp; Scenarios:&lt;/strong&gt; &lt;a href="https://www.kaggle.com/datasets/vaibhavlalwani21/kryt-northstar-panel" rel="noopener noreferrer"&gt;kaggle.com/datasets/vaibhavlalwani21/kryt-northstar-panel&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Release Manifest Verification:&lt;/strong&gt; 71/71 integrity checks pass (&lt;code&gt;node release/kryt-kaggle-rc/verify-manifest.mjs&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Secret Leak Scan:&lt;/strong&gt; 1002 tracked repository files scanned with 0 leaks (&lt;code&gt;python tools/verify/scan_secrets.py&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Takeaways for Benchmark Authors
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Benchmark degenerate heuristics alongside models:&lt;/strong&gt; If a simple string-match heuristic scores 84%, your aggregate leaderboard will conceal whether models are genuinely reasoning or exploiting base-rate skew.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decompose raw model compliance from runner fallbacks:&lt;/strong&gt; When models emit malformed outputs, default error handling can inflate scores if the default action matches the panel's dominant class.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inspect paired error overlap:&lt;/strong&gt; Two models with identical accuracy numbers can have entirely non-overlapping failure modes. Aggregate scalar scores hide where intelligence actually breaks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Balance your test panel before claiming dynamic range:&lt;/strong&gt; A benchmark with a 68.8% refusal base rate only has 10 active scenarios to evaluate real reasoning. Matched counterfactual pairs prevent trivial heuristics from masquerading as intelligence.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
