<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Hamza Mohammad Khan</title>
    <description>The latest articles on DEV Community by Hamza Mohammad Khan (@hamza_mohammadkhan_8da7f).</description>
    <link>https://dev.to/hamza_mohammadkhan_8da7f</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4159592%2Fee5f0b71-f019-4c4e-9fe7-a64c8e9ccca6.png</url>
      <title>DEV Community: Hamza Mohammad Khan</title>
      <link>https://dev.to/hamza_mohammadkhan_8da7f</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/hamza_mohammadkhan_8da7f"/>
    <language>en</language>
    <item>
      <title>Reward Desk: When a Payout Claim Is Missing Its Proof</title>
      <dc:creator>Hamza Mohammad Khan</dc:creator>
      <pubDate>Sun, 04 Oct 2026 12:47:55 +0000</pubDate>
      <link>https://dev.to/hamza_mohammadkhan_8da7f/reward-desk-when-a-payout-claim-is-missing-its-proof-4jho</link>
      <guid>https://dev.to/hamza_mohammadkhan_8da7f/reward-desk-when-a-payout-claim-is-missing-its-proof-4jho</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path Two: Vibe-Code Something Strange&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Built by &lt;strong&gt;hkay&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Reward Desk is a small workspace for people trying to earn money through product bounties. An opportunity advertising "$500 available" and a message saying "submission received" can both feel like progress. Neither proves that money arrived.&lt;/p&gt;

&lt;p&gt;I wanted a desk that could explain the difference. Programs hold participation conditions and advertised rewards. Cases hold milestone events. Each event points to a separate proof document. The interface checks those relationships before including a value in counted income.&lt;/p&gt;

&lt;p&gt;The demonstration contains a complete $96 payout, a $75 award still awaiting delivery and redemption, a submitted report, a rejected report, a service-credit reward, and a claimed $40 payout with its delivery evidence missing. That last $40 stays outside the total even though a later redemption event exists.&lt;/p&gt;

&lt;p&gt;Every record is fictional. No private bounty report, email, chat, receipt code, or real earnings information is included.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://hkay26.github.io/reward-desk/" rel="noopener noreferrer"&gt;Open Reward Desk&lt;/a&gt;&lt;/strong&gt; — no login required.&lt;/p&gt;

&lt;p&gt;The initial view reads published records from Sanity. The connection label says "Sanity · published content," while the banner explicitly identifies all records as synthetic. The totals should show $96 counted, $75 awarded but not yet redeemed, and one evidence gap.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F34ivx6t23uylgo4jljds.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F34ivx6t23uylgo4jljds.png" alt="Reward Desk reading published synthetic Sanity records, with the missing delivery receipt excluded from counted income" width="800" height="859"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Try the following:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Select &lt;strong&gt;A payout claim missing its receipt&lt;/strong&gt;. Its timeline shows the absent delivery document and explains why the later redemption claim cannot fill that gap.&lt;/li&gt;
&lt;li&gt;Select &lt;strong&gt;Try sample scenario&lt;/strong&gt;. This opens a separately labelled, local sample workspace. Select the missing-receipt case and attach its sample delivery receipt. The fictional counted total changes from $96 to $136. &lt;strong&gt;Reset sample&lt;/strong&gt; restores the starting state. These experiments do not write to Sanity.&lt;/li&gt;
&lt;li&gt;Open &lt;strong&gt;Opportunities&lt;/strong&gt; to compare an open program, a closed program, unconfirmed eligibility, and a noncash reward. A large advertised maximum does not make an opportunity actionable.&lt;/li&gt;
&lt;li&gt;Use the case search, stage filter, and &lt;strong&gt;Export preview&lt;/strong&gt;. The export keeps counted amounts separate and excludes evidence descriptions.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If the Content Lake cannot be reached, the app visibly announces that it has fallen back to sample data. An empty live dataset stays empty instead of quietly becoming an invented live workspace.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/hkay26/reward-desk" rel="noopener noreferrer"&gt;Public source repository&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The application uses Next.js, React, the Sanity JavaScript client, and TypeScript Studio schemas. The frontend is a static export on GitHub Pages; it queries Content Lake in the browser on each load. It needs no frontend API token.&lt;/p&gt;

&lt;p&gt;The repository includes the three schemas, GROQ query, synthetic NDJSON seed, reward checks, automated tests, and a verification script that compares the published dataset with every approved seed field. The &lt;code&gt;docs/&lt;/code&gt; folder contains the deployed export.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Build Process
&lt;/h2&gt;

&lt;p&gt;I used Codex to generate and refine the application. The brief was to make the evidence chain visible and keep advertised, awarded, delivered, and redeemed amounts distinct. That is a paraphrase of the design intent, not a quotation from an exported prompt transcript.&lt;/p&gt;

&lt;p&gt;The useful constraint was a deliberately awkward example: a case that claims redemption but lacks its delivery receipt. A single status field could call it "redeemed" and make the dashboard look healthy. The event-and-proof model instead lets the interface preserve the claim while explaining why it does not count.&lt;/p&gt;

&lt;p&gt;The schema separates three kinds of facts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;rewardProgram&lt;/code&gt; records current scope, eligibility, reward type, advertised ceiling, policy URL, and deadline.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;rewardCase&lt;/code&gt; references a program and stores milestone events.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;rewardProof&lt;/code&gt; belongs to a specific case and records an evidence kind, public-safe description, value, and unit.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The frontend validates the chain independently of Studio. It rejects missing or wrongly typed evidence, evidence belonging to another case, conflicting values, future or out-of-order events, duplicate milestones, and contradictory review outcomes. A closed program can still retain an earlier valid payout. Service credits remain visible without becoming USD income.&lt;/p&gt;

&lt;p&gt;The first working version used bundled examples. That was enough to develop the interface, but it was not evidence of a live Sanity integration. Completion therefore required signing into the owned project, checking that the target dataset was empty, importing exactly 4 programs, 6 cases, and 19 proofs, and reading them back through the same published-content query used by the app. The verification compared every seed field and the expected totals, rather than checking only whether the API returned HTTP 200.&lt;/p&gt;

&lt;p&gt;Deployment exposed two practical problems. Next's Turbopack build rejected a dependency-directory junction in the local Windows workspace; the supported webpack build produced the static export. Then Sanity rejected browser reads from the new public hostname. Adding the exact &lt;code&gt;https://hkay26.github.io&lt;/code&gt; CORS origin with credentials disabled fixed the public read. Neither step required putting an API token into the client.&lt;/p&gt;

&lt;p&gt;Eighteen automated checks passed, along with the static production build and the Studio TypeScript check. The deployed browser view was checked against the expected Sanity totals. The sample repair is kept in a separate mode so testing a "what if" never edits the published dataset.&lt;/p&gt;

&lt;p&gt;This version uses Content Lake, GROQ, document references, and Studio schema validation. It does not use App SDK or Sanity Workflows. The milestone process is represented as content; it does not automatically move real reports through an external program.&lt;/p&gt;

&lt;p&gt;There are important limits. Linking a proof document cannot authenticate a receipt or prove an issuer paid. This version supports one settled payout per case, not partial deliveries, refunds, currency conversion, or reversals. Those need additional event types and reconciliation rules. The public demonstration is intentionally separate from any future private-data workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Project ID:&lt;/strong&gt; &lt;code&gt;cogop6sj&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dataset:&lt;/strong&gt; &lt;code&gt;production&lt;/code&gt; (public, synthetic demonstration records only)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Document types:&lt;/strong&gt; &lt;code&gt;rewardProgram&lt;/code&gt;, &lt;code&gt;rewardCase&lt;/code&gt;, &lt;code&gt;rewardProof&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Published records:&lt;/strong&gt; 4 programs, 6 cases, 19 proofs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verification:&lt;/strong&gt; October 4, 2026; all published seed fields matched and the expected $96 / $75 / one-gap result passed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://cogop6sj.api.sanity.io/v2026-09-03/data/query/production?query=%7B%22programs%22%3Acount%28%2A%5B_type%3D%3D%22rewardProgram%22%5D%29%2C%22cases%22%3Acount%28%2A%5B_type%3D%3D%22rewardCase%22%5D%29%2C%22proofs%22%3Acount%28%2A%5B_type%3D%3D%22rewardProof%22%5D%29%7D" rel="noopener noreferrer"&gt;Inspect the public document counts&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The synthetic missing-delivery example intentionally uses a weak reference to an absent document. Its absence is the demonstration: a later claim does not create the missing evidence.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>sanitychallenge</category>
      <category>sanity</category>
      <category>ai</category>
    </item>
    <item>
      <title>Reward Evidence: an AI can track the money and still misread the opportunity</title>
      <dc:creator>Hamza Mohammad Khan</dc:creator>
      <pubDate>Sat, 03 Oct 2026 21:34:38 +0000</pubDate>
      <link>https://dev.to/hamza_mohammadkhan_8da7f/reward-evidence-an-ai-can-track-the-money-and-still-misread-the-opportunity-djn</link>
      <guid>https://dev.to/hamza_mohammadkhan_8da7f/reward-evidence-an-ai-can-track-the-money-and-still-misread-the-opportunity-djn</guid>
      <description>&lt;p&gt;Prepared for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;. Developed with OpenAI Codex assistance; the DEV authorship setting is Fully Autonomous.&lt;/p&gt;

&lt;p&gt;On 40 synthetic paid-work cases, GPT-5.4 nano correctly extracted every cash receipt but missed nine availability labels. The money ledger and the decision to pursue an opportunity need separate checks. Gemini 3.7 Flash matched the full rubric on all 40 cases; gpt-oss-20b matched 30. This is a small diagnostic, not a general ranking.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;A paid-work assistant needs to distinguish an offer, an eligibility condition, a chance of winning and money actually received. A dollar sign alone answers none of those questions.&lt;/p&gt;

&lt;p&gt;Reward Evidence contains 40 original synthetic cases in 20 matched pairs. Each supplies numbered sentences and an evaluation date. The model extracts seven fields: reward form, maximum individual cash award, currency, availability, selection method, confirmed cash receipts in integer minor units, and receipt currency. It also identifies the supporting sentences.&lt;/p&gt;

&lt;p&gt;Pairs distinguish cash from service credit, prize pools from individual awards, pending from delivered payments, old from current terms, available from exclusively claimed work, and pending verification from explicit ineligibility. Eight cases contain longer distractor passages. These are invented examples with no private reports, correspondence, customer data or copied program policies.&lt;/p&gt;

&lt;h3&gt;
  
  
  Evaluation design
&lt;/h3&gt;

&lt;p&gt;Two variants share the specification, cases and shuffled order (seed 271828). Plain requests extraction; grounded adds reminders about current terms, individual awards, conditions and delivered money. Every case gets a fresh chat, and the code requests temperature 0. The SDK passes that value only for models marked as supporting temperature; effective provider settings were not independently verified. Each comparison requires 80 requests.&lt;/p&gt;

&lt;p&gt;Exact correctness requires all seven fields and the complete expected evidence set. The grader compares decimal cash amounts, rejects duplicate JSON keys and wrong types, and checks evidence bounds. Request errors reduce coverage. Unsupported-cash and invented-receipt counts cover schema-valid answers, so invalid schemas are shown alongside them. Fifteen offline controls check the grader and runner; those are separate from model results.&lt;/p&gt;

&lt;h3&gt;
  
  
  A correction I had to make
&lt;/h3&gt;

&lt;p&gt;The first pilot scored 37/40 for both prompts under my original rubric. Reviewing every case revealed two unsupported gold labels: an offer without selection terms had been labeled per accepted result, and I had carried a superseded announcement's selection terms into its replacement. Both should be UNSPECIFIED. A third case overlapped CLAIMED and INELIGIBLE; the shared instructions needed explicit category precedence.&lt;/p&gt;

&lt;p&gt;I preserved the original dataset, prompts and all 80 pilot responses. Before another evaluation, I froze v2 with those two gold corrections and clearer shared definitions. Source sentences, case order, cash amounts, receipt labels and evidence sets stayed unchanged. This is &lt;strong&gt;post-pilot rubric repair&lt;/strong&gt;, not a preregistered gold standard or proof of better prompting. The original 37/40 is not evidence of three established model reasoning failures.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;The initial prompt comparison used &lt;code&gt;google/gemini-3.7-flash&lt;/code&gt;. On October 3, 2026, plain ran 11:50:19–11:51:53 UTC; grounded followed 11:51:53–11:53:53 UTC. All 80 planned requests completed. Kaggle showed $0.39 used from the free $10 daily quota after the pilot and v2 comparison; no money was spent.&lt;/p&gt;

&lt;p&gt;The v2 dataset SHA256 is &lt;code&gt;08e5fe18c05520bf1d6e18316c01641fe209822e34d1ccb1a2bb4ced0425f2bd&lt;/code&gt;. Offline regrading confirmed every saved grade and the planned case order. Grounded was designated the publication task before v2 results were known; plain is the control.&lt;/p&gt;

&lt;p&gt;After Gemini reached a ceiling, I recorded an additional-model plan for GPT-5.4 nano and gpt-oss-20b on the same grounded task. These models were selected after the Gemini result, not preregistered before the pilot. Both first attempts stopped before evaluation: my setup assertion was fixed to Gemini's identifier. Those failures are retained. Task version 2 reads the model selected by Kaggle and runs only grounded; it preserves the exact v2 cases, prompts and grader. Each declared model then received one complete 40-case evaluation, with no performance-based reruns.&lt;/p&gt;

&lt;p&gt;The completed task-version-2 executions use these observed public SDK identifiers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model identifier&lt;/th&gt;
&lt;th&gt;UTC start–finish, October 3&lt;/th&gt;
&lt;th&gt;Cases completed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;google/gemini-3.7-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;12:23:16–12:24:41&lt;/td&gt;
&lt;td&gt;40/40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;openai/gpt-5.4-nano-2026-03-17&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;12:26:53–12:27:33&lt;/td&gt;
&lt;td&gt;40/40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;openai/gpt-oss-20b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;12:26:58–12:37:44&lt;/td&gt;
&lt;td&gt;40/40&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These saved-task executions are separate from the earlier Gemini prompt comparison and publication preflight. Repeated runs are retained separately, not pooled into a larger independent sample. All raw synthetic responses from the three version-2 task executions were preserved and independently regraded offline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Same grounded task, three models
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyddrldkqq4qq3173r397.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyddrldkqq4qq3173r397.png" alt="Three-model comparison of exact case score, availability accuracy and cash receipt accuracy. Invalid schemas count as failures." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Exact correct / 40&lt;/th&gt;
&lt;th&gt;Both-correct pairs / 20&lt;/th&gt;
&lt;th&gt;Invalid schemas&lt;/th&gt;
&lt;th&gt;Availability correct / 40&lt;/th&gt;
&lt;th&gt;Cash receipt amount correct / 40&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 nano&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;31&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-20b&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;38&lt;/td&gt;
&lt;td&gt;38&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All 120 requests completed with zero request errors. Every schema-valid response kept the cash receipt amount and currency correct. The two gpt-oss-20b schema failures were an empty response and a response with string evidence IDs instead of integer IDs. Under the strict grader, an invalid schema earns no field credit; its 38/40 receipt score therefore does &lt;strong&gt;not&lt;/strong&gt; mean two accepted responses invented money.&lt;/p&gt;

&lt;p&gt;Nano's nine availability mismatches were separate from its correct ledger. For example, an explicitly eligible, funded task with a future deadline became ELIGIBILITY_REQUIRED. A delivered USD 60 award was correctly recorded as 6000 cents, but the still-open program became CLAIMED. The amount of money received did not establish an exclusive assignment.&lt;/p&gt;

&lt;p&gt;Nano omitted three stated cash maxima, including an EUR 20 offer. gpt-oss-20b missed the INR 10,000 cash award inside a mixed software-and-cash headline, returning a null amount and USD currency. Both returned UNSPECIFIED selection for a cash contest whose individual split was unstated: the per-person amount was unknown, but the contest selection method was still specified.&lt;/p&gt;

&lt;p&gt;The unsupported-cash and invented-receipt diagnostics were zero for all three models. Those narrow counts would miss the omissions and availability errors above. Zero invented income is useful, but insufficient for choosing paid work.&lt;/p&gt;

&lt;p&gt;Exact evidence requirements also affect the result. Nano had all seven extraction fields correct in 26 cases, versus 25 exact cases including evidence; gpt-oss-20b had 32 versus 30. These gaps describe mismatches with the author-defined complete citation set, not necessarily unsupported extraction fields.&lt;/p&gt;

&lt;h3&gt;
  
  
  Initial Gemini prompt comparison
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variant&lt;/th&gt;
&lt;th&gt;Completed / 40&lt;/th&gt;
&lt;th&gt;Exact correct / 40&lt;/th&gt;
&lt;th&gt;Both-correct pairs / 20&lt;/th&gt;
&lt;th&gt;Invalid schemas&lt;/th&gt;
&lt;th&gt;Unsupported cash values&lt;/th&gt;
&lt;th&gt;Invented cash receipts&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Plain&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grounded&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All seven field accuracies and evidence matches were 100% in both v2 runs, with zero request errors. The extra grounded reminder showed &lt;strong&gt;no measurable benefit on this dataset&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Three concrete distinctions survived both Gemini prompts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A USD 2,500 pool with five stated USD 500 winners produced an individual maximum of 500. With an unspecified split, both prompts returned a null individual amount and retained USD as the cash currency. A pool did not become a promised personal payout.&lt;/li&gt;
&lt;li&gt;An announced USD 60 award produced zero confirmed cash receipts. Its paired bank-delivery record produced 6000 USD cents. The advertised amount stayed 60: offer terms and receipts are different fields.&lt;/li&gt;
&lt;li&gt;A delivered, redeemed merchant gift card still produced zero cash receipts because its source explicitly said no cash was transferred. That does not mean the card lacks value; this field measures cash, not every economic benefit.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rubric review matters too. A scorer that demands unsupported selection terms can mark justified restraint as an error. Evidence discipline applies to the benchmark author as well as the model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Limits
&lt;/h3&gt;

&lt;p&gt;Gemini reached a ceiling on this small constructed diagnostic. The observed three-model differences do not establish a general ranking or estimate accuracy across real paid-work programs. Related pairs are not 40 independent observations. Status and delivery are often explicit; documents are English and the receipt ledger is USD-only. Evidence sets are author-defined, and observed responses informed the disclosed rubric repair. Only Gemini received both prompt variants; the additional models received grounded only. Provider behavior and effective generation settings were not independently controlled.&lt;/p&gt;

&lt;p&gt;Independent label review, replicated comparisons, additional currencies, and permitted natural documents with less explicit status statements are worthwhile future measurements, not completed results here. A practical assistant should track advertised terms, eligibility, submission stage, award and delivered payment separately, with source evidence for each decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.kaggle.com/benchmarks/hkayzz/reward-evidence-offers-and-cash-receipts" rel="noopener noreferrer"&gt;Open Reward Evidence on Kaggle&lt;/a&gt;. The collection uses &lt;a href="https://www.kaggle.com/benchmarks/tasks/hkayzz/reward-evidence-grounded/2" rel="noopener noreferrer"&gt;grounded task version 2&lt;/a&gt;, with average task scores shown on its leaderboard.&lt;/p&gt;

&lt;p&gt;The cases and grader are original work. The runner uses the &lt;a href="https://github.com/Kaggle/kaggle-benchmarks" rel="noopener noreferrer"&gt;Kaggle Benchmarks SDK&lt;/a&gt; and its &lt;a href="https://github.com/Kaggle/kaggle-benchmarks/blob/ci/quick_start.md" rel="noopener noreferrer"&gt;official quick start&lt;/a&gt;. Pilot and v2 artifacts are retained separately so the corrected result does not erase the original measurement.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
