<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Hoa Nguyen</title>
    <description>The latest articles on DEV Community by Hoa Nguyen (@shenjun93).</description>
    <link>https://dev.to/shenjun93</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4165305%2Fdc8e4aea-fb4e-43ad-b45f-7ecd6fbb0cb1.png</url>
      <title>DEV Community: Hoa Nguyen</title>
      <link>https://dev.to/shenjun93</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shenjun93"/>
    <language>en</language>
    <item>
      <title>Can an AI catch the catch? I benchmarked 14 models on bounty fine print</title>
      <dc:creator>Hoa Nguyen</dc:creator>
      <pubDate>Tue, 06 Oct 2026 10:38:43 +0000</pubDate>
      <link>https://dev.to/shenjun93/can-an-ai-catch-the-catch-i-benchmarked-14-models-on-bounty-fine-print-19hb</link>
      <guid>https://dev.to/shenjun93/can-an-ai-catch-the-catch-i-benchmarked-14-models-on-bounty-fine-print-19hb</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;I'm a solo developer in Vietnam. For the last few weeks I've been hunting paid work online: bug bounties, hackathons, open-source bounties, community contests. The hardest part isn't the work. It's reading the listing.&lt;/p&gt;

&lt;p&gt;The headline says &lt;strong&gt;"free"&lt;/strong&gt;, &lt;strong&gt;"open"&lt;/strong&gt;, &lt;strong&gt;"global"&lt;/strong&gt;, &lt;strong&gt;"deadline Oct 11"&lt;/strong&gt;. Then the fine print says:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the free $5 credit only arrives &lt;em&gt;after you add a card&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;"no interviews, just ship" ... except section 9 says winners do a &lt;em&gt;live verification video call&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;the issue is open ... but it's already assigned to someone&lt;/li&gt;
&lt;li&gt;the deadline is Oct 11 &lt;em&gt;PDT&lt;/em&gt;, which is Oct 12 afternoon where I live&lt;/li&gt;
&lt;li&gt;a sponsor comment quietly moved the deadline a week earlier&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I started asking AI assistants to read listings for me, so I wanted to know: &lt;strong&gt;can a model read a listing and catch the catch?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bounty Triage&lt;/strong&gt; is 40 short listings, each with one question and one correct answer from a fixed set. Every case is fictional (names and wording changed), but most are modelled on real listings I checked while hunting. Four categories:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Answers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hidden gate&lt;/td&gt;
&lt;td&gt;Besides submitting, what extra requirement is there?&lt;/td&gt;
&lt;td&gt;none / live_interview / in_person / card_required / hired_first / own_cloud_account&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Real deadline&lt;/td&gt;
&lt;td&gt;When do submissions actually close, in my timezone (UTC+7)?&lt;/td&gt;
&lt;td&gt;multiple choice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Still open?&lt;/td&gt;
&lt;td&gt;Can I take this and expect to be the one paid?&lt;/td&gt;
&lt;td&gt;open / claimed / contested / closed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Am I eligible?&lt;/td&gt;
&lt;td&gt;Individual, 25, Vietnam, not a student&lt;/td&gt;
&lt;td&gt;eligible / not_eligible / unclear&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Design choices:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Controls.&lt;/strong&gt; About half the cases have no trap. A model that shouts "trap!" at everything loses points. Some cases are &lt;em&gt;reverse traps&lt;/em&gt;: the scary-sounding thing is optional (office hours on video, an optional city meetup, an optional card for a bonus credit).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Exact grading, no LLM judge.&lt;/strong&gt; The model answers through a structured schema; I compare the answer to the key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two tiers.&lt;/strong&gt; 25 core cases, then 15 hard ones I added after three models scored 100% on the core set.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;14 models on Kaggle Benchmarks, from Anthropic, Google, OpenAI, DeepSeek, Alibaba (Qwen) and Zhipu (GLM), from frontier models down to the cheapest tier. I picked them in pairs on purpose: a vendor's big model next to its cheap one (Opus vs Haiku, GPT-6.1 Sol vs GPT-5.4 nano, Gemini Flash vs Flash-Lite), plus open-weights models I could run myself. The question I care about is not "which model is best" but "which model is good enough to read listings for me". Scores are the share of the 40 cases answered correctly (one run each).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Overall&lt;/th&gt;
&lt;th&gt;Core (25)&lt;/th&gt;
&lt;th&gt;Hard (15)&lt;/th&gt;
&lt;th&gt;Hidden gate&lt;/th&gt;
&lt;th&gt;Deadline&lt;/th&gt;
&lt;th&gt;Still open?&lt;/th&gt;
&lt;th&gt;Eligible?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5.5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6.1 Sol&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3 Flash&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B (open weights)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5&lt;/td&gt;
&lt;td&gt;0.975&lt;/td&gt;
&lt;td&gt;0.96&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;0.90&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5&lt;/td&gt;
&lt;td&gt;0.975&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;0.93&lt;/td&gt;
&lt;td&gt;0.93&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-20b (open weights)&lt;/td&gt;
&lt;td&gt;0.975&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;0.93&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;0.90&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.1 Flash-Lite&lt;/td&gt;
&lt;td&gt;0.925&lt;/td&gt;
&lt;td&gt;0.96&lt;/td&gt;
&lt;td&gt;0.87&lt;/td&gt;
&lt;td&gt;0.93&lt;/td&gt;
&lt;td&gt;0.89&lt;/td&gt;
&lt;td&gt;0.90&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen 3 Next 80B&lt;/td&gt;
&lt;td&gt;0.875&lt;/td&gt;
&lt;td&gt;0.92&lt;/td&gt;
&lt;td&gt;0.80&lt;/td&gt;
&lt;td&gt;0.93&lt;/td&gt;
&lt;td&gt;0.78&lt;/td&gt;
&lt;td&gt;0.80&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;0.85&lt;/td&gt;
&lt;td&gt;0.88&lt;/td&gt;
&lt;td&gt;0.80&lt;/td&gt;
&lt;td&gt;0.93&lt;/td&gt;
&lt;td&gt;0.78&lt;/td&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 nano&lt;/td&gt;
&lt;td&gt;0.80&lt;/td&gt;
&lt;td&gt;0.84&lt;/td&gt;
&lt;td&gt;0.73&lt;/td&gt;
&lt;td&gt;0.86&lt;/td&gt;
&lt;td&gt;0.78&lt;/td&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;td&gt;0.86&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek R1&lt;/td&gt;
&lt;td&gt;0.775*&lt;/td&gt;
&lt;td&gt;0.80&lt;/td&gt;
&lt;td&gt;0.73&lt;/td&gt;
&lt;td&gt;0.79&lt;/td&gt;
&lt;td&gt;0.56&lt;/td&gt;
&lt;td&gt;0.80&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;* DeepSeek R1 answered 31 of the 32 cases it formatted correctly, but 8 of its 40 replies were not valid JSON, so they count as wrong. gpt-oss-120b timed out on Kaggle's model proxy twice and is left out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. This benchmark does not separate the frontier. It separates the cheap tier.&lt;/strong&gt; Six models scored 40/40. 22 of the 25 wrong answers (not counting DeepSeek's format failures, below) came from four models you would pick because they are cheap and fast: Haiku, nano, Flash-Lite and Qwen 3 Next. That is exactly the tier I would run over hundreds of listings a week, so that is where it matters. The surprise: &lt;strong&gt;Gemma 4 31B, an open-weights model, scored 40/40&lt;/strong&gt;, and gpt-oss-20b missed one. Kaggle's score-vs-cost chart puts Gemma 4 31B on the efficient frontier: 40/40 for about one cent for the whole run, while the other perfect scores cost roughly 3 to 10 cents. If I want a cheap or local model to pre-filter listings, that is the one I'd try first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. "Free" is the most effective trap.&lt;/strong&gt; The hardest case (H1, missed by 6 of 14) is a contest whose rule says your entry must run on one specific cloud's free tier, right under the line "No card needed for the free tier". Five models answered &lt;code&gt;none&lt;/code&gt;: the word &lt;em&gt;free&lt;/em&gt; closed the question for them. Qwen went the other way and answered &lt;code&gt;card_required&lt;/code&gt;, falling for the optional card mentioned on the next line. Both mistakes come from reading the loudest line instead of the rule.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. "Who already holds this?" is hard.&lt;/strong&gt; The next hardest case (H10, missed by 5 of 14): a bounty issue that is open with no assignee, but a linked pull request is approved and the maintainer wrote "merging after the release freeze". The right answer is &lt;code&gt;claimed&lt;/code&gt;, because someone is about to be paid. Models answered &lt;code&gt;contested&lt;/code&gt;, &lt;code&gt;open&lt;/code&gt; and even &lt;code&gt;closed&lt;/code&gt;. A plainer version (open issue, but assigned to someone) still fooled 3 models into saying &lt;code&gt;open&lt;/code&gt;. Status labels win over evidence for small models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Timezones fail halfway.&lt;/strong&gt; "October 11, 11:59 PM PDT" in Vietnam time is October 12, 13:59. Haiku and Flash-Lite picked 12 Oct 06:59, which is the UTC time: they converted once and stopped. The AoE deadline ("Anywhere on Earth", UTC-12) tripped 3 models. For anyone outside the US, this is the bug that makes you miss a deadline by a day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. One run is not a verdict.&lt;/strong&gt; I ran the same cases on some models twice (locally and on Kaggle). Gemini 3 Flash missed H1 in one run and got it in the other; Claude Sonnet 5 missed C1 on Kaggle and not locally. At the top, 0.975 vs 1.00 is noise. The gap between 1.00 and 0.80 is not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Format failures are real failures.&lt;/strong&gt; DeepSeek R1 was nearly perfect &lt;em&gt;when it answered in the requested format&lt;/em&gt; (31 of 32), and still came last overall, because 8 of 40 replies came back as reasoning text instead of the JSON I asked for. In an agent that reads listings for me, an answer I can't parse is the same as a wrong one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;7. A strong model found bugs in my answer key.&lt;/strong&gt; On my first local run, one model disagreed with two of my keys, and it was right: my listing text was ambiguous. I rewrote both cases. Lesson: run a benchmark on a strong model &lt;em&gt;before&lt;/em&gt; trusting it to judge the weak ones, and read every disagreement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I'd measure next:&lt;/strong&gt; listings with three or four catches at once (a timezone &lt;em&gt;and&lt;/em&gt; a track-specific extension &lt;em&gt;and&lt;/em&gt; a claim that expired), real page noise (navigation, comment threads, emoji), and repeated runs per model so the top of the table gets error bars instead of ties.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I'd tell another bounty hunter:&lt;/strong&gt; a frontier model (or Gemma 4 31B locally) reads fine print about as well as I do on these 40 cases. Cheap models do not: on the hard tier the cheap models got 2 to 4 of 15 wrong, and their mistakes lean toward "it's open, it's free, go". Use the cheap ones to sort, never to say yes.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Kaggle benchmark: &lt;a href="https://www.kaggle.com/benchmarks/shenjun93/bounty-triage" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/shenjun93/bounty-triage&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Task: &lt;a href="https://www.kaggle.com/benchmarks/tasks/shenjun93/bounty-triage" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/shenjun93/bounty-triage&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example case (hidden gate, hard tier):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;No interviews. No pitch. Just ship.&lt;/strong&gt; Shipfast Weekend, $10,000 in prizes. Build anything in 48 hours and submit a repo and a 2-minute video. Judges score online. ...&lt;br&gt;
&lt;em&gt;Terms, section 9 (Payout): before any prize is released, each winner must complete a 20-minute live verification video call with our team, camera on.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q:&lt;/strong&gt; Besides submitting your work, which ONE extra requirement applies to receiving the prize money?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expected:&lt;/strong&gt; &lt;code&gt;live_interview&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;&lt;em&gt;I built the cases and the task code with help from an AI coding assistant; the listings, the traps and the judgment about what counts as a catch come from my own hunting.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
