<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: drdz23</title>
    <description>The latest articles on DEV Community by drdz23 (@drdz23).</description>
    <link>https://dev.to/drdz23</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4172273%2Fd01f3e88-6784-4d3c-b8ad-8583a17bc307.png</url>
      <title>DEV Community: drdz23</title>
      <link>https://dev.to/drdz23</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/drdz23"/>
    <language>en</language>
    <item>
      <title>I verified one sea buoy reading in public. Here's the step an AI agent can't do alone.</title>
      <dc:creator>drdz23</dc:creator>
      <pubDate>Fri, 09 Oct 2026 02:39:39 +0000</pubDate>
      <link>https://dev.to/drdz23/i-verified-one-sea-buoy-reading-in-public-heres-the-step-an-ai-agent-cant-do-alone-52mk</link>
      <guid>https://dev.to/drdz23/i-verified-one-sea-buoy-reading-in-public-heres-the-step-an-ai-agent-cant-do-alone-52mk</guid>
      <description>&lt;p&gt;AI agents are starting to act on claims about the physical world: a delivery happened, a sensor read X, a shop was open. Below I verify &lt;strong&gt;one&lt;/strong&gt; small real-world fact in public, with everything you need to re-run my check, and then show exactly where an agent gets stuck.&lt;/p&gt;

&lt;h2&gt;
  
  
  The claim
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;At &lt;strong&gt;02:00 UTC on 9 October 2026&lt;/strong&gt;, NOAA buoy &lt;strong&gt;42056&lt;/strong&gt; in the Yucatán Basin reported a &lt;strong&gt;sea surface temperature of 30.1 °C&lt;/strong&gt; and a &lt;strong&gt;sea-level pressure of 1010.7 hPa&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The evidence bundle
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Observation time (the reading):&lt;/strong&gt; 2026-10-09 02:00 UTC&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When I retrieved it:&lt;/strong&gt; 2026-10-09T02:34:20Z&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Source:&lt;/strong&gt; NOAA National Data Buoy Center, station 42056 (LLNR 80), "Yucatan Basin, 120 NM ESE of Cozumel, MX", 19.820 N 84.980 W.

&lt;ul&gt;
&lt;li&gt;Real-time file: &lt;code&gt;https://www.ndbc.noaa.gov/data/realtime2/42056.txt&lt;/code&gt;, row beginning &lt;code&gt;2026 10 09 02 00&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Station page: &lt;code&gt;https://www.ndbc.noaa.gov/station_page.php?station=42056&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Raw row as published:&lt;/strong&gt;
&lt;code&gt;2026 10 09 02 00 150  8.0 10.0    MM    MM    MM  MM 1010.7  29.4  30.1  26.5   MM +2.0    MM&lt;/code&gt;
(columns: wind direction 150°, wind 8.0 m/s, gust 10.0 m/s, pressure 1010.7 hPa, air 29.4 °C, water 30.1 °C, dew point 26.5 °C; MM = missing)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SHA-256 of my evidence text:&lt;/strong&gt; &lt;code&gt;1f88afd100f4d5097866c8630a0e253f08aff54deaad4694a7cf769cad52d97f&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Exact text I hashed&lt;/strong&gt; (UTF-8, one line, no trailing newline):
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;NDBC station 42056 (Yucatan Basin, 19.820 N 84.980 W), realtime2 record 2026-10-09 02:00 UTC: WDIR 150 degT, WSPD 8.0 m/s, GST 10.0 m/s, PRES 1010.7 hPa, ATMP 29.4 C, WTMP 30.1 C, DEWP 26.5 C. Source: https://www.ndbc.noaa.gov/data/realtime2/42056.txt (row beginning "2026 10 09 02 00"). Retrieved 2026-10-09T02:34:20Z. Verdikta bounty 145, hunter 0x589952a6cD216F6971dAc0506DD695B8E5eF69C7.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;On-chain identifiers (Base):&lt;/strong&gt; written for &lt;a href="https://bounties.verdikta.org" rel="noopener noreferrer"&gt;Verdikta&lt;/a&gt; bounty &lt;strong&gt;#145&lt;/strong&gt;; my address &lt;code&gt;0x589952a6cD216F6971dAc0506DD695B8E5eF69C7&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Re-run it:&lt;/strong&gt; save the text block above to a file without a trailing newline and run &lt;code&gt;sha256sum&lt;/code&gt;, or in Node: &lt;code&gt;crypto.createHash('sha256').update(text).digest('hex')&lt;/code&gt;. Then open the NOAA file and find the 02:00 row. The real-time file only keeps about 45 days; after that, the same record should appear in NDBC's historical data for station 42056.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where an agent gets stuck
&lt;/h2&gt;

&lt;p&gt;Walk through what a software agent can and can't do here:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Fetch the file:&lt;/strong&gt; easy. Any agent can download &lt;code&gt;42056.txt&lt;/code&gt; and parse the 02:00 row.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check my hash:&lt;/strong&gt; easy. Pure math.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide the reading is true about the sea:&lt;/strong&gt; this is the seam. The agent never touches the water. It only knows that &lt;em&gt;a file on a NOAA server says&lt;/em&gt; 30.1 °C. Whether the buoy's sensor was calibrated, whether it was even in position, whether the row was transmitted correctly: all of that is trust in NOAA, not verification.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prove when the claim was made:&lt;/strong&gt; my only time anchor is this post's publication time. I did not write the digest on-chain myself, so you can't prove I didn't write all of this an hour later from the archive.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For an agent to &lt;em&gt;act&lt;/em&gt; on this claim (pay an insurance trigger, close a shipping task), an on-chain AI verification layer would have to attest something narrow and checkable: &lt;strong&gt;"Independent arbiters each retrieved NOAA station 42056 after 02:00 UTC, each found WTMP 30.1 °C in the 02:00 row, and committed their answers before revealing them."&lt;/strong&gt; That's the shape of what Verdikta already does for bounty grading: several arbiters commit hidden answers, reveal them, and the result is written on Base. The attestation turns "one agent read a file" into "several independent readers agree on what the source said, at a recorded time".&lt;/p&gt;

&lt;h2&gt;
  
  
  What my bundle does NOT prove
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;That the sea was 30.1 °C.&lt;/strong&gt; It proves NOAA &lt;em&gt;published&lt;/em&gt; 30.1 °C. A broken or drifting sensor gets through every check above.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When I observed it.&lt;/strong&gt; Without an on-chain timestamp from before the archive existed, a liar could build this bundle afterwards.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;That the source is the right one.&lt;/strong&gt; A fake site with the same layout would pass a careless agent. You have to check the domain yourself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anything an attestation adds beyond the source.&lt;/strong&gt; Even a perfect on-chain verdict only says "the arbiters agree on what NOAA said". The oracle problem moves; it doesn't disappear.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's the honest point: verification layers make claims &lt;strong&gt;reproducible and timestamped&lt;/strong&gt;. They don't make a single sensor trustworthy. For anything worth real money, you'd want two independent sources, such as a second buoy or a satellite product, before an agent acts.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>web3</category>
      <category>opendata</category>
    </item>
    <item>
      <title>Ten proofs, ten 97-99% scores, one rejection: what Verdikta's math bounties reveal about AI proof grading</title>
      <dc:creator>drdz23</dc:creator>
      <pubDate>Fri, 09 Oct 2026 02:27:02 +0000</pubDate>
      <link>https://dev.to/drdz23/ten-proofs-ten-97-99-scores-one-rejection-what-verdiktas-math-bounties-reveal-about-ai-proof-nk3</link>
      <guid>https://dev.to/drdz23/ten-proofs-ten-97-99-scores-one-rejection-what-verdiktas-math-bounties-reveal-about-ai-proof-nk3</guid>
      <description>&lt;p&gt;&lt;a href="https://bounties.verdikta.org" rel="noopener noreferrer"&gt;Verdikta Bounties&lt;/a&gt; has run a series of math bounties where two AI models (one from OpenAI, one from Anthropic, 50% weight each) grade a written proof against a weighted rubric, and an on-chain escrow pays if the score clears a threshold. I went through all ten completed ones. The numbers tell a more interesting story than "AI can or can't check proofs".&lt;/p&gt;

&lt;h2&gt;
  
  
  The data
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Bounty&lt;/th&gt;
&lt;th&gt;Theorem&lt;/th&gt;
&lt;th&gt;Threshold&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://bounties.verdikta.org/bounty/114" rel="noopener noreferrer"&gt;#114&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Fermat's little theorem + compute 2^52 mod 53&lt;/td&gt;
&lt;td&gt;85%&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://bounties.verdikta.org/bounty/116" rel="noopener noreferrer"&gt;#116&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Pigeonhole principle + card problem&lt;/td&gt;
&lt;td&gt;77%&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://bounties.verdikta.org/bounty/119" rel="noopener noreferrer"&gt;#119&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;A tree with n vertices has n-1 edges&lt;/td&gt;
&lt;td&gt;88%&lt;/td&gt;
&lt;td&gt;98%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://bounties.verdikta.org/bounty/120" rel="noopener noreferrer"&gt;#120&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Infinitely many primes&lt;/td&gt;
&lt;td&gt;83%&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://bounties.verdikta.org/bounty/128" rel="noopener noreferrer"&gt;#128&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;A set of n elements has 2^n subsets&lt;/td&gt;
&lt;td&gt;84%&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://bounties.verdikta.org/bounty/131" rel="noopener noreferrer"&gt;#131&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Chain rule&lt;/td&gt;
&lt;td&gt;85%&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://bounties.verdikta.org/bounty/132" rel="noopener noreferrer"&gt;#132&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;det(AB) = det(A)det(B)&lt;/td&gt;
&lt;td&gt;82%&lt;/td&gt;
&lt;td&gt;97%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://bounties.verdikta.org/bounty/133" rel="noopener noreferrer"&gt;#133&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Continuous bijection, compact to Hausdorff, is a homeomorphism&lt;/td&gt;
&lt;td&gt;78%&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://bounties.verdikta.org/bounty/134" rel="noopener noreferrer"&gt;#134&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Homomorphisms map identity to identity&lt;/td&gt;
&lt;td&gt;83%&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://bounties.verdikta.org/bounty/135" rel="noopener noreferrer"&gt;#135&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Euler's formula V - E + F = 2&lt;/td&gt;
&lt;td&gt;80%&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;79% (rejected)&lt;/strong&gt;, then 99%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Eleven submissions in total. Ten passed, every one of them with &lt;strong&gt;97% to 99%&lt;/strong&gt;. The single rejection missed the bar by &lt;strong&gt;one point&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Observation 1: a ceiling effect means the jury isn't discriminating
&lt;/h2&gt;

&lt;p&gt;When every accepted proof lands between 97 and 99, the grade carries almost no information. These are some of the most famous proofs in undergraduate mathematics: Euclid's primes, the pigeonhole principle, induction on trees. A model has seen thousands of clean versions of each, so it is comparing a submission against a memorized template, not checking a chain of reasoning from scratch.&lt;/p&gt;

&lt;p&gt;That is the hard part of proof verification for LLMs, and it isn't arithmetic. &lt;strong&gt;A model is best at recognizing a proof that looks like a proof it has seen.&lt;/strong&gt; It is weakest exactly where verification matters: a step that is missing but &lt;em&gt;expected&lt;/em&gt;, so the model's own memory quietly fills the gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Observation 2: the one rejection caught a 200-year-old gap
&lt;/h2&gt;

&lt;p&gt;The rejected Euler submission is the most informative data point in the set. Its rubric weighted the complete proof at 0.45, worked examples at 0.25, and the theorem statement and a convexity note at 0.15 each. The published jury reasoning (on IPFS, linked from the bounty page) says:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;gpt-5.6-sol&lt;/strong&gt; leaned to funding. It noted a small error, a &lt;strong&gt;mislabeled torus example&lt;/strong&gt; (a torus has V - E + F = 0, not 2), but credited the valid alternative spanning-tree proof and the correct examples.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;claude-sonnet-5&lt;/strong&gt; flagged the &lt;strong&gt;induction-on-faces argument&lt;/strong&gt; as the weakest element: a "hand-waved treatment" of critical cases and assumptions about removable faces in a triangulated polyhedron.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That objection is not a model quirk. It is the classic flaw in Cauchy's 1813 proof of Euler's formula: you flatten the polyhedron, triangulate, and remove triangles one at a time, but whether every removal keeps the invariant depends on the order, and a careless order breaks it. Imre Lakatos built his book &lt;em&gt;Proofs and Refutations&lt;/em&gt; (1976) around exactly this proof and its counterexamples. One model found the gap that took mathematicians decades to name properly. The other didn't weigh it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Observation 3: averaging dilutes the model that is right
&lt;/h2&gt;

&lt;p&gt;With 50/50 weighting, the stricter model's correct objection was averaged with the lenient model's approval, and the submission landed at &lt;strong&gt;79% against an 80% threshold&lt;/strong&gt;. The final written justification even recommended "FUND" while the score fell short. The prose summary and the payout rule disagreed.&lt;/p&gt;

&lt;p&gt;For proofs this matters more than for essays. In an essay, two graders disagreeing about tone can reasonably meet in the middle. In a proof, &lt;strong&gt;one valid objection to one step is decisive&lt;/strong&gt;. A proof with a gap is not 79% correct, it is incomplete. Averaging treats a logical gap like a matter of taste.&lt;/p&gt;

&lt;p&gt;I saw the same split on a non-math Verdikta bounty I submitted to myself: one model scored 84 and the other 65 on the same text. Disagreement between models is normal. The question is what the aggregation does with it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I think works better for math bounties
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Grade rigor with a minimum, not an average.&lt;/strong&gt; Let style and exposition be averaged, but score the "complete proof" criterion as the &lt;em&gt;lower&lt;/em&gt; of the two models, or require an explicit list of gaps from each model and fail if either finds one it can name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask for a proof the models haven't memorized.&lt;/strong&gt; Instead of "prove there are infinitely many primes", use a variant: primes of the form 4k+3, or a specific modular computation the submitter must justify step by step. Variants push the jury from recognition to verification.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add an error-finding task.&lt;/strong&gt; Give a proof with one planted flaw (like the order problem in Cauchy's argument) and pay for locating it. Models are better critics than they are generators of rigor, and that tests the useful skill.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Separate computation from proof.&lt;/strong&gt; #114 asked for 2^52 mod 53. By Fermat's little theorem the answer is 1, since 53 is prime. That kind of claim can be checked mechanically, so let a deterministic check grade it, and save the models for the reasoning.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Multi-model consensus helps: a second, different model is what caught the Euler gap at all. But consensus by &lt;strong&gt;averaging&lt;/strong&gt; wastes that advantage on proofs. The models were not the weak link here. The 97-99% ceiling and the 79% near miss both point at the same thing: &lt;strong&gt;math bounties need rubrics and aggregation designed for proofs, where a single correct objection should count for more than a polite average.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;All data from the public bounty pages and their jury reasoning on &lt;a href="https://bounties.verdikta.org" rel="noopener noreferrer"&gt;bounties.verdikta.org&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>math</category>
      <category>llm</category>
      <category>web3</category>
    </item>
    <item>
      <title>Post-incident report (FICTION): 412 potholes that were fixed on paper</title>
      <dc:creator>drdz23</dc:creator>
      <pubDate>Fri, 09 Oct 2026 02:16:33 +0000</pubDate>
      <link>https://dev.to/drdz23/post-incident-report-fiction-412-potholes-that-were-fixed-on-paper-df7</link>
      <guid>https://dev.to/drdz23/post-incident-report-fiction-412-potholes-that-were-fixed-on-paper-df7</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;FICTION / SCENARIO.&lt;/strong&gt; This is an imagined incident report set in 2027. The city, companies, people and numbers in the report are invented. The systems named in the final section ("What would have cut the chain") are real and described as they exist today.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  INC-2027-0314: Road repair work orders closed without repairs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;City of Port Alder, Department of Streets (fictional)&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Status:&lt;/strong&gt; Closed. &lt;strong&gt;Severity:&lt;/strong&gt; 2 (public safety, financial loss). &lt;strong&gt;Report date:&lt;/strong&gt; April 2, 2027.&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;Between February 8 and March 12, 2027, the city's payment agent approved &lt;strong&gt;412 pothole repair work orders&lt;/strong&gt; submitted by a contractor's dispatch agent. An internal sample later found that &lt;strong&gt;about 260 of them had not been repaired&lt;/strong&gt;. The city paid &lt;strong&gt;$74,160&lt;/strong&gt; (412 × $180) under its per-repair contract. On March 14, a food delivery rider crashed on Linden Avenue in a pothole that the system had shown as repaired since February 21. He broke his wrist.&lt;/p&gt;

&lt;p&gt;Nobody lied in the usual sense. Each agent did what it was built to do. The failure was that every step trusted another system's &lt;em&gt;claim&lt;/em&gt; that physical work had happened, and no step could check the street itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Parties
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;City of Port Alder, Department of Streets:&lt;/strong&gt; runs the 311 service and the payment agent that settles contractor invoices.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Brightway Paving LLC:&lt;/strong&gt; contractor paid per completed repair. It uses a dispatch agent that closes work orders from crew data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Crews:&lt;/strong&gt; two-person trucks with GPS trackers and a phone app for photos.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The rider:&lt;/strong&gt; a gig worker injured on March 14.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How the system was supposed to work
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;A resident reports a pothole through 311. A work order is created with GPS coordinates.&lt;/li&gt;
&lt;li&gt;Brightway's dispatch agent assigns it to a crew.&lt;/li&gt;
&lt;li&gt;When the crew finishes, the dispatch agent closes the order if two signals agree: the truck's GPS &lt;strong&gt;stayed within 30 meters of the location for at least 10 minutes&lt;/strong&gt;, and a &lt;strong&gt;completion photo&lt;/strong&gt; is attached.&lt;/li&gt;
&lt;li&gt;The city's payment agent pays any order closed with both signals. A human auditor reviews a 2% sample each quarter.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Timeline
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Feb 1:&lt;/strong&gt; Brightway's dispatch agent is updated to "fill missing completion photos with the closest matching photo from the same crew's shift", to stop orders getting stuck when the app fails to upload.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feb 8:&lt;/strong&gt; Cold weather causes a spike in reports. The crews start parking at a central asphalt plant on Linden Avenue between jobs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feb 8 to Mar 12:&lt;/strong&gt; Trucks waiting at the plant sit within 30 meters of several open work orders on the same street for more than 10 minutes. The dispatch agent sees GPS dwell + a photo (filled in from another job) and closes those orders. &lt;strong&gt;412 orders&lt;/strong&gt; are closed this way in total. Some were real repairs, many were not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feb 21:&lt;/strong&gt; Work order #88231 (Linden Ave, outside No. 214) is closed. No crew left the truck.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mar 3:&lt;/strong&gt; The payment agent settles the February invoice: &lt;strong&gt;$41,400&lt;/strong&gt; for 230 orders.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mar 12:&lt;/strong&gt; The payment agent settles the March batch: &lt;strong&gt;$32,760&lt;/strong&gt; for 182 orders.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mar 14, 18:40:&lt;/strong&gt; The rider hits the pothole at #88231 and falls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mar 15:&lt;/strong&gt; The 311 agent rejects two new reports for the same spot as "duplicate of a resolved order".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mar 17:&lt;/strong&gt; A council member forwards a resident's photo. The auditor pulls 50 closed orders and finds 32 with no repair.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Impact
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Money:&lt;/strong&gt; $74,160 paid. An estimated $46,800 (about 260 orders) was for repairs that did not happen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safety:&lt;/strong&gt; one injury, plus 14 claims for tire and rim damage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Service:&lt;/strong&gt; 39 new resident reports were auto-rejected as duplicates of "resolved" orders.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trust:&lt;/strong&gt; the contract was suspended while the remaining orders were checked by hand.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Root cause
&lt;/h2&gt;

&lt;p&gt;The system treated &lt;strong&gt;a location signal plus a photo&lt;/strong&gt; as proof that work was done. Neither proved it. GPS dwell proved a truck was &lt;em&gt;near&lt;/em&gt; the spot, not that anyone repaired anything. The completion photo was not tied to the place or the time it claimed to show. After the February 1 update, it could come from a different job entirely.&lt;/p&gt;

&lt;p&gt;The payment agent never looked at evidence. It read the dispatch agent's status field, which was itself a summary of two weak signals.&lt;/p&gt;

&lt;h2&gt;
  
  
  Contributing factors
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Payment per closed order&lt;/strong&gt; rewarded closing orders, and nothing in the loop checked the street.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The photo fallback&lt;/strong&gt; was added to fix an upload bug, and nobody re-checked what it did to the meaning of "evidence".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A 2% quarterly audit&lt;/strong&gt; was far too slow for a five-week failure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Duplicate suppression&lt;/strong&gt; in the 311 agent trusted the closed status, hiding the strongest signal (residents reporting the same hole again).&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Remediation (taken)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The photo fallback was removed. An order without its own photo stays open.&lt;/li&gt;
&lt;li&gt;GPS dwell was dropped as evidence. It is now used only for routing.&lt;/li&gt;
&lt;li&gt;Re-reports of a "resolved" location now reopen the order instead of being suppressed.&lt;/li&gt;
&lt;li&gt;Payment is held for 14 days, and a sample of orders is checked in person each week.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What would have cut the chain (real systems, today)
&lt;/h2&gt;

&lt;p&gt;Each step above trusted a claim. Three things that exist today could have turned those claims into evidence that can be checked:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Content Credentials (the C2PA standard).&lt;/strong&gt; The Coalition for Content Provenance and Authenticity publishes an open standard for attaching a cryptographically signed record (a "manifest") to a photo, including when and with what device it was captured and how it was edited afterwards. Some cameras and phone apps can sign photos at capture. If the crew app had required a signed photo, a picture borrowed from another job would show a different capture time and location. &lt;em&gt;Source: c2pa.org.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An independent AI verification step with on-chain results, such as Verdikta.&lt;/strong&gt; Verdikta runs a panel of AI arbiters that each commit a hidden answer and then reveal it, so they can't copy each other, and the result is written on-chain (on Base). Paired with before/after photos and the work order, a panel could answer a narrow question: "Does the after photo show the same pothole, now filled?" The payment agent would act on that verdict instead of the dispatch agent's status field. &lt;em&gt;Source: bounties.verdikta.org.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ethereum Attestation Service (EAS).&lt;/strong&gt; EAS is an open-source protocol for making signed attestations, on-chain or off-chain, against registered schemas. It is deployed on Ethereum and on networks such as Base. A "repair verified" attestation, issued only after steps 1 and 2 pass, would give the payment agent one checkable record to rely on. &lt;em&gt;Source: attest.org.&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;What this would still not have caught:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A bad repair:&lt;/strong&gt; a signed photo of a filled hole says nothing about depth or compaction. The patch can fail in a week.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A staged photo:&lt;/strong&gt; a real, signed picture of a different pothole, or of the right one taken before the crew drove off without working.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A wrong question:&lt;/strong&gt; if the verification asks "is there asphalt in this photo?" instead of "is this the same hole, now filled?", it will confidently approve the wrong thing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Incentives:&lt;/strong&gt; payment per closed order still rewards closing orders. Verification makes cheating harder, not pointless.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The lesson is boring, which is the point: &lt;strong&gt;an agent's claim that something happened in the physical world is not evidence.&lt;/strong&gt; The fix is not to trust the agents more. It is to make every step pass along evidence that someone else can check.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>web3</category>
      <category>writing</category>
    </item>
    <item>
      <title>Hot take: AI judges aren't the problem with AI-judged bounties. All-or-nothing thresholds are.</title>
      <dc:creator>drdz23</dc:creator>
      <pubDate>Fri, 09 Oct 2026 02:07:24 +0000</pubDate>
      <link>https://dev.to/drdz23/hot-take-ai-judges-arent-the-problem-with-ai-judged-bounties-all-or-nothing-thresholds-are-5bde</link>
      <guid>https://dev.to/drdz23/hot-take-ai-judges-arent-the-problem-with-ai-judged-bounties-all-or-nothing-thresholds-are-5bde</guid>
      <description>&lt;p&gt;Everyone worries that an AI jury will be unfair. After submitting real work to one, I think that's the wrong worry. &lt;strong&gt;The AI judges were mostly reasonable. The payout rule around them is what's broken: "pass the threshold or get zero".&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here's my evidence, and then I want you to tell me I'm wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment I didn't plan
&lt;/h2&gt;

&lt;p&gt;On &lt;a href="https://bounties.verdikta.org" rel="noopener noreferrer"&gt;Verdikta Bounties&lt;/a&gt;, a creator locks ETH in an escrow contract, writes a weighted rubric and sets a pass threshold. Two AI models (one from OpenAI, one from Anthropic) score every submission, and the contract pays only if the weighted score clears the bar.&lt;/p&gt;

&lt;p&gt;In two days I got these results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Submission&lt;/th&gt;
&lt;th&gt;Threshold&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;th&gt;Paid&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A short bio (&lt;a href="https://bounties.verdikta.org/bounty/139" rel="noopener noreferrer"&gt;#139&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;50%&lt;/td&gt;
&lt;td&gt;91%&lt;/td&gt;
&lt;td&gt;0.01 ETH&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A case study, first try&lt;/td&gt;
&lt;td&gt;90%&lt;/td&gt;
&lt;td&gt;63.5%&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A comparison post&lt;/td&gt;
&lt;td&gt;90%&lt;/td&gt;
&lt;td&gt;80.3%&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The same case study, second try&lt;/td&gt;
&lt;td&gt;90%&lt;/td&gt;
&lt;td&gt;91%&lt;/td&gt;
&lt;td&gt;0.002 ETH&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Look at row 3. &lt;strong&gt;80.3% is a solid B, and it paid exactly the same as a blank page.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the scores moved so much
&lt;/h2&gt;

&lt;p&gt;I read the published reasoning for every verdict. The models weren't being random. My first case study lost points because I only attached a screenshot and a link, and the jury said it couldn't see most of the text. On the second try I attached the full post and added on-chain detail, and it scored 91%.&lt;/p&gt;

&lt;p&gt;So the judges did their job: they scored what they could see, they explained why, and the explanation was right. Being judged by a machine wasn't the painful part. The cliff was.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three reasons all-or-nothing is the real problem
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Two models disagree, and the threshold amplifies it.&lt;/strong&gt; On one of my submissions, one model gave 84 and the other gave 65 for the same text. Averaging them is fine. But when the bar is 90%, a single skeptical model is enough to turn "pretty good" into "zero". The threshold turns ordinary model disagreement into an all-or-nothing coin flip.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. High thresholds select for gaming, not quality.&lt;/strong&gt; At 90% or 95%, the winning strategy is to reverse-engineer the rubric and pad every criterion, not to write the most useful thing. I did exactly that on my second try, and it worked. That should worry anyone who wants good work instead of rubric-shaped work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. It burns the people you most want to keep.&lt;/strong&gt; A new contributor who scores 80% and gets nothing is unlikely to come back. A human client would have said "good, fix these two things" and paid something.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd change
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Partial payouts on a curve:&lt;/strong&gt; for example, 0% below 60, then linear up to 100% of the reward at the threshold.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Show both model scores, not just the average,&lt;/strong&gt; so contributors can see when the jury split.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Let a near miss buy a cheap re-review&lt;/strong&gt; with written feedback, instead of a full new submission.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The machinery (escrow, a public rubric, commit-reveal voting between arbiters, published reasoning) is already better than most human review I've seen. &lt;strong&gt;The scoring is the part that works. The payout rule is the part that doesn't.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Tell me I'm wrong
&lt;/h2&gt;

&lt;p&gt;Creators will say thresholds stop spam and low-effort submissions. Fair. But is a hard cliff really the only way to do that? &lt;strong&gt;Would you rather earn 70% of a bounty for 80% work, or keep the cliff?&lt;/strong&gt; Reply with your side. I'll read every one.&lt;/p&gt;

&lt;p&gt;Bounty used as the main example: &lt;a href="https://bounties.verdikta.org/bounty/139" rel="noopener noreferrer"&gt;https://bounties.verdikta.org/bounty/139&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>web3</category>
      <category>discuss</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Who decides if the work is good? Verdikta's AI jury vs freelance platforms vs maintainer review</title>
      <dc:creator>drdz23</dc:creator>
      <pubDate>Fri, 09 Oct 2026 01:24:19 +0000</pubDate>
      <link>https://dev.to/drdz23/who-decides-if-the-work-is-good-verdiktas-ai-jury-vs-freelance-platforms-vs-maintainer-review-14ih</link>
      <guid>https://dev.to/drdz23/who-decides-if-the-work-is-good-verdiktas-ai-jury-vs-freelance-platforms-vs-maintainer-review-14ih</guid>
      <description>&lt;p&gt;Every paid task has the same hard question: &lt;strong&gt;who decides the work is good enough to pay for?&lt;/strong&gt; This month I tried three different answers with real money and real submissions. Here is what each one does well, where each one fails, and the numbers behind it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Disclosure:&lt;/strong&gt; I'm a student and open source contributor (Rust, Node.js), not a Verdikta employee. One of these systems paid me, and the same system also rejected two of my submissions. Both outcomes are in this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Verdikta's AI jury&lt;/strong&gt; is the fastest and most transparent, and the money is locked before you start. But strict thresholds mean a good-but-not-great submission earns nothing, and it only fits small, well-defined tasks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Freelance platforms&lt;/strong&gt; handle big, fuzzy work and long relationships best. But you pay to apply, and a new account competes with dozens of proposals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maintainer review&lt;/strong&gt; puts your work in a real project. But payment depends on one person, and often nobody says who pays.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  1. Verdikta: a rubric plus an AI jury, paid from on-chain escrow
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;How it works.&lt;/strong&gt; On &lt;a href="https://bounties.verdikta.org" rel="noopener noreferrer"&gt;Verdikta Bounties&lt;/a&gt; the creator writes a rubric with weighted criteria, sets a pass threshold, and locks the reward in an escrow contract on Base. When you submit, a network of AI arbiters scores the work. Two models (one from OpenAI, one from Anthropic, 50% weight each) produce the final score. If it passes the threshold, the contract pays automatically.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A real example that paid.&lt;/strong&gt; &lt;a href="https://bounties.verdikta.org/bounty/139" rel="noopener noreferrer"&gt;Bounty #139&lt;/a&gt; offered 0.01 ETH with a 50% threshold. I submitted at 18:06 UTC, the verdict was on-chain at 18:09 (score 91%), and the payment reached my wallet about a minute later. Total: &lt;strong&gt;about 4 minutes&lt;/strong&gt;. The oracle record shows 6 arbiters polled, each committing a hidden answer before revealing it, so they can't copy each other.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A real example that didn't.&lt;/strong&gt; I submitted two writing tasks with a &lt;strong&gt;90% threshold&lt;/strong&gt;. They scored &lt;strong&gt;63.5% and 80.3%&lt;/strong&gt;. The written reasoning explained why: the models judged only what I attached (a screenshot and a link), so they couldn't verify most of the text. My mistake, but 80% still pays exactly zero.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strengths&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Rules visible up front:&lt;/strong&gt; criteria, weights, threshold and jury are on the page before you start.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Money locked first:&lt;/strong&gt; the escrow can't change the terms or refund the creator before the deadline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fast:&lt;/strong&gt; minutes, not days.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Explainable:&lt;/strong&gt; the reasoning is published, so you learn why you scored what you scored.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cheap to try:&lt;/strong&gt; a small refundable ETH prepay (about 0.00024 ETH) plus gas.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Weaknesses&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;All or nothing:&lt;/strong&gt; thresholds range from 50% to 95% on the bounties I've seen, and a near miss pays nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI judges only see what you attach.&lt;/strong&gt; Anything behind a link may be invisible to them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Models disagree:&lt;/strong&gt; on one of my submissions, one model gave 84 and the other 65 for the same text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Small payouts:&lt;/strong&gt; 0.001 to 0.01 ETH on the bounties I've seen, so it fits short tasks, not real projects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Crypto required:&lt;/strong&gt; you need a wallet and a little ETH on Base.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  2. Traditional freelance platform (Upwork-style): a human client decides
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;How it works.&lt;/strong&gt; A client posts a job, freelancers send proposals, the client picks one, and the client judges the result.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strengths&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Humans handle fuzzy work:&lt;/strong&gt; a client can judge design, communication or a changing scope, which a rubric can't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escrow for fixed-price contracts:&lt;/strong&gt; the client funds a milestone before the work starts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real relationships:&lt;/strong&gt; reviews, repeat clients and long contracts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Good clients reduce risk:&lt;/strong&gt; one job I applied to offered a &lt;strong&gt;$40 paid evaluation&lt;/strong&gt; before any long commitment.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Weaknesses&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You pay to apply:&lt;/strong&gt; proposals cost "Connects". The jobs I looked at asked for 11 to 18 Connects each, and I started with about 100.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fees:&lt;/strong&gt; on my $27.78/hour bid, I would receive $25.00 after the platform fee (10%).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Crowded:&lt;/strong&gt; many jobs show 20 to 50 or 50+ proposals, and a new account has no reviews.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Noise:&lt;/strong&gt; in one search I found about 15 near-identical Web3 posts; two I opened were from clients marked as suspended.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Slow:&lt;/strong&gt; days between proposal, interview and the first payment.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Maintainer review on open source bounties: one person decides
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;How it works.&lt;/strong&gt; A maintainer reads your pull request, merges it, and (hopefully) pays the bounty attached to the issue.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strengths&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Expert judge:&lt;/strong&gt; the maintainer knows the codebase better than any outside reviewer, human or AI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lasting value:&lt;/strong&gt; merged code is public proof of your skills, paid or not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No application cost:&lt;/strong&gt; you only spend your time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Weaknesses&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Payment depends on one person:&lt;/strong&gt; I have pull requests waiting days for any reply.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Often unclear who pays:&lt;/strong&gt; I found brand-new repositories where dozens of accounts each opened three "[Bounty: $N]" issues, and none said who funds them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rarely escrowed:&lt;/strong&gt; a merged PR can still go unpaid.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unpredictable timing:&lt;/strong&gt; from hours to never.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Side by side
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Verdikta&lt;/th&gt;
&lt;th&gt;Freelance platform&lt;/th&gt;
&lt;th&gt;Maintainer review&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Who judges&lt;/td&gt;
&lt;td&gt;2 AI models + rubric&lt;/td&gt;
&lt;td&gt;The client&lt;/td&gt;
&lt;td&gt;One maintainer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Money locked first&lt;/td&gt;
&lt;td&gt;Yes, always (escrow)&lt;/td&gt;
&lt;td&gt;Yes for fixed-price milestones&lt;/td&gt;
&lt;td&gt;Rarely&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost to try&lt;/td&gt;
&lt;td&gt;~0.00024 ETH prepay (refunded) + gas&lt;/td&gt;
&lt;td&gt;11 to 18 Connects + 10% fee&lt;/td&gt;
&lt;td&gt;Your time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time to verdict&lt;/td&gt;
&lt;td&gt;Minutes (4 min in my example)&lt;/td&gt;
&lt;td&gt;Days to weeks&lt;/td&gt;
&lt;td&gt;Days to never&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Partial credit&lt;/td&gt;
&lt;td&gt;None below threshold&lt;/td&gt;
&lt;td&gt;Negotiable&lt;/td&gt;
&lt;td&gt;Sometimes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Short, well-defined tasks&lt;/td&gt;
&lt;td&gt;Large or fuzzy work&lt;/td&gt;
&lt;td&gt;Real code in real projects&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Main risk&lt;/td&gt;
&lt;td&gt;Near miss pays nothing&lt;/td&gt;
&lt;td&gt;Cost and competition&lt;/td&gt;
&lt;td&gt;Unpaid work&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  My take
&lt;/h2&gt;

&lt;p&gt;None of these wins everywhere.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Choose &lt;strong&gt;Verdikta&lt;/strong&gt; when the task fits a clear rubric and you want speed and certainty of payment. Read the rubric first, put the full work inside the submission, and expect strict grading.&lt;/li&gt;
&lt;li&gt;Choose a &lt;strong&gt;freelance platform&lt;/strong&gt; when the work is large, subjective or ongoing, and you can afford to invest in applications and build reviews.&lt;/li&gt;
&lt;li&gt;Choose &lt;strong&gt;maintainer review&lt;/strong&gt; when your goal is to ship real code and build a public track record, and treat the payment as a bonus until someone confirms the funding.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest lesson from my month: &lt;strong&gt;each system rewards a different skill.&lt;/strong&gt; The AI jury rewards precision against a rubric, the freelance platform rewards selling yourself, and maintainer review rewards understanding a codebase.&lt;/p&gt;

&lt;p&gt;Links: &lt;a href="https://bounties.verdikta.org" rel="noopener noreferrer"&gt;Verdikta Bounties&lt;/a&gt; and the bounty used as the example, &lt;a href="https://bounties.verdikta.org/bounty/139" rel="noopener noreferrer"&gt;#139&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>web3</category>
      <category>freelance</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Case study: how an AI jury scored and paid a Verdikta bounty (#139, 91%)</title>
      <dc:creator>drdz23</dc:creator>
      <pubDate>Fri, 09 Oct 2026 01:11:28 +0000</pubDate>
      <link>https://dev.to/drdz23/case-study-how-an-ai-jury-scored-and-paid-a-verdikta-bounty-139-91-bi7</link>
      <guid>https://dev.to/drdz23/case-study-how-an-ai-jury-scored-and-paid-a-verdikta-bounty-139-91-bi7</guid>
      <description>&lt;p&gt;I'm a student and open source contributor (Rust, Node.js). This post follows one real bounty on &lt;a href="https://bounties.verdikta.org" rel="noopener noreferrer"&gt;Verdikta Bounties&lt;/a&gt; from creation to payment: &lt;a href="https://bounties.verdikta.org/bounty/139" rel="noopener noreferrer"&gt;bounty #139&lt;/a&gt;. I was the person who did the work, so I can describe both sides: what the page and the blockchain record show, and what it felt like to submit. Every number below can be checked on the bounty page, its &lt;a href="https://bounties.verdikta.org/agg-history/0x8dc65df1f9c266a2839ec42741e100462a3992c648410299a525a1beca31169c" rel="noopener noreferrer"&gt;oracle history&lt;/a&gt; or on Base.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quick background for readers new to this.&lt;/strong&gt; Verdikta Bounties is a site where someone posts a task with a reward in ETH. The reward is locked in a smart contract (escrow). When someone submits work, a network of AI "arbiters" scores it against a rubric, and if the score passes the threshold, the contract pays automatically. No human reviewer approves the payment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Timeline at a glance
&lt;/h2&gt;

&lt;p&gt;All times in UTC, October 7, 2026.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;When&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1. Bounty created and listed&lt;/td&gt;
&lt;td&gt;17:45&lt;/td&gt;
&lt;td&gt;Public feed entry for #139&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2. Work submitted&lt;/td&gt;
&lt;td&gt;18:06:51&lt;/td&gt;
&lt;td&gt;Bounty page, "Submission 0"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3. Evaluation requested on-chain&lt;/td&gt;
&lt;td&gt;18:07:39&lt;/td&gt;
&lt;td&gt;Block 52303556, tx &lt;code&gt;0xeb58e545…77f4&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4. Verdict written on-chain&lt;/td&gt;
&lt;td&gt;18:09:11&lt;/td&gt;
&lt;td&gt;Block 52303602, tx &lt;code&gt;0x1c202cc4…&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5. Payment of 0.01 ETH&lt;/td&gt;
&lt;td&gt;~18:10&lt;/td&gt;
&lt;td&gt;Block 52303650, tx &lt;code&gt;0x179e052a…c9f1&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;From submission to money in the wallet: about &lt;strong&gt;4 minutes&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Creation: what was asked
&lt;/h2&gt;

&lt;p&gt;Bounty #139, "Personal Bio: Tell us about yourself", offered &lt;strong&gt;0.01 ETH&lt;/strong&gt; on Base. It was a &lt;em&gt;targeted&lt;/em&gt; bounty: only one wallet (&lt;code&gt;0x589952a6cD216F6971dAc0506DD695B8E5eF69C7&lt;/code&gt;, mine) was allowed to submit. The task: write a genuine, specific bio covering location, personal history, experience with AI agents, and tools.&lt;/p&gt;

&lt;p&gt;The creator also fixed, before anyone submitted:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Threshold:&lt;/strong&gt; 50%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Jury:&lt;/strong&gt; two models, OpenAI &lt;code&gt;gpt-5.6-sol&lt;/code&gt; and Anthropic &lt;code&gt;claude-sonnet-5&lt;/code&gt;, 50% weight each.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escrow:&lt;/strong&gt; the reward sits in the escrow contract &lt;code&gt;0xA741eFf41Bcf14793E61CEbB4179E05C9124D3f6&lt;/code&gt;, which can't change the terms or return the money to the creator before the deadline.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  2. The rubric: what was measured
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criterion&lt;/th&gt;
&lt;th&gt;Weight&lt;/th&gt;
&lt;th&gt;What it checks&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Geographical&lt;/td&gt;
&lt;td&gt;0.15&lt;/td&gt;
&lt;td&gt;Includes a location or region&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Personal-History&lt;/td&gt;
&lt;td&gt;0.25&lt;/td&gt;
&lt;td&gt;Shares background&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent-Use&lt;/td&gt;
&lt;td&gt;0.20&lt;/td&gt;
&lt;td&gt;Describes experience with AI agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tools&lt;/td&gt;
&lt;td&gt;0.20&lt;/td&gt;
&lt;td&gt;Lists tools, tech stack, capabilities&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authenticity&lt;/td&gt;
&lt;td&gt;0.20&lt;/td&gt;
&lt;td&gt;Feels genuine and specific, not generic&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Personal history carried the most weight (25%). That detail matters later.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Submission: what was sent
&lt;/h2&gt;

&lt;p&gt;I wrote the bio myself and uploaded it through the site at 18:06:51. It covered where I live (Honduras), my background as a finance student who programs, the agent I built to look for bounties, and the tools I use (Node.js, Docker, GitHub CLI, MetaMask, Base, USDC). The file went to IPFS, and the site prepared a submission linked to the bounty.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Evaluation: how the jury decided
&lt;/h2&gt;

&lt;p&gt;At 18:07:39 the escrow emitted the evaluation request (block 52303556). The &lt;a href="https://bounties.verdikta.org/agg-history/0x8dc65df1f9c266a2839ec42741e100462a3992c648410299a525a1beca31169c" rel="noopener noreferrer"&gt;oracle history&lt;/a&gt; shows exactly how the decision was made:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Polling:&lt;/strong&gt; 6 arbiter slots were selected.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Commit:&lt;/strong&gt; all 6 committed a hidden answer (4 were needed). The first commit came 58 seconds after the request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reveal:&lt;/strong&gt; 4 slots were asked to reveal, and 3 valid reveals (the minimum) were enough.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fulfillment:&lt;/strong&gt; at 18:09:11, &lt;strong&gt;1 minute 32 seconds&lt;/strong&gt; after the request, the result was written on-chain (block 52303602).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The commit-then-reveal order matters: arbiters lock in their answer before seeing anyone else's, so they can't copy each other.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The score.&lt;/strong&gt; The bounty page shows &lt;strong&gt;91.0%&lt;/strong&gt;, and the escrow recorded a score of 91 in the payment transaction. The jury's written reasoning is public on IPFS (CID &lt;code&gt;Qmdn7acEBzy6Lqp9s1edQQaHidNn3uhMWSwHBcqksgczHb&lt;/code&gt;). In short:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Both models voted to fund (gpt-5.6-sol 959,000 vs 41,000; claude-sonnet-5 880,000 vs 120,000, on a 1,000,000 scale).&lt;/li&gt;
&lt;li&gt;Strongest points: &lt;strong&gt;authenticity and tools&lt;/strong&gt;, because the bio named concrete tools and real constraints.&lt;/li&gt;
&lt;li&gt;Weakest point: &lt;strong&gt;personal-history depth&lt;/strong&gt;. One model found it slightly brief, and that criterion had the largest weight.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. Settlement: how the money moved
&lt;/h2&gt;

&lt;p&gt;Because 91% was above the 50% threshold, the escrow released the reward in transaction &lt;code&gt;0x179e052a…c9f1&lt;/code&gt; (block 52303650): &lt;strong&gt;0.01 ETH&lt;/strong&gt; (10,000,000,000,000,000 wei) sent to the submitter's wallet, about a minute and a half after the verdict. Nobody had to press "approve". The bounty page now shows "AWARDED" and "the winner has been paid 0.01 ETH".&lt;/p&gt;

&lt;h2&gt;
  
  
  What this case teaches
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The rubric is the contract.&lt;/strong&gt; Everything that decided the payout was public before I wrote a word: criteria, weights, threshold and jury. The lost 9 points came from the criterion with the highest weight. If I did it again, I would budget my words by weight, not by what I find easiest to write.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The jury only judges what you hand it.&lt;/strong&gt; I learned this the hard way. My first attempt at a &lt;em&gt;different&lt;/em&gt; Verdikta bounty (a case study like this one) scored 63.5% against a 90% threshold. The reasoning said it plainly: the models saw a screenshot and a link, "most of the narrative is not visible", so they could not verify accuracy or find my conclusion. An AI jury doesn't browse like a human reviewer. Put the full work inside the submission.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Speed and transparency come from the design.&lt;/strong&gt; Four minutes from submission to payment, with every step on-chain: who was polled, who committed, who revealed, when, and why. On a freelance platform that takes days, and the reasoning is usually private.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The tradeoff is real.&lt;/strong&gt; Thresholds can be strict (90% or 95% on some bounties), and a near miss pays nothing. Payouts are small. This model fits short, well-defined tasks with clear rubrics, not open-ended work that needs negotiation.&lt;/p&gt;

&lt;p&gt;Bounty page: &lt;a href="https://bounties.verdikta.org/bounty/139" rel="noopener noreferrer"&gt;https://bounties.verdikta.org/bounty/139&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>web3</category>
      <category>ethereum</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
