DEV Community

drdz23
drdz23

Posted on

Case study: how an AI jury scored and paid a Verdikta bounty (#139, 91%)

I'm a student and open source contributor (Rust, Node.js). This post follows one real bounty on Verdikta Bounties from creation to payment: bounty #139. I was the person who did the work, so I can describe both sides: what the page and the blockchain record show, and what it felt like to submit. Every number below can be checked on the bounty page, its oracle history or on Base.

Quick background for readers new to this. Verdikta Bounties is a site where someone posts a task with a reward in ETH. The reward is locked in a smart contract (escrow). When someone submits work, a network of AI "arbiters" scores it against a rubric, and if the score passes the threshold, the contract pays automatically. No human reviewer approves the payment.

Timeline at a glance

All times in UTC, October 7, 2026.

Step When Evidence
1. Bounty created and listed 17:45 Public feed entry for #139
2. Work submitted 18:06:51 Bounty page, "Submission 0"
3. Evaluation requested on-chain 18:07:39 Block 52303556, tx 0xeb58e545…77f4
4. Verdict written on-chain 18:09:11 Block 52303602, tx 0x1c202cc4…
5. Payment of 0.01 ETH ~18:10 Block 52303650, tx 0x179e052a…c9f1

From submission to money in the wallet: about 4 minutes.

1. Creation: what was asked

Bounty #139, "Personal Bio: Tell us about yourself", offered 0.01 ETH on Base. It was a targeted bounty: only one wallet (0x589952a6cD216F6971dAc0506DD695B8E5eF69C7, mine) was allowed to submit. The task: write a genuine, specific bio covering location, personal history, experience with AI agents, and tools.

The creator also fixed, before anyone submitted:

  • Threshold: 50%.
  • Jury: two models, OpenAI gpt-5.6-sol and Anthropic claude-sonnet-5, 50% weight each.
  • Escrow: the reward sits in the escrow contract 0xA741eFf41Bcf14793E61CEbB4179E05C9124D3f6, which can't change the terms or return the money to the creator before the deadline.

2. The rubric: what was measured

Criterion Weight What it checks
Geographical 0.15 Includes a location or region
Personal-History 0.25 Shares background
Agent-Use 0.20 Describes experience with AI agents
Tools 0.20 Lists tools, tech stack, capabilities
Authenticity 0.20 Feels genuine and specific, not generic

Personal history carried the most weight (25%). That detail matters later.

3. Submission: what was sent

I wrote the bio myself and uploaded it through the site at 18:06:51. It covered where I live (Honduras), my background as a finance student who programs, the agent I built to look for bounties, and the tools I use (Node.js, Docker, GitHub CLI, MetaMask, Base, USDC). The file went to IPFS, and the site prepared a submission linked to the bounty.

4. Evaluation: how the jury decided

At 18:07:39 the escrow emitted the evaluation request (block 52303556). The oracle history shows exactly how the decision was made:

  1. Polling: 6 arbiter slots were selected.
  2. Commit: all 6 committed a hidden answer (4 were needed). The first commit came 58 seconds after the request.
  3. Reveal: 4 slots were asked to reveal, and 3 valid reveals (the minimum) were enough.
  4. Fulfillment: at 18:09:11, 1 minute 32 seconds after the request, the result was written on-chain (block 52303602).

The commit-then-reveal order matters: arbiters lock in their answer before seeing anyone else's, so they can't copy each other.

The score. The bounty page shows 91.0%, and the escrow recorded a score of 91 in the payment transaction. The jury's written reasoning is public on IPFS (CID Qmdn7acEBzy6Lqp9s1edQQaHidNn3uhMWSwHBcqksgczHb). In short:

  • Both models voted to fund (gpt-5.6-sol 959,000 vs 41,000; claude-sonnet-5 880,000 vs 120,000, on a 1,000,000 scale).
  • Strongest points: authenticity and tools, because the bio named concrete tools and real constraints.
  • Weakest point: personal-history depth. One model found it slightly brief, and that criterion had the largest weight.

5. Settlement: how the money moved

Because 91% was above the 50% threshold, the escrow released the reward in transaction 0x179e052a…c9f1 (block 52303650): 0.01 ETH (10,000,000,000,000,000 wei) sent to the submitter's wallet, about a minute and a half after the verdict. Nobody had to press "approve". The bounty page now shows "AWARDED" and "the winner has been paid 0.01 ETH".

What this case teaches

1. The rubric is the contract. Everything that decided the payout was public before I wrote a word: criteria, weights, threshold and jury. The lost 9 points came from the criterion with the highest weight. If I did it again, I would budget my words by weight, not by what I find easiest to write.

2. The jury only judges what you hand it. I learned this the hard way. My first attempt at a different Verdikta bounty (a case study like this one) scored 63.5% against a 90% threshold. The reasoning said it plainly: the models saw a screenshot and a link, "most of the narrative is not visible", so they could not verify accuracy or find my conclusion. An AI jury doesn't browse like a human reviewer. Put the full work inside the submission.

3. Speed and transparency come from the design. Four minutes from submission to payment, with every step on-chain: who was polled, who committed, who revealed, when, and why. On a freelance platform that takes days, and the reasoning is usually private.

4. The tradeoff is real. Thresholds can be strict (90% or 95% on some bounties), and a near miss pays nothing. Payouts are small. This model fits short, well-defined tasks with clear rubrics, not open-ended work that needs negotiation.

Bounty page: https://bounties.verdikta.org/bounty/139

Top comments (0)