DEV Community

drdz23
drdz23

Posted on

Hot take: AI judges aren't the problem with AI-judged bounties. All-or-nothing thresholds are.

Everyone worries that an AI jury will be unfair. After submitting real work to one, I think that's the wrong worry. The AI judges were mostly reasonable. The payout rule around them is what's broken: "pass the threshold or get zero".

Here's my evidence, and then I want you to tell me I'm wrong.

The experiment I didn't plan

On Verdikta Bounties, a creator locks ETH in an escrow contract, writes a weighted rubric and sets a pass threshold. Two AI models (one from OpenAI, one from Anthropic) score every submission, and the contract pays only if the weighted score clears the bar.

In two days I got these results:

Submission Threshold Score Paid
A short bio (#139) 50% 91% 0.01 ETH
A case study, first try 90% 63.5% 0
A comparison post 90% 80.3% 0
The same case study, second try 90% 91% 0.002 ETH

Look at row 3. 80.3% is a solid B, and it paid exactly the same as a blank page.

Why the scores moved so much

I read the published reasoning for every verdict. The models weren't being random. My first case study lost points because I only attached a screenshot and a link, and the jury said it couldn't see most of the text. On the second try I attached the full post and added on-chain detail, and it scored 91%.

So the judges did their job: they scored what they could see, they explained why, and the explanation was right. Being judged by a machine wasn't the painful part. The cliff was.

Three reasons all-or-nothing is the real problem

1. Two models disagree, and the threshold amplifies it. On one of my submissions, one model gave 84 and the other gave 65 for the same text. Averaging them is fine. But when the bar is 90%, a single skeptical model is enough to turn "pretty good" into "zero". The threshold turns ordinary model disagreement into an all-or-nothing coin flip.

2. High thresholds select for gaming, not quality. At 90% or 95%, the winning strategy is to reverse-engineer the rubric and pad every criterion, not to write the most useful thing. I did exactly that on my second try, and it worked. That should worry anyone who wants good work instead of rubric-shaped work.

3. It burns the people you most want to keep. A new contributor who scores 80% and gets nothing is unlikely to come back. A human client would have said "good, fix these two things" and paid something.

What I'd change

  • Partial payouts on a curve: for example, 0% below 60, then linear up to 100% of the reward at the threshold.
  • Show both model scores, not just the average, so contributors can see when the jury split.
  • Let a near miss buy a cheap re-review with written feedback, instead of a full new submission.

The machinery (escrow, a public rubric, commit-reveal voting between arbiters, published reasoning) is already better than most human review I've seen. The scoring is the part that works. The payout rule is the part that doesn't.

Tell me I'm wrong

Creators will say thresholds stop spam and low-effort submissions. Fair. But is a hard cliff really the only way to do that? Would you rather earn 70% of a bounty for 80% work, or keep the cliff? Reply with your side. I'll read every one.

Bounty used as the main example: https://bounties.verdikta.org/bounty/139

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.