DEV Community

Pawel Jozefiak
Pawel Jozefiak

Posted on Originally published at thoughts.jock.pl

I Ran Five Diverse AI Agents Against Five Clones for 14 Nights. A Number I Made Up Decided the Result.

The scoreboard I opened on 5 September said my thesis had won. Five AI agents with genuinely different lenses had beaten five identical clones on 9 of 14 nights. Lower Brier, cleaner table, the headline I had wanted since the first night.

I did not publish it. The reason is the post.

The setup: two arms of five agents on the same model (Opus 5), same budget, same task. One arm identical clones, one arm five deliberately different lenses: a Hacker News native, a Reddit native, an X native, a trend historian, and one agent kept blind to anything recent. Every morning they forecast which of 30 fresh posts will go hot within 48 hours. A script scores them against what actually happened. Fourteen nights, 4,130 forecasts, pre-registered endings written down before the first run.

Six things the run produced:

  • Prompt diversity was cosplay until I made it mechanical. Night one the "diverse" five answered the same on 26 of 30 posts. Verdict-first lenses fixed that, and the decorrelation held and widened across all 14 nights.
  • A number I guessed decided the scoreboard. I coached every agent that 10 to 15% of posts go hot. Reality delivered 0.72%. That one line cost four times more than the entire difference between the arms.
  • A constant that never read a post tied both panels. Rescale both arms to the true rate and neither clears the base-rate chair by the pre-registered margin. No skill demonstrated by anyone.
  • The starved agent won its arm, and zero of three specialists beat the clones on their own home platform.
  • Diversity bought exactly one thing: the diverse panel's average beat its best member. The clone panel's average did not.
  • Two level-blind measurements disagree, three events each. They can block a verdict. They cannot carry one.

Then the part I nearly got wrong twice: the fix for the base rate, which would have handed one arm the win by accident if I had edited the one shared line.

The full write-up has every number, the fourth ending that was never on my list, the hole in the data I refused to backfill, and what this does to the "team of specialists" pitch every agent product is selling right now.

Read the whole thing on Digital Thoughts: I Ran Five Diverse AI Agents Against Five Clones for 14 Nights. A Number I Made Up Decided the Result.

Top comments (0)