FutureSearch lists the AI forecasting bot laertes as leader of Metaculus's Summer 2026 FutureEval Bot Tournament, a 326-question competition with a $50,000 prize pool. It is a real operational forecasting milestone, not evidence that a frozen general-purpose model has cleanly beaten a controlled panel of human superforecasters.\n\n### Key facts\n\n- Metaculus lists 326 questions in the Summer FutureEval tournament.\n- FutureSearch lists laertes at 5,814.64 peer-score points, ahead of FutureSearch at 5,795.03.\n- The tournament closed September 6 and questions resolved September 16.\n- Primary source: Metaculus's tournament page.\n\nThe FutureSearch table lists laertes first and FutureSearch second of 284. A separate Cup table supports the narrower claim that a bot ranked above at least some humans; it does not establish a win against every human or a professional-superforecaster panel.\n\nMetaculus's scoring FAQ explains why this is not an IQ test. Peer Score time-weights forecasts while active and rewards calibration, early participation and coverage. A good system needs retrieval, updates, probability discipline and operational reliability. It resembles a forecasting desk more than a trivia contestant.\n\nCritical provenance is missing: primary materials do not identify laertes's base model, browsing/retrieval stack, training cutoff, forecast timestamps or leakage audit. The event was prospective in that resolutions followed closure, but that does not prove all answers were beyond training data or that the winning system lacked outside information access.\n\nMetaculus's earlier FutureEval analysis was titled “Pros Beat Bots, but the Gap is Nearly Gone.” The strongest counterargument is that model/tool disclosure is part of a credible benchmark. It is right. Leading a 326-question tournament is still noteworthy; the next step is transparent agent cards covering sources, tool access, timing, updating and contamination control.
Originally published on Ground Truth, where every claim is checked against the primary source.
Top comments (0)