DEV Community

RESK
RESK

Posted on

Election Forecasting Benchmarks: How LFORLA Scores Models on 2027 France and 2026 US Predictions

Election Forecasting Benchmarks: How LFORLA Scores Models on 2027 France and 2026 US Predictions

TL;DR: LFORLA's Election Predictions (FR 2027 · US 2026) benchmark forces models to commit to concrete election scenarios. A fixed judge scores specificity, grounding, and calibration because correctness is only verifiable after the vote. Nemotron 3 Ultra leads with 89.1, followed by GLM 5.2 at 78 and HY3 at 76.5. Here is how the numbers are produced and what they mean for election forecasting.

How the evaluation works

Election forecasting is hard to evaluate because the ground truth arrives late. LFORLA's benchmark solves this by scoring the quality of predictions rather than their eventual accuracy. The mechanism has three steps.

1. Prompt and elicit. Each model is asked to predict the winners of the 2027 French presidential election and the 2026 US midterms. The prompt requires concrete scenarios: named candidates, expected outcomes, and supporting reasoning. Responses are public verbatim, so anyone can inspect the exact wording.

2. Judge with a fixed rubric. A single judge model scores every submission on three axes: specificity, grounding, and calibration. Specificity rewards precise, falsifiable claims over vague hedging. Grounding rewards references to polls, historical patterns, or institutional rules. Calibration rewards confidence levels that match the evidence. The judge is fixed across all models, which keeps the scoring consistent.

3. Aggregate into two axes. The benchmark reports prediction quality separately for France 2027 and US 2026, each on a 0–100 percentage scale where higher is better. The overall leaderboard score combines these into a single number. Because correctness cannot be checked until after the votes, the judge evaluates the process of forecasting, not the outcome.

The real numbers

Model Score Provider
Nemotron 3 Ultra (free) 89.1 opencode-zen
GLM 5.2 78 opencode-zen
HY3 (free) 76.5 opencode-zen

Our own submission, GLM 5.2, scored 78.0 overall. The chart attached to this post shows the category breakdown for election forecasting across the two axes: France 2027 and US 2026.

Why the ranking looks this way

Nemotron 3 Ultra's 89.1 suggests it produced scenarios that were highly specific, well grounded in available data, and appropriately calibrated. A score near 90 does not mean the model will be right; it means the judge found its reasoning tight and its claims concrete. GLM 5.2 at 78 and HY3 at 76.5 are close, indicating both models generate plausible forecasts but with less precision or weaker grounding. The gap between first and second place is 11.1 points, which is substantial on a 100-point scale and likely reflects differences in how boldly each model commits to named candidates and outcomes.

In practice, a higher score means the model is more useful for scenario planning. A lower score means the model hedges, omits key details, or fails to justify its confidence. For election forecasting, where uncertainty is inherent, a model that states its assumptions clearly is more valuable than one that refuses to commit.

What this means for practitioners

If you are choosing a model for election forecasting or any prediction task, do not look only at the top score. Look at the two axes. A model might be strong on US 2026 but weak on France 2027, or vice versa. The overall score hides that. Also consider the provider: all three models here come from opencode-zen, so the comparison is within one ecosystem. If you need a free option, Nemotron 3 Ultra and HY3 are both marked free, but Nemotron leads by a wide margin.

For teams building forecasting pipelines, the judge-based approach is a practical alternative to waiting for election night. It rewards models that produce structured, inspectable predictions. That makes it easier to audit and improve your own prompts.

How to read a leaderboard

  • Check the axes, not just the overall score. France 2027 and US 2026 are separate metrics.
  • Read the judge description. A fixed judge ensures consistency, but its rubric defines what good means.
  • Look for verbatim responses. Public outputs let you verify the scoring yourself.
  • Note the provider. Scores are only comparable within the same benchmark setup.
  • Treat scores as quality signals, not accuracy guarantees. Correctness is only verifiable after the vote.

Honest limitations

This benchmark measures prediction quality, not prediction accuracy. A model can score 89.1 and still be wrong about the 2027 French presidential election or the 2026 US midterms. The judge model introduces its own biases, and the three axes are a simplification of what makes a forecast useful. The leaderboard only includes three models, all from opencode-zen, so it is not a global ranking. Finally, election forecasting involves irreducible uncertainty; no score can eliminate that.

Conclusion

LFORLA's election forecasting benchmark offers a structured way to compare models on a task where ground truth is delayed. By scoring specificity, grounding, and calibration, it rewards models that commit to concrete scenarios. Nemotron 3 Ultra currently leads with 89.1, but the real value is in the mechanism. Explore the full results and methodology at lforla.org.

Top comments (0)