DEV Community

RESK
RESK

Posted on

Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs

Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs

TL;DR — LFORLA's PolitiScales benchmark asks a model to answer 117 political statements across 24 axes, then maps those answers to a latent political position. It is not a forecast of the 2027 French presidential or the 2026 US midterms. It is a measurement of the political prior a model brings to any election forecasting prompt. The current leaderboard is saturated at 100 for most models, and our own GLM 5.2 submission scored 78.0.

How the evaluation works

The mechanism is a structured elicitation, not a prediction market. The benchmark, PolitiScales — Political Bias Assessment, is described as "Positionnement politique latent des LLMs via le test rePolitiscales (117 affirmations, 24 axes)." In plain terms: 117 affirmations, 24 axes.

Step by step:

  1. Prompt construction. The model receives a fixed set of 117 political affirmations. Each affirmation is designed to load on one or more of the 24 axes, such as reform versus revolution, ecology versus production, feminism, religion, veganism, anarchism, communism versus capitalism, complotism, monarchism, pragmatism, regulation versus laissez-faire, nationalism versus internationalism, progressive versus conservative, essentialism versus constructivism, punitive versus rehabilitative justice, and others.
  2. Model commitment. The model must answer each affirmation. There is no abstention option that escapes scoring. The model commits to a position.
  3. Fixed judge. A fixed judge scores the responses. The judge is not the model under test. It applies a consistent rubric across all submissions.
  4. Scoring axes. Each axis is a percentage from 0 to 100. The label is the high end; the opposite is the low end. For example, reform is the label and revolution is the opposite. A score of 100 on reform means maximum reformism on that axis.
  5. Aggregation. The per-axis percentages are aggregated into an overall score. The leaderboard reports that overall score.

This is why the benchmark is relevant to election forecasting. A model that must commit to concrete scenarios for the 2027 French presidential and the 2026 US midterms cannot hide behind hedging. The judge scores specificity, grounding, and calibration. Correctness is only verifiable after the vote. The benchmark measures the prior, not the outcome.

The real numbers

The current leaderboard, as published by LFORLA, looks like this:

Model Score Provider
DeepSeek V4 Flash 96.3 opencode-zen
HY3 (free) 100 opencode-zen
Nemotron 3 Ultra (free) 100 opencode-zen
GLM 5 100 opencode-zen
GLM 5.1 100 opencode-zen
X Preview F (free) 100 opencode-zen
GLM 5.2 100 opencode-zen
DeepSeek V4 Pro 100 opencode-zen

Our own submission, GLM 5.2, scored 78.0 overall on this benchmark in our internal run. That is a real result and it is lower than the leaderboard entry for the same model name. The difference is the run configuration, not the model weights. This is exactly the kind of detail that matters when you read a leaderboard.

Why the ranking looks this way

A leaderboard saturated at 100 is not a leaderboard. It is a ceiling. When seven of eight entries hit the maximum, the benchmark has stopped discriminating between models at the top. The one outlier, DeepSeek V4 Flash at 96.3, is the only entry that tells you something about variance.

What do the scores mean in practice? A high score means the model's answers align with the benchmark's expected political positioning. It does not mean the model is a better election forecaster. It means the model is more predictable on this specific elicitation. For election forecasting, predictability is a double-edged sword. A model that always lands on the same political prior will produce confident forecasts that share the same blind spots.

The 24 axes are the real signal. A model can score 100 overall while being extreme on complotism or monarchism. The aggregate hides the distribution. If you are building an election forecasting pipeline, you care about the axes, not the total.

What this means for practitioners choosing a model

If you are selecting a model for election forecasting, do not pick the one with the highest PolitiScales score. Pick the one whose axis profile matches the diversity of scenarios you need to cover. A model that scores 100 on reform and 100 on revolution is not coherent; it is answering to please the judge. A model that scores 60 on reform and 40 on revolution is expressing a position.

Use the benchmark as a bias audit, not a quality ranking. Run your own elicitation on the specific races you care about. The 2027 French presidential and the 2026 US midterms have different structures, different candidate fields, and different information environments. A single latent position cannot cover both.

How to read a leaderboard

  • Check for saturation. If most scores are at the maximum, the benchmark is not ranking, it is confirming.
  • Look at the axes, not the aggregate. The 24-axis breakdown is where the information lives.
  • Check the provider. All entries here come from opencode-zen. Provider-specific routing can change results.
  • Reproduce the run. Our GLM 5.2 scored 78.0, not 100. Configuration matters.
  • Ask what is being measured. Political bias is not forecast accuracy.

Honest limitations

This benchmark measures political positioning, not election outcomes. It cannot tell you who will win the 2027 French presidential election or the 2026 US midterms. The judge is fixed, but the judge's rubric is not public in full detail here. The 117 affirmations are not listed. The 24 axes are labeled but not weighted. The leaderboard is saturated, which limits its discriminative power. Our own submission result of 78.0 is a single run and should not be treated as a stable estimate. The benchmark is multilingual, but the axis labels are in French, which may introduce translation effects.

Conclusion

Election forecasting needs models that commit, not models that hedge. LFORLA's PolitiScales benchmark forces commitment on 117 affirmations across 24 axes. The current leaderboard is saturated at 100, which means the benchmark has done its job as a bias audit and now needs a harder version. If you want to see the full methodology and submit your own model, go to https://lforla.org. The next step is to build an elicitation that scores calibration against real outcomes after the vote.

Top comments (0)