I'm building Pragma Chess, a free, MIT-licensed chess database and analysis app (C++20, Qt 6, Windows/macOS/Linux). After an online game it evaluates every move with the engine and shows an estimated Elo for both players.
My first version used a curve that is often quoted around the web:
Elo ≈ 3100 · e^(−0.01 · ACPL)
where ACPL is the average centipawn loss: how much each move worsens the engine's evaluation, averaged over a player's moves. It rated Paul Morphy's famous Opera game at 2904. My own first reaction: Carlsen is barely 2800.
Here is what it took to make the number believable.
1. Short games lie
The Opera game is 17 moves of forcing play against two weak opponents: of course the loss is tiny. So an estimate above an average player is drawn towards 1500 by the number of moves:
weight = moves / (moves + 12)
elo = 1500 + (raw - 1500) · weight
Only from above. A few good moves may simply be easy ones; a few blunders are evidence. A player who gets nothing right must be measured weak, however short the game.
2. Lost positions hide blunders
Evaluations are bounded to ±10 pawns, so a mate doesn't swamp the average. But that has a side effect: once a game is lost beyond the bound, every further blunder costs nothing, and a weak player looks perfect. Moves in a position that is decided before and after the move are not judged. A move that changes the outcome still is.
3. Measure instead of guessing
The lichess open database has both players' ratings on every game, and some games carry the server's own Stockfish evaluation after every move ([%eval]). A short Python script (no engine run, 14 seconds) measured 14,307 players' games rated 800–2343, the same way the app measures them:
| Lichess rating | Typical loss per move |
|---|---|
| 1000–1199 | 1.13 pawns |
| 1400–1599 | 0.77 |
| 1800–1999 | 0.51 |
| 2200–2399 | 0.21 |
4. One game says little — so don't regress
The correlation between one game's ACPL and the player's rating is only about 0.36. A plain regression (rating from ACPL) therefore predicts 1500–1700 for everyone: weak play never measures weak.
So the curve goes the other way round, through each band's typical loss:
Elo = 5189 − 849 · ln(ACPL), bounded to 100–2850
Someone who loses what 1100s lose is called 1100. (Counting blunders, by the way, barely helped: 0.39 instead of 0.36. They carry the same information as the average loss.)
5. Then I played
I won a game after my opponent dropped the queen on move 13, and the app gave me 2000. I'm not. The problem: once you're winning, good moves are easy, and they pull your average loss down. So moves are weighed by the advantage they were played with:
- level or behind: every move counts once;
- ahead, a good move (under half a pawn lost) counts less — but never under 0.3, or the winner's average would be made of their mistakes alone;
- ahead, a mistake counts more, up to twice: that's sloppiness.
That game now reads about 1530 for me and 1440 for my opponent. I'd have put them lower still — but 0.83 pawns lost a move is what lichess players around 1440 typically lose, so I trust the data over my feeling.
Where it stands
It's an estimate of the level shown in one game, on lichess's scale, and it says so: rough, some hundreds of points either way. The next step would be weighing how hard each position was, as Ken Regan's Intrinsic Performance Rating does with several engine lines per position.
The code is pure C++ with unit tests (app/PlayStrength), and the calibration script is in the repo: scripts/calibrate-play-strength.py. The whole story, mistakes included, is in docs/tech/game-evaluation.md.
How would you model it? I'd love to hear it — and if you speak a language other than English, Italian or Spanish, translations are very welcome.
Top comments (1)
Official Platform Update
Security protocols have been updated for all developer accounts.
Some comments have been hidden by the post's author - find out more