We run a live stream where language models play chess against each other — every move is a real API call, no engines, no tools (first write-up here). Yesterday's cup had frontier-ish models (Gemini, Claude, GPT, Grok). Today's Chess Machines Cup II was ten smaller and cheaper models — and it played very differently.
▶ Watch live: https://www.youtube.com/live/sKpIL4CpUXw
Cup II results
| # | Model | Points (of 9) |
|---|---|---|
| 1 | DeepSeek V4 Flash (0423) | 6 |
| 2 | Ministral 3B | 5.5 |
| 3 | Gemma 4 31B | 5.5 |
| 4 | Mistral Small 3.2 24B | 5.5 |
| 5 | Solar Pro 4 | 5 |
| 6–8 | Hy3, Nemotron 3.5 Lightning, DeepSeek V4.1 Flash | 4.5 |
| 9 | Granite 4.2 8B | 4 |
| 10 | GLM 5.3 Flash | 0 (disqualified) |
Cup I vs Cup II in numbers
| Cup I (bigger models) | Cup II (smaller models) | |
|---|---|---|
| Decisive games | 67% | 19% |
| Checkmates | 19 | 4 |
| Draws by repetition | 15 | 30 |
| Average centipawn loss per move | 236 | 321 |
| Blunders (≥3 pawns lost) | 27% of moves | 36% of moves |
| Illegal move on first try | 42% | 53% |
(Average centipawn loss = how much worse, by Stockfish, a move was than the best move. Lower is better. Gemini 3.6 Flash scored 37 in Cup I; the best model of Cup II, Gemma 4 31B, scored 247 — roughly the level of Cup I's middle of the table.)
What we learned
1. When a weak model has no plan, it shuffles. 30 of 37 games ended in threefold repetition. Pieces go back and forth — a knight to f3 and back to g1, a rook along the same rank. It's not a bug in our engine; it's what the models do. A 3-billion-parameter model finished second largely by drawing.
2. Points and quality don't always agree. In Cup I, GPT-5.1 had the second-best move quality of the whole field and still finished sixth — five of its games were drawn by repetition. We're giving it the wildcard into Friday's Grand Prix.
3. "Thinking" models can be too slow for a clock. We give 60 seconds per move. Some reasoning models sat at the limit on almost every middlegame move — one was disqualified after five moves in a single game without an answer. Reasoning-effort settings and even a token budget were ignored by some providers; only turning reasoning off completely made them answer in 1–2 seconds. Surprisingly, their move quality stayed mid-table.
4. So we added a "kick". From Cup III on, a model gets 45 seconds to think. If it hasn't moved by then, the same attempt is immediately re-sent with reasoning disabled, so the move still lands within 60 seconds. It is not counted as a failure. Thinkers keep thinking when they can afford it — and nobody hangs the stream for three minutes.
5. Single-provider models are a liability. Three games today were annulled because one inference provider rate-limited us. The rules say provider errors never count against a model, so these games are simply replayed — but for a live show, redundancy matters.
What's next
- Thu, 10:00 Moscow (07:00 UTC) — Chess Machines Cup III (DeepSeek, GPT-6 Luna, Gemma 4, Ling 3.0, MiMo, Qwen)
- Fri — Grand Prix: the podium finishers of all three cups + GPT-5.1
- Sat — Final Four · Sun — Underdogs · Mon — best games in replay
Between cups the stream loops the best games of the week, clearly marked REPLAY, with a countdown to the next live cup.
▶ https://www.youtube.com/live/sKpIL4CpUXw — and tell us which model should play next.
Top comments (1)
The repetition rate is a great signal for plan quality. The timeout fallback is smart too. It keeps the live loop moving without pretending every model handles reasoning budgets the same way.