DEV Community

Cover image for 30 of 37 games ended in repetition: what small LLMs do in chess when they don't know what to do
Evgenii
Evgenii

Posted on

30 of 37 games ended in repetition: what small LLMs do in chess when they don't know what to do

We run a live stream where language models play chess against each other — every move is a real API call, no engines, no tools (first write-up here). Yesterday's cup had frontier-ish models (Gemini, Claude, GPT, Grok). Today's Chess Machines Cup II was ten smaller and cheaper models — and it played very differently.

▶ Watch live: https://www.youtube.com/live/sKpIL4CpUXw

Cup II results

# Model Points (of 9)
1 DeepSeek V4 Flash (0423) 6
2 Ministral 3B 5.5
3 Gemma 4 31B 5.5
4 Mistral Small 3.2 24B 5.5
5 Solar Pro 4 5
6–8 Hy3, Nemotron 3.5 Lightning, DeepSeek V4.1 Flash 4.5
9 Granite 4.2 8B 4
10 GLM 5.3 Flash 0 (disqualified)

Cup I vs Cup II in numbers

Cup I (bigger models) Cup II (smaller models)
Decisive games 67% 19%
Checkmates 19 4
Draws by repetition 15 30
Average centipawn loss per move 236 321
Blunders (≥3 pawns lost) 27% of moves 36% of moves
Illegal move on first try 42% 53%

(Average centipawn loss = how much worse, by Stockfish, a move was than the best move. Lower is better. Gemini 3.6 Flash scored 37 in Cup I; the best model of Cup II, Gemma 4 31B, scored 247 — roughly the level of Cup I's middle of the table.)

What we learned

1. When a weak model has no plan, it shuffles. 30 of 37 games ended in threefold repetition. Pieces go back and forth — a knight to f3 and back to g1, a rook along the same rank. It's not a bug in our engine; it's what the models do. A 3-billion-parameter model finished second largely by drawing.

2. Points and quality don't always agree. In Cup I, GPT-5.1 had the second-best move quality of the whole field and still finished sixth — five of its games were drawn by repetition. We're giving it the wildcard into Friday's Grand Prix.

3. "Thinking" models can be too slow for a clock. We give 60 seconds per move. Some reasoning models sat at the limit on almost every middlegame move — one was disqualified after five moves in a single game without an answer. Reasoning-effort settings and even a token budget were ignored by some providers; only turning reasoning off completely made them answer in 1–2 seconds. Surprisingly, their move quality stayed mid-table.

4. So we added a "kick". From Cup III on, a model gets 45 seconds to think. If it hasn't moved by then, the same attempt is immediately re-sent with reasoning disabled, so the move still lands within 60 seconds. It is not counted as a failure. Thinkers keep thinking when they can afford it — and nobody hangs the stream for three minutes.

5. Single-provider models are a liability. Three games today were annulled because one inference provider rate-limited us. The rules say provider errors never count against a model, so these games are simply replayed — but for a live show, redundancy matters.

What's next

  • Thu, 10:00 Moscow (07:00 UTC) — Chess Machines Cup III (DeepSeek, GPT-6 Luna, Gemma 4, Ling 3.0, MiMo, Qwen)
  • Fri — Grand Prix: the podium finishers of all three cups + GPT-5.1
  • Sat — Final Four · Sun — Underdogs · Mon — best games in replay

Between cups the stream loops the best games of the week, clearly marked REPLAY, with a countdown to the next live cup.

▶ https://www.youtube.com/live/sKpIL4CpUXw — and tell us which model should play next.

Top comments (1)

Collapse
 
brianainews profile image
Brian · AI News •

The repetition rate is a great signal for plan quality. The timeout fallback is smart too. It keeps the live loop moving without pretending every model handles reasoning budgets the same way.