What happens when six small AI models — four of them running on a laptop — are asked to solve the same scrambled cube, one move at a time?
Most "AI plays a game" demos hide the interesting part. They show a model winning, and you are left to guess how much of that was the model and how much was the harness. This project does the opposite: it puts six models in a glass box, gives them the exact same scrambled Rubik's cube, the exact same question every turn, and shows you — live, in one screen — exactly where each one succeeds and where it falls apart.
The result is uncomfortable and honest. A Rubik's cube is a planning problem: you need lookahead and search. The models in this arena are decision models: fast, calibrated judgments about the current state. They are not the same thing, and the arena makes that difference visible.
1. The premise: decision models vs. planning
A "decision model" (sometimes called a system-one model) does one thing well: it looks at a situation, weighs a closed set of options, and returns a choice with a confidence. It is fast and it is often surprisingly well-calibrated.
A Rubik's cube is not a decision problem. It is a search problem. From a scrambled state there is an optimal solution length, and finding it means exploring a tree of moves. A model that only ever looks at the current state — with no memory of the plan and no lookahead — is being asked to do something it was not built for.
So the arena asks a precise question:
Given the same cube and the same one-move question, how good are six real decision models at picking moves that actually help?
Everything else in the project exists to answer that question fairly.
2. The two hard rules
The whole comparison is only meaningful if the conditions are identical, so the code enforces two rules everywhere:
- Every engine gets the identical state and the identical question. No engine gets extra context, a better prompt, retries, or a second look. Same 54 stickers, same 18 options, same instructions.
-
Nothing is ever faked. If a model or its server is missing, its panel reads
unavailablewith the exact reason, and the race simply runs without it. There are no placeholder scores and no synthetic results.
That second rule sounds obvious until you build a demo. It is very tempting to fill a dead panel with something plausible. The arena refuses.
3. Architecture
┌──────────────────────────────────┐
│ browser http://127.0.0.1:8077 │
│ ui/index.html (CSS-3D cubes) │
└───────────────┬──────────────────┘
GET / GET /api/engines │ POST /api/run
GET /api/run/<id> GET /api/chart/<id>
│
┌───────────────▼──────────────────┐
│ server.py (stdlib HTTP) │
│ Hub: loaded engines + active run │
│ CHART_LOCK serialises matplotlib │
└───────┬───────────────────┬───────┘
│ │
one thread per engine (run in parallel) │
▼ ▼
┌────────────────────┐ ┌─────────────────────────────┐
│ arena.run() │ │ charts.build_scorecard() │
│ guided | greedy │ │ -> results/chart_<id>.png │
│ | staged │ │ progress · latency · │
└─────────┬──────────┘ │ guided · final │
│ └─────────────────────────────┘
each turn: decide(state, question)
▼
┌────────────────────────────────────────────────────────────────┐
│ engines.py — one contract, six adapters │
│ laya MLX, in-process gliner CPU, in-process │
│ strands HTTP 127.0.0.1:8010 kev HTTP 127.0.0.1:8008
│ clef / clef-flash Cloudflare Workers AI (HTTPS) │
└────────────────────────────────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────┐
│ cube.py (pycuber state + metrics) │
│ solver.py (kociemba plan / first move) │
└────────────────────────────────────────────────────────────────┘
The stack is deliberately boring: a stdlib HTTP server, a single HTML file, and plain Python. There is no framework, no bundler, and no build step. You can read the whole thing in an afternoon.
Tech stack
| Layer | Choice |
|---|---|
| Cube simulation | pycuber |
| Optimal search |
kociemba (two-phase) |
| Local runtimes | MLX (laya), torch/MPS (strands), transformers/CPU (gliner2) |
| Cloud inference | Cloudflare Workers AI (clef, clef-flash) |
| HTTP client | httpx |
| Web server | Python stdlib http.server + ThreadingHTTPServer
|
| UI | single-file HTML + CSS 3D transforms + vanilla JS |
| Charts |
matplotlib (Agg) |
| Terminal | rich |
| Automation | Playwright |
| Env |
uv + Python 3.12 |
4. The cube model
Everything starts with a deterministic scramble. Given a seed, every engine sees the same cube — and the same cube every time you rerun the demo.
# cube.py
FACES = "URFDLB"
MOVES = [f + s for f in FACES for s in ("", "'", "2")] # 18 quarter/half turns
def scramble(seed: int, n: int = 25) -> tuple[Cube, list[str]]:
rng = random.Random(seed)
seq: list[str] = []
prev_face = None
while len(seq) < n:
mv = rng.choice(MOVES)
if mv[0] == prev_face: # avoid trivial same-face repeats
continue
prev_face = mv[0]
seq.append(mv)
c = Cube()
for mv in seq:
c.perform_algo(mv)
return c, seq
Two metrics are tracked. A sticker is correct when it matches its own face's centre colour (centres never move, so the centre is the solved colour). That gives a smooth 0–54 progress line; complete faces give a chunky 0–6 milestone.
# cube.py
def solved_facelets(cube: Cube) -> int:
n = 0
for f in FACES:
ctr = centre_colour(cube, f)
face = _face(cube, f)
for r in range(3):
for c in range(3):
if face[r][c].colour == ctr:
n += 1
return n
The model never sees a Cube object. It sees a compact text state, identical for every engine:
# cube.py
def state_text(cube: Cube) -> str:
rows = facelet_rows(cube)
body = "\n".join(f"{f}: {rows[f]}" for f in FACES)
return (
"3x3 Rubik's cube. Faces U R F D L B, each 9 stickers read row by row "
"seen from outside. Letters are sticker colours (W Y G B R O). A face is "
"solved when all 9 stickers match its fixed centre sticker.\n"
f"{body}\n"
f"Correct stickers: {solved_facelets(cube)}/54. "
f"Complete faces: {solved_faces(cube)}/6."
)
5. The solver bridge (and its encoding trap)
Guided mode needs an optimal-ish solver. The project uses kociemba, but wiring it up has a classic gotcha: pycuber colours its stickers with colour names, and Kociemba wants a 54-character string of face letters in URFDLB order.
Hardcoding the colour→face map is wrong, because pycuber's default cube is not the usual white-up arrangement. The fix is to read the map off the centres, which never move:
# solver.py
FACES = "URFDLB"
def colour_map(cube: Cube) -> dict[str, str]:
"""colour name -> face letter, taken from the fixed centres."""
return {cube.get_face(f)[1][1].colour: f for f in FACES}
def facelet_string(cube: Cube) -> str:
"""54-char Kociemba facelet string, URFDLB order, row-major from outside."""
m = colour_map(cube)
out: list[str] = []
for f in FACES:
face = cube.get_face(f)
for r in range(3):
for c in range(3):
out.append(m[face[r][c].colour])
return "".join(out)
@lru_cache(maxsize=200_000)
def _solve_cached(state: str) -> str:
return kociemba.solve(state)
Every state is cached by its facelet string. The race revisits states constantly (neutral moves, repeated positions), and Kociemba is by far the slowest part of the loop — the cache is the difference between a snappy demo and a stall.
Two more things worth knowing if you build on this:
-
kociembahas no Apple-silicon wheel, so it always compiles a small C extension. If that build fails, the module degrades cleanly: greedy and staged modes still work, and guided mode says exactly what is missing. - The solver's solution length is near-optimal, not provably optimal. That is why guided runs take
optimal + a fewmoves rather than exactlyoptimal.
6. One contract, six engines
This is the heart of the project. Six very different runtimes — an in-process MLX model, a torch server, a CPU classifier, two cloud endpoints, and a local OpenAI-ish server — are all reduced to one method:
# engines.py
def decide(self, state: str, question: dict, qid: str = "move") -> Decision:
...
@dataclass
class Decision:
engine: str
choice: str # e.g. "R2"
confidence: float | None
probabilities: dict[str, float] # over the 18 moves (may be empty)
latency_ms: float
error: str | None = None
The question is a closed choice over the 18 legal moves, with a description for each — the shape that strands-decider requires, and that the others accept too:
# engines.py
_FACE_NAME = {"U": "up", "R": "right", "F": "front",
"D": "down", "L": "left", "B": "back"}
MOVE_CRITERIA = {}
for _f, _n in _FACE_NAME.items():
MOVE_CRITERIA[_f] = f"turn the {_n} face clockwise 90 degrees"
MOVE_CRITERIA[_f + "'"] = f"turn the {_n} face anticlockwise 90 degrees"
MOVE_CRITERIA[_f + "2"] = f"turn the {_n} face 180 degrees"
def move_question(instructions: str) -> dict:
return {"type": "choice", "instructions": instructions,
"criteria": dict(MOVE_CRITERIA)}
Most engines share an HTTP transport, so adding one is usually a tiny subclass:
# engines.py
class HttpEngine(Engine):
def probe(self) -> tuple[bool, str | None]:
"""A dead server must read as unavailable, not as 'ready'."""
if not self.health_url:
return True, None
try:
r = httpx.get(self.health_url, timeout=3.0)
except Exception as e:
return False, f"no server at {self.health_url} ({type(e).__name__})"
return True, None
def decide(self, state, question, qid="move") -> Decision:
body = {"state": state, "model": self.model, "questions": {qid: question}}
t0 = time.perf_counter()
try:
r = self.client.post(self.url, json=body)
r.raise_for_status()
resp = r.json()
except Exception as e:
return Decision(self.name, "", None,
latency_ms=(time.perf_counter() - t0) * 1000,
error=f"{type(e).__name__}: {e}")
dt = (time.perf_counter() - t0) * 1000
choice, conf, probs = _parse_answers(resp, qid)
return Decision(self.name, choice, conf, probs, dt)
The probe() method is why a dead Kev server shows unavailable: no server at 127.0.0.1:8008 instead of silently reporting "ready" and failing on every move.
The six models
| Model | Params | Where | How |
|---|---|---|---|
Laya (aac6fef/laya-mlx) |
421M | local, Apple silicon | in-process MLX |
| Strands Decider 2B | 1.9B | local, torch/MPS | HTTP POST /v1/systemone
|
| GLiNER2.5-Decide | 340M | local, CPU | in-process classify_text
|
| Clef-Flash | 9B | Cloudflare Workers AI | HTTPS |
| Clef | 27B | Cloudflare Workers AI | HTTPS |
| Kev-0.8B | 0.8B | local server :8008
|
HTTP POST /v1/systemone
|
Why two run in the cloud: this was built on an Apple M2 mini with 8 GB of RAM. A 27B model does not fit at any quantization, so Clef and Clef-Flash are called on Workers AI where they cost no local RAM. The other four run locally — including on a 16 GB laptop — which is the point of the demo.
7. The guided loop: the solver plans, the model decides
The unaided models cannot solve a cube. That is measured, not assumed: after 24 moves each, three different architectures all landed on exactly the sticker count they started from. So guided mode separates the two jobs cleanly — the solver is the search, the model is the decision layer inside it.
# arena.py (abridged)
plan = solution(cube) # a concrete, known-good move list
correct = plan[0] # the solver's optimal next move
d = engine.decide(state_text(cube), move_question(GREEDY_Q))
mv = pick_move(d, policy, rng) # sample the distribution, or take argmax
play, tag, replan = correct, "corrected", False
if mv == correct:
tag = "hit" # matched the solver exactly
elif mv in MOVES:
probe = cube.copy(); apply_move(probe, mv)
d_m = len(solution(probe))
if d_m < len(plan):
play, tag, replan = mv, "accepted", True # a different move that helped
elif d_m == len(plan) and neutral_budget > 0:
play, tag, replan = mv, "neutral", True # wasted turn (budgeted)
neutral_budget -= 1
else:
play, tag = correct, "corrected" # solver overrides
else:
play, tag = correct, "corrected" # unusable answer
apply_move(cube, play)
plan = solution(cube) if replan else plan[1:]
The model gets the same 18-move question as the unaided race, so the two are directly comparable. Each pick is scored:
| Outcome | Meaning |
|---|---|
hit |
the model picked exactly the solver's next move |
accepted |
a different move that still made strict progress |
neutral |
the move did not change the distance — a wasted turn (budgeted) |
corrected |
the model's move was worse or unusable, so the solver played its own move |
The cube is guaranteed to advance because the solver always holds a valid plan, and after the neutral budget is spent every move must strictly reduce the distance.
But here is the subtlety that trips people up: a corrected turn still costs one of the moves you allocated. The solver is doing the solving, and the model is being measured — but the clock runs on every turn, including the ones where the model was overruled. With an optimal distance of 22 and a neutral budget of 6, a guided run can legitimately need up to 28 moves. Set the cap below that and slower models get cut off unsolved while still showing the done status.
Policy: how a distribution becomes a move
Most of these models return a probability for each of the 18 moves. Policy decides what to do with it:
# arena.py
def pick_move(d, policy: str, rng: random.Random) -> str:
if policy == "sample" and d.probabilities:
opts = [(m, v) for m, v in d.probabilities.items() if m in MOVES and v > 0]
if opts:
total = sum(v for _, v in opts)
r = rng.random() * total
acc = 0.0
for m, v in opts:
acc += v
if r <= acc:
return m
return opts[-1][0]
return d.choice # argmax
-
sampledraws the move from the model's own distribution. This is the honest way to run a decision model in a loop, and it prevents the collapse you get from always taking the top choice. -
argmaxalways takes the most likely move. In practice it collapses — the model repeats one move forever — so it is shown mainly for contrast.
8. The server: threads, polling, and a chart race
The server is stdlib-only. Its job is small: hold the loaded engines, run one race at a time, and serve the UI. The interesting part is concurrency.
A run spawns one thread per ready engine, so all six models think in parallel:
# server.py
def _drive(self, run: dict) -> None:
threads = []
for panel in run["panels"]:
if panel["status"] != "ready":
continue
t = threading.Thread(target=self._one_engine, args=(run, panel), daemon=True)
t.start()
threads.append(t)
for t in threads:
t.join()
# ... persist the run + render the chart, THEN flip status to "done"
Every turn, the engine thread fires an on_move callback that updates the shared panel under a lock. The browser polls /api/run/<id> every 350 ms and redraws the cubes.
There is one nasty bug worth calling out, because it is the kind of thing that only shows up in a demo: matplotlib's pyplot is not thread-safe. The scorecard is built on the run thread when a race finishes, and on demand by the /api/chart/<id> route. If those two ever overlap, the render corrupts and the browser shows a broken image.
The fix is two-fold — a lock, and getting the order right:
# server.py
CHART_LOCK = threading.Lock()
# ... in _drive(), after all engine threads join:
snapshot = {**run, "status": "done", "finished_at": finished_at}
(RESULTS / f"ui_run_{run['id']}.json").write_text(json.dumps(snapshot, indent=2))
with CHART_LOCK:
build_scorecard(run, RESULTS / f"chart_{run['id']}.png")
with self.lock:
run["status"] = "done" # flipped LAST, only after the PNG exists
The status flips to done after the PNG is on disk, so the page can never request a chart that does not exist yet. The front-end adds a retry with a cache-buster as a belt-and-braces measure.
The API is four routes:
| Route | Purpose |
|---|---|
GET / |
the single-page UI |
GET /api/engines |
per-engine load status (polled on startup) |
POST /api/run |
start a race {seed, mode, policy, moves, engines}
|
GET /api/run/<id> |
current run state (polled while running) |
GET /api/chart/<id> |
the scorecard PNG |
9. The UI: a real 3D cube with no library
The cube is not an image or a canvas — it is six DOM faces positioned in 3D space with CSS transforms, slowly rotating. That is the whole trick:
/* ui/index.html */
.scene { height: 164px; perspective: 1000px; perspective-origin: 50% 45%; }
.cube3d {
position: relative; width: var(--cube); height: var(--cube);
transform-style: preserve-3d;
animation: spin 46s linear infinite;
}
.face3d.U { transform: rotateX(90deg) translateZ(calc(var(--cube)/2)); }
.face3d.D { transform: rotateX(-90deg) translateZ(calc(var(--cube)/2)); }
.face3d.F { transform: translateZ(calc(var(--cube)/2)); }
.face3d.B { transform: rotateY(180deg) translateZ(calc(var(--cube)/2)); }
.face3d.L { transform: rotateY(-90deg) translateZ(calc(var(--cube)/2)); }
.face3d.R { transform: rotateY(90deg) translateZ(calc(var(--cube)/2)); }
Each face is a 3×3 grid of stickers whose colours come straight from the model's net (the per-face colour strings). The layout is three columns by two rows, with a short-screen media query that shrinks the cube so all six stay in one frame.
Two front-end details earned their keep during development:
-
Persistence. When a run finishes,
runIdis cleared and the engine poll resumes. The first version re-rendered empty placeholders and every cube vanished at the moment of victory. The fix keeps the last run's panels and re-renders those instead, so the final solved state stays on screen. -
The chart is the true "done" signal.
startRace()is async, so the status text briefly still reads "models ready". Tests wait on the chart image actually loading (naturalWidth > 0), not on the status string.
10. The scorecard
When a race finishes, the server renders a four-panel dark scorecard with
matplotlib:
- Progress — correct stickers over moves, one line per model.
- Latency — median ms per move, log scale.
-
Guided resolution — stacked
hit/accepted/neutral/correctedbars, so you can see at a glance how much of the solve was the model. - Final — stickers correct at the end, solved models marked.
Colours are stable per engine key, so a model keeps its colour across runs.
11. What we actually measured
Short version; the full write-up is in REPORT.md.
Unaided (greedy / staged), same scramble, seed 20261002: after 24 moves each, three different architectures landed on exactly the sticker count they started from — zero net progress, and no face ever completed. Even 60 moves of subgoal-named play (staged mode) topped out at 17 of 54 stickers.
Guided mode, same scramble (optimal distance 22): every model solves the cube, because the solver is doing the solving. The honest reading is in the counters:
| model | solved | moves | hits | accepted | neutral | solver corrected |
|---|---|---|---|---|---|---|
| laya-mlx | yes | 28 | 0 | 1 | 6 | 21 |
| strands-decider | yes | 28 | 0 | 2 | 6 | 20 |
| GLiNER2.5-Decide | yes | 27 | 4 | 0 | 5 | 18 |
Read it carefully:
- 18 to 21 of every 27–28 moves were the solver overriding the model. laya and strands never once picked the solver's move.
-
GLiNER's 4 hits are an artefact, not skill. It returns a single label with no distribution, so it cannot be sampled — it plays
Ualmost every turn, and the optimal plan happened to start withUfour times. - Wall clock is the real race. Same solve, but laya finishes in ~5 s against strands' ~64 s and GLiNER's ~95 s.
The headline is not "who won". It is that small decision models do not plan, and the arena shows exactly where that breaks down — while never pretending the model did something it didn't.
12. Hard-won lessons
A few things that cost real debugging time and are worth stealing:
-
kociembahas no Apple-silicon wheel. It always compiles a C extension. Keep it in its ownpip installline, or a failed build aborts the whole command and takespycuberdown with it. -
matplotlib
pyplotis not thread-safe. Serialise every render with a lock, and write the PNG before you advertise that the run is done. -
A dead server must be probed, not assumed.
probe()is what turns a silent failure into an honestunavailablepanel. - Never wipe the final state on the next poll. Keep the last run's panels and re-render those, or the cubes disappear the instant the race ends.
-
The move cap is part of the experiment. In guided mode a
correctedturn still costs a move, so a cap belowoptimal + neutral_budgetsilently cuts models off unsolved. -
argmaxon a decision model collapses. Sampling from the returned distribution is the honest way to run a model in a loop. - Build on the centre stickers, not on assumed colours. pycuber's default cube is not the usual white-up arrangement, so derive the colour→face map from the fixed centres.
13. Run it, fork it, extend it
git clone https://github.com/harishkotra/cube-arena.git && cd cube-arena
uv venv --python 3.12 && source .venv/bin/activate
uv pip install laya-mlx strands-decider gliner2 pycuber rich matplotlib requests httpx
uv pip install kociemba
python server.py # open http://127.0.0.1:8077
Everything is optional except pycuber: a model without credentials or a running server simply shows unavailable. ./run.sh starts the strands server and the UI together; ./run_kev.sh starts the Kev server.
There is no framework and no build step, so contributing is easy. The most common contribution is a new engine: subclass HttpEngine (or Engine), register it in ENGINE_SPECS in server.py, and the UI, charts, and CLI pick it up automatically.
Other good first contributions: per-move regret and calibration metrics, a
best-of-N mode, an OpenAI-compatible adapter, a /api/runs history index, and a multi-seed suite. The full list is in the README's "Feature ideas" section.
If you build something on top of this, I'd love to see it. The whole point is that comparison under identical conditions is more interesting than a leaderboard.
Screenshots
Watch
Code & more: https://www.dailybuild.xyz/project/277-cube-arena




Top comments (0)