DEV Community

Cover image for I Built an AI Chess Tutor Inspired by Wizard's Chess, Here's How It Works
Samuel Oziegbe
Samuel Oziegbe

Posted on

I Built an AI Chess Tutor Inspired by Wizard's Chess, Here's How It Works

I'm a chess enthusiast and also a Harry Potter fan. Right from when I was a kid and I saw the scene in the very first part of the Harry Potter franchise (The Philosopher's stone), I just thought it to be cool and that I'd want something like that for myself. Basically a chess game where I could just speak out loud and my commands would be executed on the board, almost like magic.

So, I put my AI engineering skillset to work and the result is thechesstutor.com, a chess app where you play against an AI engine, get real-time coaching from an AI tutor, and can control the board entirely by voice.

This post is a breakdown of how I built it, the architecture decisions I made, and some issues I faced along the way.

The Idea

I wanted more than a chess engine. Chess engines are everywhere. What I wanted was something that could:

Play against me
Explain why a move is good or bad
Compare moves
Let me ask questions mid-game ("how should I develop my bishop from here?" or "do you think I should take their knight with my pawn or queen?")
Accept voice commands the way wizard's chess pieces accept commands from their player in Harry Potter, for example Ron shouting at the top of his voice "Castle to E4", still gives me the chills, yeah I'm a nerd, I know 😅😅.

The Stack

Frontend: React + Vite
Backend: FastAPI
Agent Framework: LangGraph
LLM: OpenAI (GPT-4.1)
Chess Engine: Stockfish
Voice: Whisper (STT) + Web Speech API (TTS)
Background Jobs: Celery + Redis
Database: PostgreSQL
Monitoring: LangSmith
Infrastructure: Docker + Docker Compose

The Architecture

The architectural diagram

The heart of the app is a multi-agent system built with LangGraph. There are three agents:

The Orchestrator Agent which is kind of like the brain. It receives every user query, creates a plan for solving it, and decides which agent(s) to delegate to.

The Move Agent that is responsible for executing moves on the board. It validates that a move is legal, executes it, and returns the resulting board state.

The Tutor Agent, this is the chess coach. It answers questions, suggests moves, analyzes positions, and explains strategy. It has access to Stockfish through tools.

The orchestrator + worker agents

The orchestrator routes queries like "play e4" to the move agent, queries like "what's the best move?" to the tutor agent, and complex queries like "find the best move and play it" to both agents in sequence. It knows never to call an agent whose result already exists in the conversation history, a guardrail enforced both at the prompt level and at the graph edge level in code.

Voice Mode (the feature that gives the Wizard's chess feel)

The chess tutor voice mode

Checkout the voice mode demo

At the bottom right you can see the AI listening to you and speaking back.

Voice mode puts you in constant conversation with the AI, you can leave it on and talk naturally, giving move commands, asking for explanations, or just thinking out loud about your position. If you prefer one-off messages, hold the space bar to record and release to send.

So you can literally say "move my knight to C3" or "destroy their queen" and watch the piece slide across the board to the described position.

Background Move Analysis with Celery + Redis

I implemented real-time game analysis. So basically, after every move, the app classifies it as a blunder, mistake, inaccuracy, or good move and then displays your accuracy stats on the left panel.

The naive approach would be to run this analysis when the user asks for it. But that means running Stockfish evaluations for every move in the game history on demand. Depending on how far into the game and how many moves already played, this can get really slow and would keep the user waiting for quite a while.

Instead, I run the analysis in the background after every move using Celery and cache the result in Redis. The flow looks like this:

Player makes a move
Frontend fires a POST to /analyze/{game_id} — fire and forget, no awaiting
FastAPI queues a Celery task immediately and returns
The Celery worker picks up the task, runs Stockfish evaluation on only the latest move (not the full history), and appends the result to the cached analysis in Redis.
When the user asks "where did I go wrong?", the tutor agent reads from Redis instantly

The incremental analysis was key, instead of re-evaluating every move from scratch each time, I track how many moves have already been analyzed and only compute the difference.

The Hallucination Problem

One thing that genuinely shocked me was the high level of hallucination I faced while working with the LLM.

My tutor agent kept confidently describing pieces that weren't there. It would tell me my queen was on D4 when she was clearly on d1. It would say I had no bishop when my bishop was sitting right there on f1, I almost went crazy at some point 😂.

The root cause is that FEN strings, the standard notation for board positions require spatial reasoning that LLMs struggle with. The model has to parse rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR and mentally construct a board from it. That's a lot of inference with a lot of room for error.

Here's how I reduced hallucinations significantly:

i) Send a pre-parsed piece map alongside the FEN

Instead of just the FEN string, I also send a dictionary of square-to-piece mappings:

json
{
"e1": "K",
"d1": "Q",
"a2": "P"
}

The model no longer has to parse the FEN, it can just look up what's on any square directly.

ii) Force tool use before any positional claim

I added a hard rule to the tutor agent's prompt: before making any claim about what pieces are on the board or what moves are available, it must call get_legal_moves. This grounds every positional statement in actual engine output rather than the model's potentially faulty FEN interpretation.

iii) Upgrade the tutor model

I use GPT-4.1-mini for the orchestrator and move agent since speed matters there. But for the tutor agent I switched to GPT-4.1, which is noticeably better at staying grounded to structured data.

The combination of all three significanly reduced the hallucination issues that were encountered.

The UI

The chess tutor text mode

The layout is split into three sections:

Left panel: player profile, captured pieces, and game stats (accuracy, good moves, mistakes, blunders)
Center: the chess board, the move history and the AI strength indicator.
Right: the sidebar containing the chat and voice controls

The tutor renders board positions inline in the chat as a Unicode chess board so you can see exactly what position it's describing without leaving the conversation.

Monitoring with LangSmith

Every agent call is traced through LangSmith, which gives me visibility into what the orchestrator planned, which agent it delegated to, what tools were called, and how long each step took. This was invaluable for debugging the hallucination issues, I could see exactly what the model was seeing and where it was going wrong.

My next steps

Adding an ELO rating system, right now players are "unrated." I want to implement a proper rating system that adjusts based on game results against the AI at different Stockfish depths.
Opening explorer, I want it such that when you play a recognized opening, the tutor should name it and give you context about the theory behind it.
Multiplayer, two human players with a shared AI tutor commenting on the game.

I do hope you try it out, remember its thechesstutor.com, if you have any questions or comments please do not hesitate 😁.

Top comments (0)