This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
The tower says: "Foxtrot Alpha Bravo Charlie Delta, Isola Ground, runway two seven, QNH one zero one three, taxi holding point Alpha One via Bravo." You hold the space bar and read it back, and you say one zero zero three. The tower comes straight back: "Foxtrot Alpha Bravo Charlie Delta, negative, QNH one zero one three."
That tower is a synthetic voice running on a laptop, the aerodrome doesn't exist, and the whole thing is a small game I built this weekend for a friend.
What I Built
My friend just got his private pilot's licence. To fly abroad he needs an English language proficiency endorsement on it, and he's wondering whether to go for it, but he's shy about speaking English on the radio. I passed mine two years ago, which doesn't make my English good enough to coach anyone: I am just a private pilot myself. The plan is simple: when he has it, we fly abroad together.
He knows the procedures. What he hasn't done much is say them in English, out loud, in the right order, with the numbers right, while someone waits for the answer. You can read the phraseology manual on the sofa. You can't rehearse the radio alone, because nobody talks back. This won't prepare him for the proficiency check; it gives him somewhere to say the words, alone, as many times as he wants, before he says them to someone.
So I built him something that talks back: phraseology, a practice game for one traffic circuit at Isola, a made-up controlled aerodrome. You're F-ABCD at the flying club. You call Ground, taxi to the holding point, change to Tower, line up, take off, fly a touch and go, then a second circuit to land, vacate and taxi back. Runway, wind, QNH, squawk and the other traffic change every game, and up to three surprises can change the plan: "number 2, follow the Cherokee on base", "extend downwind", etc. A second aircraft shares the frequency, with its own voice.
You answer with the space bar as your press-to-talk. After each transmission the page tells you what was missing or wrong, in plain words:
- "QNH: you said QNH 1003, the tower said QNH 1013."
- "Missing: runway 27 (the runway in use is always read back)."
- "Put your callsign at the end of a readback."
There is no score and no level. If the readback is wrong, the tower says "negative" and gives the correct part, and you read it back again.
It's a practice game, not training. It isn't an approved training device or an assessment, and it doesn't prepare anyone for a language proficiency check. The aerodrome, frequencies and traffic are invented, and no real ATC recording is used. Phraseology also differs from one country to another. Your instructor and your country's official manuals are the reference.
Demo
The video is one real game, played on a small demo server and recorded in a headless browser: speech recognition, the checker and the tower are the real ones. The pilot is synthetic too, Kokoro's French voice reading English, so you hear the whole loop with an accent (I have to say I am too shy myself to publish a video with my own voice in this context!). The waits while the server transcribes are cut.
Code
phraseology
Practice the radio calls of a traffic circuit in English, out loud, on your own laptop.
The tower talks to you with a synthetic voice. You hold the space bar like a press-to-talk switch and answer. An open-weight speech model transcribes you on your machine, a set of plain rules checks your readback, and the circuit moves on. If the readback is wrong, the tower answers "negative" and gives the correct version, and you read it back again.
A practice game, not training. This is not an approved training device, not an assessment, and not a preparation for any language proficiency check. The aerodrome, its frequencies and the traffic are made up. Phraseology differs from one country to the next: your flight instructor and your national authority's manual are the reference.
What a game looks like
You are F-ABCD, a light aircraft parked at the flying club of Isola…
Python with no web framework, GPL-3.0-or-later. pip install ., python -m phraseology download once, then python -m phraseology serve --open. Any recent laptop works, no GPU. Add ?seed=71 to the page's address to play the game from the video.
How I Built It
The loop is short:
microphone → browser (16 kHz WAV) → faster-whisper, on the CPU
→ normalizer → readback rules ← scenario (state machine)
→ next tower message → Piper → speaker
No language model decides anything. Phraseology is meant to be said word for word, so the moment the tower speaks, the game knows exactly what a correct readback contains: the runway, the QNH, the holding point, the callsign at the end. Checking it is a matter of rules. The open models hear and speak; the rules decide. That also means every message the tower says comes from templates I wrote with made-up values, never from generated text.
The rules are the ones a pilot learns, in my own words. The runway in use, clearances on a runway, altimeter settings, transponder codes and new frequencies are always read back. Other instructions can be acknowledged ("wilco"), though the game asks for taxi instructions in full, as the UK manual does. A readback ends with your callsign, a new call starts with it, and you only shorten F-ABCD to F-CD after the station has.
Most of the work went into the space between what you say and what Whisper writes:
- The prompt must not contain the values. faster-whisper takes an initial prompt, and giving it the station names, the callsign and the radio words helps a lot. Giving it the QNH would be a mistake: the transcript leans towards what's in the prompt, so a wrong readback would come out right. The prompt only holds what never changes during a game.
- Order matters in that prompt. When it ended with "say again", Whisper dropped a "say again" spoken at the start of the clip, as if it had already written it. The callsign now goes last.
- Whisper has opinions about "niner". The base.en model wrote "niner niner eight" as "9 and 9 are 8", or "9R9R8", or turned "zero niner" into "Zero Minor". The bigger small.en heard "nine-oh" and wrote 9098. "Four" came out as "for", "via" as "wire". A normalizer undoes these slips, but only where a slip can't change a value: "for" becomes 4 between digits or after "squawk", and stays "for" in "cleared for take-off"; a zero after a nine is dropped only when the QNH would otherwise read above 1100 hPa, which no altimeter setting does. Then values are compared exactly.
- Mishearings have to fail safe. The tests close the loop with the real models: Piper speaks correct and wrong readbacks (wrong QNH, wrong squawk, wrong frequency, "cleared to land" instead of "cleared for take-off", wrong callsign, "roger" instead of a readback), Whisper transcribes them, and the checker has to get the verdict right. With small.en, the model the demo server ran, and two synthetic voices, 14 of 14 wrong readbacks were rejected and 14 of 16 correct ones were accepted (15 of 16 with the smaller base.en). The misses were correct readbacks misheard ("A.K. Wright" for "vacate right"), so the tower asked again. Across every run, a wrong readback was never accepted.
The demo server started on base.en, the smaller model, because it answered faster. The first real test, with a French accent, broke it: "at the flying club, request taxi" came out as "de France, clamps, stochasticity". On readbacks spoken with a strong French accent, small.en got 5 of 6 right and base.en 3 of 6, at 4.2 seconds per readback on two vCPUs instead of 1.4. I took the slower one: on a practice tool, a readback that is right but rejected is worse than a wait.
I built it over the weekend with an AI coding agent.
Credits, because none of this works without them: Whisper by OpenAI (MIT, open weights), faster-whisper and CTranslate2 (MIT), the Silero VAD model (MIT), Piper with espeak-ng (GPL-3.0-or-later) on ONNX Runtime (MIT), and two voices from rhasspy/piper-voices, en_US-norman for the tower and en_GB-cori for the other aircraft, both trained from scratch on public-domain LibriVox recordings; the video's pilot is Kokoro-82M (Apache 2.0), French voice ff_siwis (SIWIS database, CC BY 4.0).
Why Does Open Innovation Matter?
Practicing a language means sounding unsure, repeating yourself, getting numbers wrong. That recording belongs on my friend's laptop. With open weights, Whisper runs there, on the CPU, and the voice never leaves the machine. After the first download it works with no internet, so it works at the airfield, and it costs nothing to run.
Having the model on my side of the wire also let me do the work above. I could hand Whisper a prompt, measure what that prompt did to a wrong readback, and decide what goes in it. I could pin the version, decode in int8 on a CPU, set the beam size and the voice-activity filter, and run a test suite against the real model as often as I liked.
The voices too: Piper's voices come with model cards, and I could read where each one was trained. Two that I first planned to use turned out to be fine-tuned from a voice whose dataset license doesn't fit this project, so I picked ones trained from scratch on public-domain recordings. And because Piper is GPL, so is this game.
Prize Categories
Best Use of DigitalOcean. The game in the video was served by a DigitalOcean Droplet with 2 vCPUs, 4 GB of memory and no GPU. Whisper small.en ran on the Droplet's CPU and Piper spoke the tower, as a systemd service behind Caddy, which got the HTTPS certificate on its own (the browser only opens the microphone over HTTPS). The Droplet only let in SSH, HTTP and HTTPS, transcribed one recording at a time with a short queue, capped recordings at 10 seconds, rate-limited each address, and never wrote audio to disk. The setup is two scripts in the repository, so anyone can run their own.

Top comments (2)
That cover image alone sold me before I even read a word. Flight sim nerds are going to love this one.
Thanks! Flight sim pilots welcome: this tower never sighs at "say again"!