We've released Ito's inference engine, ESP32-S3 firmware and two English voice models. The latest demo lists 4.05M parameters, about 4M. The chip weights occupy 4.89 MB per voice. The engine produces 24 kHz speech using integer arithmetic.
What runs where
Text becomes phonemes on the host using espeak-ng. Token IDs are sent to the firmware. A streaming acoustic front predicts duration, pitch, energy and a mel spectrogram. A Vocos-style vocoder turns that into audio using ConvNeXt blocks, a harmonic pitch source and an inverse STFT.
The C99 engine builds on a laptop and for ESP32-S3. The target has a 240 MHz dual-core CPU, no neural accelerator, and 8 MB PSRAM; the recommended board is N16R8 with 16 MB flash.
What we've verified
Espressif's QEMU runs the firmware and produces PCM that is bit-identical to the host engine. The public samples are that engine's output. This checks the implementation and arithmetic.
Nothing has run on a physical board yet. Time to first audio and real-time factor are estimates, so real-time playback on silicon is still an open question.
The latest firmware uses a 125 ms first audio chunk. The current firmware documentation estimates first audio at 137-143 ms optimistically, 200-207 ms centrally and 311-318 ms pessimistically. Estimated RTF is 0.53-0.54, 0.77-0.80 and 1.18-1.27 respectively. An RTF above 1 falls behind playback. The pessimistic case does that, so the range matters more than the central number.
These estimates use exact QEMU instruction counts with assumed CPI and PSRAM bandwidth. They are not measured board latency. First audio also differs from the buffering needed for gapless playback: the central estimate for that start delay is 215-240 ms.
Quality checks and limits
The repo contains automatic evaluations and a small listening exercise. The latter used one founder as listener and four sentences per system from the float model. It doesn't establish an independent MOS score, and the exact chip configuration hasn't been blind-rated.
Ito is English-only, with one voice per weights file and a fixed speaking style on the chip. We'd particularly like outside feedback on pronunciation, intonation and the flashing instructions.
Try the chip arithmetic on a computer
After accepting the weight terms at https://huggingface.co/lokutor-ai/ito and following the repo's dependency setup:
cd esp32/host && make && cd ../..
python esp32/tools/chip_wav.py "Good morning! The coffee is ready." hello_chip.wav
Full setup and flashing instructions: https://github.com/lokutor-ai/ito
Samples and comparisons: https://lokutor-ai.github.io/ito/
Licensing
Engine, firmware and Python code are GPLv3. The voice weights and Ito audio are CC BY-NC-SA 4.0 plus additional terms, with gated access on Hugging Face. Commercial use needs a written license from Lokutor. The GPL on the code doesn't make the weights commercially unrestricted. Training code and recipe are not public.
A separate Spanish ASR update
Oído now includes oido_es.tnm and oido_es.tlm, a Spanish Conformer model and language model for the same target. Together they occupy 14.0 + 1.3 MB. Spanish weights are CC BY 4.0; engine code is GPLv3.
Our own host evaluation with the language model reports WER 13.8% on Common Voice, 10.9% on MLS, 15.7% on VoxPopuli and 11.3% on FLEURS across the full test sets. Oído was fine-tuned on training data from these corpora, so this doesn't establish zero-shot generalization or performance on arbitrary regional accents. Spanish noise benchmarking is still missing, and board speed remains unmeasured.
Oído: https://github.com/lokutor-ai/oido
These are separate engines. We haven't validated a complete simultaneous STT/TTS loop on one ESP32-S3.
Top comments (0)