DEV Community

Daniel Varela
Daniel Varela

Posted on Originally published at lokutor.com

Oído: Open-Source Speech Recognition on a $5 Chip

Speech recognition on a microcontroller has, in practice, meant one of two things: a wake word, or a short list of commands the device was programmed to expect. Say anything else and it hears nothing.

Today we're releasing Oído: speech recognition for any English sentence, running entirely on an ESP32-S3. That's a $5 microcontroller with a 240 MHz dual-core CPU, 8 MB of PSRAM, 16 MB of flash and no neural accelerator. No cloud, no command list, and it's open source.

Watch the Oído demo on the original post

The transcripts in the video are Oído's chip-exact output, sped up. Footage from a physical board is coming.

¡Oído! is what cooks call out in a Spanish kitchen to confirm an order: heard, got it.

How accurate it is

Word error rate on LibriSpeech, the standard English benchmark (lower is better):

System Runs on test-clean test-other
Oído, int8 ESP32-S3 3.7% 8.2%
Oído, int4 (8.3 MB) ESP32-S3 4.6% 10.0%
Espressif MultiNet7 ESP32-S3 8.5% 21.3%
Moonshine tiny laptop 5.0% 12.1%
Whisper tiny.en laptop 6.3% 15.9%

On the same chip, Oído makes 2.3 to 2.6 times fewer errors than Espressif's own recognizer, which matches speech against a list of up to 200 predefined commands. It also makes fewer errors than Whisper tiny and Moonshine tiny running at full precision on a laptop. That holds in noise too: across 14 conditions built from real car, kitchen and cafeteria recordings, background chatter and room echo, Oído averages 8.4% against 12.1% for Whisper tiny.

Squeezing the model onto the chip costs almost nothing. The int8 version is within about a tenth of a point of the original full-precision model (3.70% vs 3.68% on test-clean).

The model is NVIDIA's. The work is the engine.

Oído runs NVIDIA's openly licensed Conformer-CTC Small, a 13-million-parameter speech model. We didn't retrain or distill it. The hard part was making it run on a chip whose fast internal memory is 512 KB, when the model alone is 14 MB.

The obstacle isn't arithmetic, it's memory bandwidth. The weights live in flash, which the chip reads through a small cache at a few tens of megabytes per second. So we wrote a new inference engine in C for the ESP32-S3's vector instructions, designed around that limit:

  • Integer matrix kernels that do 16 multiply-adds per instruction.
  • Integer attention, softmax included, so the whole model runs in int8 arithmetic.
  • A bandwidth-aware schedule that processes 64 frames of audio at a time, so each weight is read from flash once per block instead of once per frame. This cut weight traffic from 18 to 7.5 MB per second of audio.
  • Both CPU cores working in parallel.

There's also an int4 version that takes 8.3 MB of flash instead of 14 MB, leaving 6 MB for your own application, and an optional on-chip language model that brings the error rate down to 3.3% and 7.2%.

Where it stands

We want to be precise about this. Every transcript and accuracy number above comes from the exact arithmetic of the on-chip engine, and the real firmware produces the same transcripts, word for word, in Espressif's QEMU emulator. Speed is estimated from exact instruction counts at 0.7 to 0.95 times real time. Boards arrive this week, and we'll publish measured numbers here.

Two other limits: it's English only for now, and it transcribes after each utterance rather than word by word, so text appears about 3 seconds after you stop talking on a short command.

Why we built it

Voice is moving off the cloud. A device that streams audio to a GPU pays for inference for as long as it exists, stops understanding people when the connection drops, and sends their voice somewhere else. On-device recognition removes all three problems, but until now it also meant giving up open-ended speech. Oído shows that a commodity $5 chip is enough for the real thing.

It's the listening half of the on-device voice stack we're building at Lokutor for this class of hardware.

Try it

Everything is on GitHub at github.com/lokutor-ai/oido, and the models are on Hugging Face: int8 and int4.

You don't need a board to try it. The host build runs the chip's exact arithmetic on your laptop, including a live microphone demo (Python with numpy, soundfile, sentencepiece and sounddevice):

git clone https://github.com/lokutor-ai/oido
cd oido/esp32/host && make
python live_demo.py
Enter fullscreen mode Exit fullscreen mode

To run it on hardware you need an ESP32-S3-DevKitC-1 N16R8 and an INMP441 microphone. The README has the wiring and a one-line flash script.

Open source, with a commercial option

The code is licensed under GPLv3. The int8 model is CC-BY-4.0 and the int4 model is CC-BY-SA-4.0. For products that can't meet GPLv3 terms, we offer commercial licenses, models for other languages and integration support. Write to us at contact@lokutor.com.

We're also writing up the full technical details as a paper. If you build something with Oído, we'd love to hear about it.

Try the voices in your browser on Hugging Face. Lokutor 2.0 launches on Product Hunt October 13.

Top comments (0)