Oído is an open-source speech-to-text engine built by Lokutor for the ESP32-S3. It accepts open-vocabulary English speech rather than a fixed list of commands, with the recognition computation kept on the device. No cloud inference, GPU or NPU is required.
The release is available at github.com/lokutor-ai/oido.
There is an important boundary to this launch: as of September 30, 2026, the engine's arithmetic and firmware transcripts are verified on the host and in Espressif's QEMU emulator. Physical-board speed measurements are still pending. The demo uses chip-exact transcripts, but it is not footage of a physical board running in real time.
That distinction matters when evaluating an embedded speech system. Correct recognition, fitting in memory and keeping up with a microphone are separate claims.
Speech recognition without a command list
A fixed-command recognizer can map a few phrases to device actions. Open-vocabulary recognition has a different job: turn an English sentence into text without requiring the developer to enumerate every possible sentence first.
Oído targets the second case on an ESP32-S3 N16R8: a 240 MHz dual-core Xtensa LX7 microcontroller with 16 MB flash and 8 MB octal PSRAM. Those memory requirements are part of the target, not an optional upgrade. This release does not mean the model fits on every ESP32 board.
The public engine runs NVIDIA's Conformer-CTC Small model, with two supplied weight formats:
- int8: a 14.0 MB model, with the stronger published accuracy.
- int4: an 8.3 MB model, with a quantization-aware fine-tune. The supplied partition layout leaves a 6 MB app partition for your own code.
The choice is a concrete tradeoff between recognition accuracy and flash space for the rest of the device.
What the Whisper comparison actually says
The repository reports lower word error rates than Whisper tiny.en on its LibriSpeech evaluation and on its noise-and-reverberation evaluation. That is an accuracy comparison, not proof that Oído has beaten Whisper in a measured physical-board speed test.
The evaluation scope also matters. Oído's chip-arithmetic rows use the full LibriSpeech test sets; the laptop baselines use 500 evenly spaced utterances per set, with the same text normalization. These are project-reported results, not an independent benchmark on identical hardware or an identical full-set workload.
The useful conclusion is narrower than "microcontrollers are better than laptops": a model and engine designed around this memory and compute budget can provide useful open-vocabulary recognition without a neural accelerator. Check the README's accuracy table and evaluation notes before carrying the comparison into your own product claims.
The engine is the embedded work
The model architecture is only part of fitting recognition onto this chip. The public implementation includes:
- A log-mel front end and convolutional subsampling to 25 Hz.
- A 16-layer Conformer encoder, followed by CTC decoding over 1,024 BPE tokens.
- C kernels that use the ESP32-S3's PIE vector unit for int8 matrix operations, alongside int4 kernels.
- Quantized relative-position attention and a lookup-table softmax.
- Scheduling across both cores and tiled access to weights stored in flash.
- A VAD/AGC segmenter that determines when an utterance is ready to recognize.
The flash-access strategy is worth looking at if you work on embedded inference. The engine tiles computation so each weight is streamed from flash once per 64 frames. On this target, moving weights is part of the workload; counting model operations alone does not describe it.
You can inspect the implementation in esp32/components/tinyasr, the ESP-IDF application in esp32/firmware, and the host build in esp32/host.
Try the same arithmetic on a laptop first
The host build lets you test recognition before wiring a board. From a checkout of the repository, the documented path is:
cd esp32/host
make
./tasr_cli ../../models/nemo8.tnm recording.wav
The recording must be a 16 kHz mono PCM16 WAV. For microphone input, the repository documents python live_demo.py, with Python dependencies including NumPy, SoundFile, SentencePiece and sounddevice.
The point of this host path is to test the firmware engine's arithmetic. It is not a laptop throughput benchmark that can substitute for board timing. The live demo includes ESP32 time estimates; estimates remain estimates.
For the hardware path, the README specifies an ESP32-S3-DevKitC-1 N16R8, an INMP441 I2S microphone and ESP-IDF v5.5. An SSD1306 OLED is optional. The repository includes flashing scripts and an emulator path, so readers can inspect and reproduce the steps rather than rely on a video alone.
What to test before putting it in a device
This release is English-only and works in utterance mode. Text appears after a pause and recognition computation, not word by word as someone speaks. It is therefore not a drop-in promise of streaming captions or instant turn-taking.
The current real-time factor is estimated from emulator instruction counts and assumptions about execution and memory stalls. Physical hardware is the next check.
For a device trial, I would start with:
- Recognition on the actual microphone, enclosure and acoustic environment.
- End-of-speech behavior, including false segmentation and the delay before text appears.
- Sustained operation, power draw and latency alongside the rest of the firmware.
- Crowded speech and reverberant rooms, which the README explicitly lists as hard cases.
Keeping recognition local removes the need to send audio to a cloud recognizer. It does not, by itself, establish the privacy behavior of a whole product; the rest of its firmware still determines what is stored or transmitted.
Code and weights have different licenses
The engine and firmware code are GPLv3. The int8 weights and tokenizer are CC-BY-4.0; the int4 weights are CC-BY-SA-4.0. The supported NVIDIA transducer weights are not bundled and have separate NVIDIA terms.
Read the repository's NOTICE and commercial licensing notes before shipping a device. Lokutor offers commercial licenses for the engine and firmware where the GPL terms do not fit the product. That is separate from the weight licenses.
Try Oído, inspect the engine and share reproducible board results. The next useful evidence is recognition and timing on real hardware, under the conditions your device will actually face.
Try the voices in your browser on Hugging Face. Lokutor 2.0 launches on Product Hunt October 13.
Top comments (0)