<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Daniel Varela</title>
    <description>The latest articles on DEV Community by Daniel Varela (@danivs10).</description>
    <link>https://dev.to/danivs10</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4131813%2F88dd238f-872e-40bf-85cf-54a85a1b3a49.jpg</url>
      <title>DEV Community: Daniel Varela</title>
      <link>https://dev.to/danivs10</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/danivs10"/>
    <language>en</language>
    <item>
      <title>Ito: streaming speech synthesis in 4.89 MB, verified in ESP32-S3 emulation</title>
      <dc:creator>Daniel Varela</dc:creator>
      <pubDate>Mon, 05 Oct 2026 09:07:59 +0000</pubDate>
      <link>https://dev.to/danivs10/ito-streaming-speech-synthesis-in-489-mb-verified-in-esp32-s3-emulation-3jp1</link>
      <guid>https://dev.to/danivs10/ito-streaming-speech-synthesis-in-489-mb-verified-in-esp32-s3-emulation-3jp1</guid>
      <description>&lt;p&gt;We've released Ito's inference engine, ESP32-S3 firmware and two English voice models. The latest demo lists 4.05M parameters, about 4M. The chip weights occupy 4.89 MB per voice. The engine produces 24 kHz speech using integer arithmetic.&lt;/p&gt;

&lt;h2&gt;
  
  
  What runs where
&lt;/h2&gt;

&lt;p&gt;Text becomes phonemes on the host using espeak-ng. Token IDs are sent to the firmware. A streaming acoustic front predicts duration, pitch, energy and a mel spectrogram. A Vocos-style vocoder turns that into audio using ConvNeXt blocks, a harmonic pitch source and an inverse STFT.&lt;/p&gt;

&lt;p&gt;The C99 engine builds on a laptop and for ESP32-S3. The target has a 240 MHz dual-core CPU, no neural accelerator, and 8 MB PSRAM; the recommended board is N16R8 with 16 MB flash.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we've verified
&lt;/h2&gt;

&lt;p&gt;Espressif's QEMU runs the firmware and produces PCM that is bit-identical to the host engine. The public samples are that engine's output. This checks the implementation and arithmetic.&lt;/p&gt;

&lt;p&gt;Nothing has run on a physical board yet. Time to first audio and real-time factor are estimates, so real-time playback on silicon is still an open question.&lt;/p&gt;

&lt;p&gt;The latest firmware uses a 125 ms first audio chunk. The current firmware documentation estimates first audio at 137-143 ms optimistically, 200-207 ms centrally and 311-318 ms pessimistically. Estimated RTF is 0.53-0.54, 0.77-0.80 and 1.18-1.27 respectively. An RTF above 1 falls behind playback. The pessimistic case does that, so the range matters more than the central number.&lt;/p&gt;

&lt;p&gt;These estimates use exact QEMU instruction counts with assumed CPI and PSRAM bandwidth. They are not measured board latency. First audio also differs from the buffering needed for gapless playback: the central estimate for that start delay is 215-240 ms.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quality checks and limits
&lt;/h2&gt;

&lt;p&gt;The repo contains automatic evaluations and a small listening exercise. The latter used one founder as listener and four sentences per system from the float model. It doesn't establish an independent MOS score, and the exact chip configuration hasn't been blind-rated.&lt;/p&gt;

&lt;p&gt;Ito is English-only, with one voice per weights file and a fixed speaking style on the chip. We'd particularly like outside feedback on pronunciation, intonation and the flashing instructions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try the chip arithmetic on a computer
&lt;/h2&gt;

&lt;p&gt;After accepting the weight terms at &lt;a href="https://huggingface.co/lokutor-ai/ito" rel="noopener noreferrer"&gt;https://huggingface.co/lokutor-ai/ito&lt;/a&gt; and following the repo's dependency setup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;esp32/host &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; ../..
python esp32/tools/chip_wav.py &lt;span class="s2"&gt;"Good morning! The coffee is ready."&lt;/span&gt; hello_chip.wav
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Full setup and flashing instructions: &lt;a href="https://github.com/lokutor-ai/ito" rel="noopener noreferrer"&gt;https://github.com/lokutor-ai/ito&lt;/a&gt;&lt;br&gt;
Samples and comparisons: &lt;a href="https://lokutor-ai.github.io/ito/" rel="noopener noreferrer"&gt;https://lokutor-ai.github.io/ito/&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Licensing
&lt;/h2&gt;

&lt;p&gt;Engine, firmware and Python code are GPLv3. The voice weights and Ito audio are CC BY-NC-SA 4.0 plus additional terms, with gated access on Hugging Face. Commercial use needs a written license from Lokutor. The GPL on the code doesn't make the weights commercially unrestricted. Training code and recipe are not public.&lt;/p&gt;

&lt;h2&gt;
  
  
  A separate Spanish ASR update
&lt;/h2&gt;

&lt;p&gt;Oído now includes &lt;code&gt;oido_es.tnm&lt;/code&gt; and &lt;code&gt;oido_es.tlm&lt;/code&gt;, a Spanish Conformer model and language model for the same target. Together they occupy 14.0 + 1.3 MB. Spanish weights are CC BY 4.0; engine code is GPLv3.&lt;/p&gt;

&lt;p&gt;Our own host evaluation with the language model reports WER 13.8% on Common Voice, 10.9% on MLS, 15.7% on VoxPopuli and 11.3% on FLEURS across the full test sets. Oído was fine-tuned on training data from these corpora, so this doesn't establish zero-shot generalization or performance on arbitrary regional accents. Spanish noise benchmarking is still missing, and board speed remains unmeasured.&lt;/p&gt;

&lt;p&gt;Oído: &lt;a href="https://github.com/lokutor-ai/oido" rel="noopener noreferrer"&gt;https://github.com/lokutor-ai/oido&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;These are separate engines. We haven't validated a complete simultaneous STT/TTS loop on one ESP32-S3.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>hardware</category>
      <category>iot</category>
      <category>software</category>
    </item>
    <item>
      <title>Oído: Open-Source Speech Recognition on a $5 Chip</title>
      <dc:creator>Daniel Varela</dc:creator>
      <pubDate>Thu, 01 Oct 2026 10:24:51 +0000</pubDate>
      <link>https://dev.to/danivs10/oido-open-source-speech-recognition-on-a-5-chip-3n9d</link>
      <guid>https://dev.to/danivs10/oido-open-source-speech-recognition-on-a-5-chip-3n9d</guid>
      <description>&lt;p&gt;Speech recognition on a microcontroller has, in practice, meant one of two things: a wake word, or a short list of commands the device was programmed to expect. Say anything else and it hears nothing.&lt;/p&gt;

&lt;p&gt;Today we're releasing &lt;strong&gt;Oído&lt;/strong&gt;: speech recognition for any English sentence, running entirely on an ESP32-S3. That's a $5 microcontroller with a 240 MHz dual-core CPU, 8 MB of PSRAM, 16 MB of flash and no neural accelerator. No cloud, no command list, and it's open source.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://lokutor.com/blog/introducing-oido-speech-recognition-on-a-5-dollar-chip/" rel="noopener noreferrer"&gt;Watch the Oído demo on the original post&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The transcripts in the video are Oído's chip-exact output, sped up. Footage from a physical board is coming.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;¡Oído!&lt;/em&gt; is what cooks call out in a Spanish kitchen to confirm an order: &lt;em&gt;heard, got it&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How accurate it is
&lt;/h2&gt;

&lt;p&gt;Word error rate on LibriSpeech, the standard English benchmark (lower is better):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;Runs on&lt;/th&gt;
&lt;th&gt;test-clean&lt;/th&gt;
&lt;th&gt;test-other&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Oído, int8&lt;/td&gt;
&lt;td&gt;ESP32-S3&lt;/td&gt;
&lt;td&gt;3.7%&lt;/td&gt;
&lt;td&gt;8.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Oído, int4 (8.3 MB)&lt;/td&gt;
&lt;td&gt;ESP32-S3&lt;/td&gt;
&lt;td&gt;4.6%&lt;/td&gt;
&lt;td&gt;10.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Espressif MultiNet7&lt;/td&gt;
&lt;td&gt;ESP32-S3&lt;/td&gt;
&lt;td&gt;8.5%&lt;/td&gt;
&lt;td&gt;21.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Moonshine tiny&lt;/td&gt;
&lt;td&gt;laptop&lt;/td&gt;
&lt;td&gt;5.0%&lt;/td&gt;
&lt;td&gt;12.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whisper tiny.en&lt;/td&gt;
&lt;td&gt;laptop&lt;/td&gt;
&lt;td&gt;6.3%&lt;/td&gt;
&lt;td&gt;15.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On the same chip, Oído makes 2.3 to 2.6 times fewer errors than Espressif's own recognizer, which matches speech against a list of up to 200 predefined commands. It also makes fewer errors than Whisper tiny and Moonshine tiny running at full precision on a laptop. That holds in noise too: across 14 conditions built from real car, kitchen and cafeteria recordings, background chatter and room echo, Oído averages 8.4% against 12.1% for Whisper tiny.&lt;/p&gt;

&lt;p&gt;Squeezing the model onto the chip costs almost nothing. The int8 version is within about a tenth of a point of the original full-precision model (3.70% vs 3.68% on test-clean).&lt;/p&gt;

&lt;h2&gt;
  
  
  The model is NVIDIA's. The work is the engine.
&lt;/h2&gt;

&lt;p&gt;Oído runs NVIDIA's openly licensed Conformer-CTC Small, a 13-million-parameter speech model. We didn't retrain or distill it. The hard part was making it run on a chip whose fast internal memory is 512 KB, when the model alone is 14 MB.&lt;/p&gt;

&lt;p&gt;The obstacle isn't arithmetic, it's memory bandwidth. The weights live in flash, which the chip reads through a small cache at a few tens of megabytes per second. So we wrote a new inference engine in C for the ESP32-S3's vector instructions, designed around that limit:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Integer matrix kernels&lt;/strong&gt; that do 16 multiply-adds per instruction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integer attention&lt;/strong&gt;, softmax included, so the whole model runs in int8 arithmetic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A bandwidth-aware schedule&lt;/strong&gt; that processes 64 frames of audio at a time, so each weight is read from flash once per block instead of once per frame. This cut weight traffic from 18 to 7.5 MB per second of audio.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Both CPU cores&lt;/strong&gt; working in parallel.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There's also an int4 version that takes 8.3 MB of flash instead of 14 MB, leaving 6 MB for your own application, and an optional on-chip language model that brings the error rate down to 3.3% and 7.2%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;We want to be precise about this. Every transcript and accuracy number above comes from the exact arithmetic of the on-chip engine, and the real firmware produces the same transcripts, word for word, in Espressif's QEMU emulator. Speed is estimated from exact instruction counts at 0.7 to 0.95 times real time. Boards arrive this week, and we'll publish measured numbers here.&lt;/p&gt;

&lt;p&gt;Two other limits: it's English only for now, and it transcribes after each utterance rather than word by word, so text appears about 3 seconds after you stop talking on a short command.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why we built it
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://lokutor.com/blog/edge-voice-ai-off-the-cloud/" rel="noopener noreferrer"&gt;Voice is moving off the cloud&lt;/a&gt;. A device that streams audio to a GPU pays for inference for as long as it exists, stops understanding people when the connection drops, and sends their voice somewhere else. On-device recognition removes all three problems, but until now it also meant giving up open-ended speech. Oído shows that a commodity $5 chip is enough for the real thing.&lt;/p&gt;

&lt;p&gt;It's the listening half of the on-device voice stack we're building at Lokutor for this class of hardware.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Everything is on GitHub at &lt;a href="https://github.com/lokutor-ai/oido" rel="noopener noreferrer"&gt;github.com/lokutor-ai/oido&lt;/a&gt;, and the models are on Hugging Face: &lt;a href="https://huggingface.co/lokutor-ai/oido-ctc-small-int8" rel="noopener noreferrer"&gt;int8&lt;/a&gt; and &lt;a href="https://huggingface.co/lokutor-ai/oido-ctc-small-int4" rel="noopener noreferrer"&gt;int4&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;You don't need a board to try it. The host build runs the chip's exact arithmetic on your laptop, including a live microphone demo (Python with numpy, soundfile, sentencepiece and sounddevice):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/lokutor-ai/oido
&lt;span class="nb"&gt;cd &lt;/span&gt;oido/esp32/host &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make
python live_demo.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To run it on hardware you need an ESP32-S3-DevKitC-1 N16R8 and an INMP441 microphone. The README has the wiring and a one-line flash script.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open source, with a commercial option
&lt;/h2&gt;

&lt;p&gt;The code is licensed under GPLv3. The int8 model is CC-BY-4.0 and the int4 model is CC-BY-SA-4.0. For products that can't meet GPLv3 terms, we offer commercial licenses, models for other languages and integration support. Write to us at &lt;a href="mailto:contact@lokutor.com"&gt;contact@lokutor.com&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;We're also writing up the full technical details as a paper. If you build something with Oído, we'd love to hear about it.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Try the voices in your browser on &lt;a href="https://huggingface.co/spaces/lokutor-ai/tts-on-cpu" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt;. Lokutor 2.0 launches on &lt;a href="https://www.producthunt.com/products/lokutor-2-0?launch=lokutor-2-0" rel="noopener noreferrer"&gt;Product Hunt&lt;/a&gt; October 13.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Oído: open-source speech recognition for the ESP32-S3, without a command list</title>
      <dc:creator>Daniel Varela</dc:creator>
      <pubDate>Wed, 30 Sep 2026 19:40:26 +0000</pubDate>
      <link>https://dev.to/danivs10/oido-open-source-speech-recognition-for-the-esp32-s3-without-a-command-list-2n48</link>
      <guid>https://dev.to/danivs10/oido-open-source-speech-recognition-for-the-esp32-s3-without-a-command-list-2n48</guid>
      <description>&lt;p&gt;Oído is an open-source speech-to-text engine built by Lokutor for the ESP32-S3. It accepts open-vocabulary English speech rather than a fixed list of commands, with the recognition computation kept on the device. No cloud inference, GPU or NPU is required.&lt;/p&gt;

&lt;p&gt;The release is available at &lt;a href="https://github.com/lokutor-ai/oido" rel="noopener noreferrer"&gt;github.com/lokutor-ai/oido&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;There is an important boundary to this launch: &lt;strong&gt;as of September 30, 2026, the engine's arithmetic and firmware transcripts are verified on the host and in Espressif's QEMU emulator. Physical-board speed measurements are still pending.&lt;/strong&gt; The demo uses chip-exact transcripts, but it is not footage of a physical board running in real time.&lt;/p&gt;

&lt;p&gt;That distinction matters when evaluating an embedded speech system. Correct recognition, fitting in memory and keeping up with a microphone are separate claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  Speech recognition without a command list
&lt;/h2&gt;

&lt;p&gt;A fixed-command recognizer can map a few phrases to device actions. Open-vocabulary recognition has a different job: turn an English sentence into text without requiring the developer to enumerate every possible sentence first.&lt;/p&gt;

&lt;p&gt;Oído targets the second case on an ESP32-S3 N16R8: a 240 MHz dual-core Xtensa LX7 microcontroller with 16 MB flash and 8 MB octal PSRAM. Those memory requirements are part of the target, not an optional upgrade. This release does not mean the model fits on every ESP32 board.&lt;/p&gt;

&lt;p&gt;The public engine runs NVIDIA's Conformer-CTC Small model, with two supplied weight formats:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;int8:&lt;/strong&gt; a 14.0 MB model, with the stronger published accuracy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;int4:&lt;/strong&gt; an 8.3 MB model, with a quantization-aware fine-tune. The supplied partition layout leaves a 6 MB app partition for your own code.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The choice is a concrete tradeoff between recognition accuracy and flash space for the rest of the device.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Whisper comparison actually says
&lt;/h2&gt;

&lt;p&gt;The repository reports lower word error rates than Whisper tiny.en on its LibriSpeech evaluation and on its noise-and-reverberation evaluation. That is an accuracy comparison, not proof that Oído has beaten Whisper in a measured physical-board speed test.&lt;/p&gt;

&lt;p&gt;The evaluation scope also matters. Oído's chip-arithmetic rows use the full LibriSpeech test sets; the laptop baselines use 500 evenly spaced utterances per set, with the same text normalization. These are project-reported results, not an independent benchmark on identical hardware or an identical full-set workload.&lt;/p&gt;

&lt;p&gt;The useful conclusion is narrower than "microcontrollers are better than laptops": a model and engine designed around this memory and compute budget can provide useful open-vocabulary recognition without a neural accelerator. Check the &lt;a href="https://github.com/lokutor-ai/oido" rel="noopener noreferrer"&gt;README's accuracy table and evaluation notes&lt;/a&gt; before carrying the comparison into your own product claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  The engine is the embedded work
&lt;/h2&gt;

&lt;p&gt;The model architecture is only part of fitting recognition onto this chip. The public implementation includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A log-mel front end and convolutional subsampling to 25 Hz.&lt;/li&gt;
&lt;li&gt;A 16-layer Conformer encoder, followed by CTC decoding over 1,024 BPE tokens.&lt;/li&gt;
&lt;li&gt;C kernels that use the ESP32-S3's PIE vector unit for int8 matrix operations, alongside int4 kernels.&lt;/li&gt;
&lt;li&gt;Quantized relative-position attention and a lookup-table softmax.&lt;/li&gt;
&lt;li&gt;Scheduling across both cores and tiled access to weights stored in flash.&lt;/li&gt;
&lt;li&gt;A VAD/AGC segmenter that determines when an utterance is ready to recognize.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The flash-access strategy is worth looking at if you work on embedded inference. The engine tiles computation so each weight is streamed from flash once per 64 frames. On this target, moving weights is part of the workload; counting model operations alone does not describe it.&lt;/p&gt;

&lt;p&gt;You can inspect the implementation in &lt;code&gt;esp32/components/tinyasr&lt;/code&gt;, the ESP-IDF application in &lt;code&gt;esp32/firmware&lt;/code&gt;, and the host build in &lt;code&gt;esp32/host&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try the same arithmetic on a laptop first
&lt;/h2&gt;

&lt;p&gt;The host build lets you test recognition before wiring a board. From a checkout of the repository, the documented path is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;esp32/host
make
./tasr_cli ../../models/nemo8.tnm recording.wav
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The recording must be a 16 kHz mono PCM16 WAV. For microphone input, the repository documents &lt;code&gt;python live_demo.py&lt;/code&gt;, with Python dependencies including NumPy, SoundFile, SentencePiece and sounddevice.&lt;/p&gt;

&lt;p&gt;The point of this host path is to test the firmware engine's arithmetic. It is not a laptop throughput benchmark that can substitute for board timing. The live demo includes ESP32 time estimates; estimates remain estimates.&lt;/p&gt;

&lt;p&gt;For the hardware path, the README specifies an ESP32-S3-DevKitC-1 N16R8, an INMP441 I2S microphone and ESP-IDF v5.5. An SSD1306 OLED is optional. The repository includes flashing scripts and an emulator path, so readers can inspect and reproduce the steps rather than rely on a video alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to test before putting it in a device
&lt;/h2&gt;

&lt;p&gt;This release is English-only and works in utterance mode. Text appears after a pause and recognition computation, not word by word as someone speaks. It is therefore not a drop-in promise of streaming captions or instant turn-taking.&lt;/p&gt;

&lt;p&gt;The current real-time factor is estimated from emulator instruction counts and assumptions about execution and memory stalls. Physical hardware is the next check.&lt;/p&gt;

&lt;p&gt;For a device trial, I would start with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Recognition on the actual microphone, enclosure and acoustic environment.&lt;/li&gt;
&lt;li&gt;End-of-speech behavior, including false segmentation and the delay before text appears.&lt;/li&gt;
&lt;li&gt;Sustained operation, power draw and latency alongside the rest of the firmware.&lt;/li&gt;
&lt;li&gt;Crowded speech and reverberant rooms, which the README explicitly lists as hard cases.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keeping recognition local removes the need to send audio to a cloud recognizer. It does not, by itself, establish the privacy behavior of a whole product; the rest of its firmware still determines what is stored or transmitted.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code and weights have different licenses
&lt;/h2&gt;

&lt;p&gt;The engine and firmware code are &lt;strong&gt;GPLv3&lt;/strong&gt;. The int8 weights and tokenizer are &lt;strong&gt;CC-BY-4.0&lt;/strong&gt;; the int4 weights are &lt;strong&gt;CC-BY-SA-4.0&lt;/strong&gt;. The supported NVIDIA transducer weights are not bundled and have separate NVIDIA terms.&lt;/p&gt;

&lt;p&gt;Read the repository's &lt;a href="https://github.com/lokutor-ai/oido/blob/main/NOTICE" rel="noopener noreferrer"&gt;NOTICE&lt;/a&gt; and &lt;a href="https://github.com/lokutor-ai/oido/blob/main/COMMERCIAL.md" rel="noopener noreferrer"&gt;commercial licensing notes&lt;/a&gt; before shipping a device. Lokutor offers commercial licenses for the engine and firmware where the GPL terms do not fit the product. That is separate from the weight licenses.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/lokutor-ai/oido" rel="noopener noreferrer"&gt;Try Oído, inspect the engine and share reproducible board results&lt;/a&gt;. The next useful evidence is recognition and timing on real hardware, under the conditions your device will actually face.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Try the voices in your browser on &lt;a href="https://huggingface.co/spaces/lokutor-ai/tts-on-cpu" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt;. Lokutor 2.0 launches on &lt;a href="https://www.producthunt.com/products/lokutor-2-0?launch=lokutor-2-0" rel="noopener noreferrer"&gt;Product Hunt&lt;/a&gt; October 13.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Voice agent pricing at real volumes: the monthly bill from 500 to 100,000 minutes</title>
      <dc:creator>Daniel Varela</dc:creator>
      <pubDate>Sun, 27 Sep 2026 09:45:38 +0000</pubDate>
      <link>https://dev.to/danivs10/voice-agent-pricing-at-real-volumes-the-monthly-bill-from-500-to-100000-minutes-460d</link>
      <guid>https://dev.to/danivs10/voice-agent-pricing-at-real-volumes-the-monthly-bill-from-500-to-100000-minutes-460d</guid>
      <description>&lt;p&gt;Per-minute list prices are the easiest numbers in voice AI to compare and the least useful. Nobody pays list price: you pay a monthly plan plus whatever your minutes spill over it, and the plan decides the real bill. So here is the real bill, at five monthly volumes, for two platforms whose pricing pages publish enough to do the math: ElevenLabs Agents and Lokutor.&lt;/p&gt;

&lt;p&gt;This comparison is deliberate. Lokutor is our product, and the math favors it at every volume below. The numbers are not ours: every ElevenLabs figure below is from elevenlabs.io/pricing, read 27 September 2026. Do your own math before you believe anyone's, including this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two plan tables
&lt;/h2&gt;

&lt;p&gt;ElevenLabs Agents, monthly billing, from their pricing page on 27 September 2026:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;$/month&lt;/th&gt;
&lt;th&gt;Agent minutes included&lt;/th&gt;
&lt;th&gt;Effective per minute&lt;/th&gt;
&lt;th&gt;Additional minute&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;td&gt;$0.080&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Starter&lt;/td&gt;
&lt;td&gt;$6&lt;/td&gt;
&lt;td&gt;75&lt;/td&gt;
&lt;td&gt;8.0c&lt;/td&gt;
&lt;td&gt;$0.080&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creator&lt;/td&gt;
&lt;td&gt;$22&lt;/td&gt;
&lt;td&gt;275&lt;/td&gt;
&lt;td&gt;8.0c&lt;/td&gt;
&lt;td&gt;$0.080&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro&lt;/td&gt;
&lt;td&gt;$99&lt;/td&gt;
&lt;td&gt;1,238&lt;/td&gt;
&lt;td&gt;8.0c&lt;/td&gt;
&lt;td&gt;$0.080&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scale&lt;/td&gt;
&lt;td&gt;$299&lt;/td&gt;
&lt;td&gt;3,738&lt;/td&gt;
&lt;td&gt;8.0c&lt;/td&gt;
&lt;td&gt;$0.080&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Business&lt;/td&gt;
&lt;td&gt;$990&lt;/td&gt;
&lt;td&gt;12,375&lt;/td&gt;
&lt;td&gt;8.0c&lt;/td&gt;
&lt;td&gt;$0.080&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One thing worth crediting: 8 cents a minute, flat, on every tier. No volume discount, but no games either. Burst pricing when you exceed your concurrent call limit is $0.160 a minute, and the LLM (Gemini 2.5 Flash in their estimator) bills separately from credits. Telephony is excluded.&lt;/p&gt;

&lt;p&gt;Lokutor, from &lt;a href="https://lokutor.com/blog/two-cents-a-minute/" rel="noopener noreferrer"&gt;the pricing post&lt;/a&gt; published 23 September 2026:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;$/month&lt;/th&gt;
&lt;th&gt;Minutes included&lt;/th&gt;
&lt;th&gt;Effective per minute&lt;/th&gt;
&lt;th&gt;Overage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;td&gt;-&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Starter&lt;/td&gt;
&lt;td&gt;$19&lt;/td&gt;
&lt;td&gt;1,000&lt;/td&gt;
&lt;td&gt;1.9c&lt;/td&gt;
&lt;td&gt;2c&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Growth&lt;/td&gt;
&lt;td&gt;$99&lt;/td&gt;
&lt;td&gt;6,000&lt;/td&gt;
&lt;td&gt;1.65c&lt;/td&gt;
&lt;td&gt;1.8c&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Business&lt;/td&gt;
&lt;td&gt;$349&lt;/td&gt;
&lt;td&gt;25,000&lt;/td&gt;
&lt;td&gt;1.4c&lt;/td&gt;
&lt;td&gt;1.5c&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The Lokutor minute includes speech recognition, the language model, and the voice. Phone minutes count double inbound and quadruple outbound, because carriers charge more to call a mobile.&lt;/p&gt;

&lt;h2&gt;
  
  
  The monthly bill at real volumes
&lt;/h2&gt;

&lt;p&gt;Cheapest legal combination of plan plus overage on each side, voice agent minutes only:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Minutes per month&lt;/th&gt;
&lt;th&gt;ElevenLabs Agents&lt;/th&gt;
&lt;th&gt;Lokutor&lt;/th&gt;
&lt;th&gt;Ratio&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;$40 (Creator + 225 over)&lt;/td&gt;
&lt;td&gt;$19 (Starter)&lt;/td&gt;
&lt;td&gt;2.1x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2,000&lt;/td&gt;
&lt;td&gt;$160 (Pro + 762 over)&lt;/td&gt;
&lt;td&gt;$99 (Growth)&lt;/td&gt;
&lt;td&gt;1.6x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;$800 (Scale + 6,262 over)&lt;/td&gt;
&lt;td&gt;$349 (Business)&lt;/td&gt;
&lt;td&gt;2.3x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50,000&lt;/td&gt;
&lt;td&gt;$4,000 (Business + 37,625 over)&lt;/td&gt;
&lt;td&gt;$724 (Business + 25,000 over)&lt;/td&gt;
&lt;td&gt;5.5x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At 100,000 minutes the ElevenLabs bill is $8,000 (Business plus 87,625 overage minutes). Lokutor's published overage covers up to one extra month of plan minutes, so 100,000 is volume-contract territory; the first 50,000 of it costs $724.&lt;/p&gt;

&lt;p&gt;The gap widens with volume because the overage rates do all the work: 8 cents against 1.5.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this math leaves out, honestly
&lt;/h2&gt;

&lt;p&gt;Concurrency. ElevenLabs Business allows 40 concurrent calls, Lokutor Business 20. If your peak is 30 simultaneous calls, the cheaper minute does not help you.&lt;/p&gt;

&lt;p&gt;Languages. ElevenLabs lists 32. Lokutor lists 9: Spanish, Catalan, Galician, Basque, Portuguese, English, French, Italian, German. The Iberian set is first-class, which is the point of the product, but 9 is fewer than 32.&lt;/p&gt;

&lt;p&gt;Catalog and features. ElevenLabs offers thousands of library voices, dubbing, sound effects, and a music product. Lokutor offers 10 voices and voice cloning on upper plans. If your product depends on a specific ElevenLabs voice, a per-minute table is not the decision.&lt;/p&gt;

&lt;p&gt;The LLM line. On ElevenLabs the language model bills from credits on top of the minute. Their estimator puts Gemini 2.5 Flash at $0.34 for 270 minutes, so it is small, but it is not zero. On Lokutor it is inside the minute.&lt;/p&gt;

&lt;p&gt;Free tiers. ElevenLabs gives 15 agent minutes a month. Lokutor gives 60, no card. For evaluation, both are enough to hear your own use case, which is the only benchmark that matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the difference comes from
&lt;/h2&gt;

&lt;p&gt;Lokutor's voice model is a few tens of millions of parameters and runs on CPUs; the category norm is models tens of times larger on GPUs, priced to match. The 1.4 to 2 cent minute is what the architecture costs to run, plus a business. Whether that trade is worth it depends on the section above, not on the headline number.&lt;/p&gt;

&lt;p&gt;The latency side of the same architecture is in &lt;a href="https://dev.to/danivs10/voice-agents-on-cpu-vs-the-gpu-incumbents-latency-cost-and-a-deliberate-comparison-hge"&gt;Voice agents on CPU vs the GPU incumbents&lt;/a&gt;, and the migration mechanics are in &lt;a href="https://dev.to/danivs10/migrate-from-elevenlabs-to-a-cpu-voice-api-in-about-20-lines-4ibe"&gt;Migrate from ElevenLabs to a CPU voice API in about 20 lines&lt;/a&gt;. Pricing sources: &lt;a href="https://elevenlabs.io/pricing" rel="noopener noreferrer"&gt;ElevenLabs pricing&lt;/a&gt;, &lt;a href="https://lokutor.com/pricing" rel="noopener noreferrer"&gt;Lokutor pricing&lt;/a&gt;, both as of late September 2026. Try the free tier at &lt;a href="https://app.lokutor.com" rel="noopener noreferrer"&gt;app.lokutor.com&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Try the voices in your browser on &lt;a href="https://huggingface.co/spaces/lokutor-ai/tts-on-cpu" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt;. Lokutor 2.0 launches on &lt;a href="https://www.producthunt.com/products/lokutor-2-0?launch=lokutor-2-0" rel="noopener noreferrer"&gt;Product Hunt&lt;/a&gt; October 13.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>pricing</category>
      <category>voiceai</category>
      <category>voiceagents</category>
      <category>tts</category>
    </item>
    <item>
      <title>Migrate from ElevenLabs to a CPU voice API in about 20 lines</title>
      <dc:creator>Daniel Varela</dc:creator>
      <pubDate>Sat, 26 Sep 2026 16:34:31 +0000</pubDate>
      <link>https://dev.to/danivs10/migrate-from-elevenlabs-to-a-cpu-voice-api-in-about-20-lines-4ibe</link>
      <guid>https://dev.to/danivs10/migrate-from-elevenlabs-to-a-cpu-voice-api-in-about-20-lines-4ibe</guid>
      <description>&lt;p&gt;Moving text-to-speech from ElevenLabs to Lokutor is a small diff: same install, same one call, same play-through-speakers result. What changes is what sits behind the call: a few-tens-of-millions-parameter voice model running on commodity CPUs, and a per-minute price that includes more of the stack.&lt;/p&gt;

&lt;p&gt;Scope note up front: this post covers the real-time path (TTS for apps and voice agents). If you use ElevenLabs for dubbing, sound effects, or its voice library breadth, this migration is not aimed at you. The broader comparison lives in &lt;a href="https://dev.to/danivs10/elevenlabs-alternative-for-real-time-voice-agents-cpu-native-2-cents-a-minute-28em"&gt;ElevenLabs alternative for real-time voice agents&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The swap, in Python
&lt;/h2&gt;

&lt;p&gt;Before, from the ElevenLabs quickstart (checked 26 September 2026):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;elevenlabs.client&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ElevenLabs&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;elevenlabs.play&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;play&lt;/span&gt;

&lt;span class="n"&gt;elevenlabs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ElevenLabs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ELEVENLABS_API_KEY&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;audio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;elevenlabs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text_to_speech&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;convert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The first move is what sets everything in motion.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;voice_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;JBFqnCBsd6RMkjVDRZzb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# George
&lt;/span&gt;    &lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eleven_v3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;output_format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mp3_44100_128&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;play&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After, with &lt;code&gt;pip install lokutor&lt;/code&gt; and a key from &lt;a href="https://app.lokutor.com" rel="noopener noreferrer"&gt;app.lokutor.com&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;lokutor&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;TTSClient&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;VoiceStyle&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Language&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;TTSClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;LOKUTOR_API_KEY&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;synthesize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The first move is what sets everything in motion.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;voice&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;VoiceStyle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;F1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;# Claire
&lt;/span&gt;    &lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;Language&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ENGLISH&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# plays through the speakers by default
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole migration for the common case. There is also a JavaScript SDK (&lt;code&gt;@lokutor/sdk&lt;/code&gt;) with the same shape.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes under you
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;ElevenLabs&lt;/th&gt;
&lt;th&gt;Lokutor&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Package&lt;/td&gt;
&lt;td&gt;&lt;code&gt;pip install elevenlabs&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;pip install lokutor&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Voice selection&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;voice_id&lt;/code&gt; string, 3,000+ library voices and clones&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;VoiceStyle&lt;/code&gt; enum, 10 voices (F1 to F5, M1 to M5)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model choice&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;model_id&lt;/code&gt;: eleven_v3, multilingual_v2, flash_v2.5&lt;/td&gt;
&lt;td&gt;One production model; &lt;code&gt;steps&lt;/code&gt; 1 to 32 as a quality dial&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;MP3 or your chosen format&lt;/td&gt;
&lt;td&gt;PCM16 stream, 44.1 kHz mono&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Languages&lt;/td&gt;
&lt;td&gt;32 on multilingual_v2&lt;/td&gt;
&lt;td&gt;9: en, es, ca, gl, eu, pt, fr, it, de&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Free tier&lt;/td&gt;
&lt;td&gt;10k characters per month&lt;/td&gt;
&lt;td&gt;60 minutes per month&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read both sides honestly: ElevenLabs gives you a much larger voice library and more than three times the languages. Lokutor gives you one streaming PCM path, one model to reason about, and latency measured in production: 139 ms median time to first audio, 211 ms p90, on a single CPU thread (&lt;a href="https://dev.to/danivs10/near-gpu-tts-latency-zero-gpu-what-voice-agents-actually-need-in-production-23h9"&gt;full benchmark&lt;/a&gt;). The platform also supports zero-shot voice cloning from an eight-second reference.&lt;/p&gt;

&lt;p&gt;The deliberate pairing, admitted again: every comparison in this post was chosen by us and favors us where it can. Bundles differ. Check both pricing pages before deciding.&lt;/p&gt;

&lt;h2&gt;
  
  
  The agent case, also about 20 lines
&lt;/h2&gt;

&lt;p&gt;If your ElevenLabs usage is Conversational AI rather than plain TTS, the equivalent object is &lt;code&gt;VoiceAgentClient&lt;/code&gt;. One WebSocket carries speech recognition, the language model, and the voice, with barge-in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;lokutor&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;VoiceAgentClient&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;VoiceStyle&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Language&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;VoiceAgentClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;LOKUTOR_API_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are the receptionist for a dental clinic. Keep answers short.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;voice&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;VoiceStyle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;F4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;# Maria
&lt;/span&gt;    &lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;Language&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;SPANISH&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;transcription&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;user:&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;response&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;agent:&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;start_conversation&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Tool calling, mid-session prompt updates, and transcripts are constructor parameters and methods on the same client. The &lt;a href="https://docs.lokutor.com/sdks/python/reference" rel="noopener noreferrer"&gt;Python SDK reference&lt;/a&gt; has the full surface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Price, from the public pages
&lt;/h2&gt;

&lt;p&gt;ElevenLabs Agents lists 8 cents per minute beyond a plan's allocation, and low-latency TTS as low as 5 cents per minute on the $990-per-month Business tier (their pricing page, 25 September 2026). Lokutor is about 2 cents per minute with recognition, the language model, and the voice included, 1.4 cents on Business, with 60 free minutes a month (details in &lt;a href="https://dev.to/danivs10/voice-agents-for-2c-a-minute-everything-included-1n5j"&gt;this pricing breakdown&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  When not to migrate
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;You need 32 languages or thousands of library voices.&lt;/li&gt;
&lt;li&gt;You rely on dubbing, sound effects, or the voice changer.&lt;/li&gt;
&lt;li&gt;You have a specific cloned voice in ElevenLabs that your product depends on.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If none of those apply and your cost per minute is becoming a problem, get a free key at &lt;a href="https://app.lokutor.com" rel="noopener noreferrer"&gt;app.lokutor.com&lt;/a&gt;, run both providers on the same script, and keep whichever wins your traffic. Docs at &lt;a href="https://docs.lokutor.com" rel="noopener noreferrer"&gt;docs.lokutor.com&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Try the voices in your browser on &lt;a href="https://huggingface.co/spaces/lokutor-ai/tts-on-cpu" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt;. Lokutor 2.0 launches on &lt;a href="https://www.producthunt.com/products/lokutor-2-0?launch=lokutor-2-0" rel="noopener noreferrer"&gt;Product Hunt&lt;/a&gt; October 13.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>voiceai</category>
      <category>tts</category>
      <category>voiceagents</category>
      <category>python</category>
    </item>
    <item>
      <title>ElevenLabs alternative for real-time voice agents: CPU-native, 2 cents a minute</title>
      <dc:creator>Daniel Varela</dc:creator>
      <pubDate>Fri, 25 Sep 2026 18:28:24 +0000</pubDate>
      <link>https://dev.to/danivs10/elevenlabs-alternative-for-real-time-voice-agents-cpu-native-2-cents-a-minute-28em</link>
      <guid>https://dev.to/danivs10/elevenlabs-alternative-for-real-time-voice-agents-cpu-native-2-cents-a-minute-28em</guid>
      <description>&lt;p&gt;ElevenLabs is the default answer for synthetic voices, and for good reason: a large voice library, strong voice cloning, dubbing, and a broad creative suite. If you are producing narration, dubbing video, or exploring voices, it is the right tool.&lt;/p&gt;

&lt;p&gt;If what you are building is a real-time voice agent, the requirements change. You need streaming time to first audio that survives a phone call, a per-minute price that holds up at call volume, and deployment that does not depend on GPU availability. That is the case for evaluating Lokutor as an ElevenLabs alternative.&lt;/p&gt;

&lt;h2&gt;
  
  
  What ElevenLabs costs, from its pricing page
&lt;/h2&gt;

&lt;p&gt;Read on 25 September 2026:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Free: $0, 10k credits per month&lt;/li&gt;
&lt;li&gt;Starter: $6 per month, 30k credits&lt;/li&gt;
&lt;li&gt;Scale: $299 per month, 1.8M credits&lt;/li&gt;
&lt;li&gt;Business: $990 per month, 6M credits, with low-latency TTS as low as 5 cents per minute at that tier&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;ElevenLabs Agents lists 8 cents per minute beyond a plan's allocation (read 23 September 2026). Credits, bundles, and what one minute buys differ between providers, so treat these as list rates rather than a like-for-like benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Lokutor does differently
&lt;/h2&gt;

&lt;p&gt;One price of about 2 cents per minute with speech recognition, the language model, and the voice included (1.4 cents on the Business plan). The free plan includes 60 minutes a month. The voice model, Versa, is a few tens of millions of parameters and runs on commodity CPUs, which is where the price comes from.&lt;/p&gt;

&lt;p&gt;Latency from production traffic, single CPU thread: 139 ms median streaming time to first audio, 211 ms p90. At the full agent level, caller-observed time to first audio is 385 to 442 ms in the best case and about 660 ms typical.&lt;/p&gt;

&lt;p&gt;Deployment runs on x86 and Arm CPUs: hosted API, private VPC, on-premises, or edge. In a measured AMD Genoa capacity test, one node handled about 19 concurrent calls, roughly $22 of infrastructure per concurrent stream per month.&lt;/p&gt;

&lt;p&gt;Languages: Spanish, Catalan, Galician, Basque, English, Portuguese, French, Italian, and German. Zero-shot voice cloning from an eight-second reference.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where ElevenLabs still wins
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Voice variety and marketplace breadth&lt;/li&gt;
&lt;li&gt;Dubbing and creative tooling&lt;/li&gt;
&lt;li&gt;Ecosystem and brand recognition&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your product is content production, stay put. If your product answers phone calls and the GPU bill is becoming a problem, test us.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evaluate in an hour
&lt;/h2&gt;

&lt;p&gt;Get a free API key at &lt;a href="https://app.lokutor.com" rel="noopener noreferrer"&gt;app.lokutor.com&lt;/a&gt;, 60 minutes included. Docs at &lt;a href="https://docs.lokutor.com" rel="noopener noreferrer"&gt;docs.lokutor.com&lt;/a&gt;. Run it against your current provider on the same script, cold and warm, and keep whichever wins your traffic.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Try the voices in your browser on &lt;a href="https://huggingface.co/spaces/lokutor-ai/tts-on-cpu" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt;. Lokutor 2.0 launches on &lt;a href="https://www.producthunt.com/products/lokutor-2-0?launch=lokutor-2-0" rel="noopener noreferrer"&gt;Product Hunt&lt;/a&gt; October 13.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>voiceai</category>
      <category>voiceagents</category>
      <category>tts</category>
      <category>ai</category>
    </item>
    <item>
      <title>Voice agents on CPU vs the GPU incumbents: latency, cost, and a deliberate comparison</title>
      <dc:creator>Daniel Varela</dc:creator>
      <pubDate>Fri, 25 Sep 2026 18:28:20 +0000</pubDate>
      <link>https://dev.to/danivs10/voice-agents-on-cpu-vs-the-gpu-incumbents-latency-cost-and-a-deliberate-comparison-hge</link>
      <guid>https://dev.to/danivs10/voice-agents-on-cpu-vs-the-gpu-incumbents-latency-cost-and-a-deliberate-comparison-hge</guid>
      <description>&lt;p&gt;Two numbers frame this post. On Lokutor, a production voice agent starts speaking in 139 ms at the median (TTS streaming time to first audio, one CPU thread), and a full minute of a running agent costs about 2 cents with speech recognition, the language model, and the voice included. On the GPU incumbents' own pricing pages, a comparable minute lists at 6 to 8 cents and covers only part of the stack.&lt;/p&gt;

&lt;p&gt;We picked this pairing on purpose, and it favors us. Below are the numbers, the sources, and the places where the comparison breaks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Latency, measured in production
&lt;/h2&gt;

&lt;p&gt;Versa 2.0, our TTS model, on live customer-facing infrastructure, single CPU thread (c8g.2xlarge, Graviton4):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;74 ms fastest streaming time to first audio&lt;/li&gt;
&lt;li&gt;139 ms median TTFA&lt;/li&gt;
&lt;li&gt;211 ms p90&lt;/li&gt;
&lt;li&gt;0.26 to 0.36 real-time factor&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At the full agent level (turn detection, language model, voice, transport), caller-observed time to first audio is 385 to 442 ms in the best case and about 660 ms typical. In three &lt;a href="https://dev.to/danivs10/three-real-sessions-on-a-cpu-voice-agent-api-d5o"&gt;recorded demo sessions&lt;/a&gt;, the server-side figure from end of turn to first audio came out between 67 and 207 ms.&lt;/p&gt;

&lt;p&gt;The GPU reference point, from &lt;a href="https://dev.to/danivs10/near-gpu-tts-latency-zero-gpu-what-voice-agents-actually-need-in-production-23h9"&gt;our earlier benchmark&lt;/a&gt;: Kokoro 82M on a spot RTX 4090 generated a short phrase in 47 ms warm. That is faster than Versa, as a good GPU model should be. Two caveats keep it honest: the 47 ms is full-utterance generation in a setup that did not stream, and the same setup needed 14.9 seconds to load the model. A CPU does not beat a 4090 on raw warm inference. The question is which infrastructure you want to operate when traffic is bursty and capacity goes cold.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost, from public pricing pages
&lt;/h2&gt;

&lt;p&gt;Each company's own pay-as-you-go rate, read from their pricing pages on 23 September 2026 (details and billing rules in &lt;a href="https://dev.to/danivs10/voice-agents-for-2c-a-minute-everything-included-1n5j"&gt;this breakdown&lt;/a&gt;):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Platform&lt;/th&gt;
&lt;th&gt;Published price per minute&lt;/th&gt;
&lt;th&gt;What that covers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Lokutor&lt;/td&gt;
&lt;td&gt;~2 cents (1.4 cents on Business)&lt;/td&gt;
&lt;td&gt;Recognition, language model, and voice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cartesia&lt;/td&gt;
&lt;td&gt;6 cents&lt;/td&gt;
&lt;td&gt;Call duration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deepgram Voice Agent&lt;/td&gt;
&lt;td&gt;7.5 cents&lt;/td&gt;
&lt;td&gt;Standard tier, pay as you go&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ElevenLabs Agents&lt;/td&gt;
&lt;td&gt;8 cents&lt;/td&gt;
&lt;td&gt;Per minute beyond a plan's allocation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Bundles differ, so this is not a like-for-like price benchmark. It is still the comparison buyers make. A Lokutor minute includes the model and the voice; the other rates cover part of the stack, so the gap widens once you add recognition and the language model, and it narrows or flips if you already operate your own models.&lt;/p&gt;

&lt;p&gt;Capacity: in a measured AMD Genoa test, one node handled about 19 concurrent calls, equal to roughly $22 of infrastructure per concurrent stream per month. That was a capacity test, not our current production deployment, so treat it as a capacity result rather than a guarantee.&lt;/p&gt;

&lt;p&gt;Quality is gated, not assumed: English WER fell from 6.51% to 3.20% between the previous Versa version and 2.0, mean WER across ten benchmarks is 11.21%, and English UTMOS is 3.14. WER from a recognizer is a proxy for intelligibility, not proof of quality; we use it to catch pronunciation regressions before latency work ships.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this comparison breaks
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Warm GPU TTS wins on raw speed. If you already operate GPU infrastructure well and need maximum throughput per node, use it.&lt;/li&gt;
&lt;li&gt;List prices move. Check the current pages before deciding.&lt;/li&gt;
&lt;li&gt;Bundles differ: support, languages, telephony, compliance, and what one minute buys.&lt;/li&gt;
&lt;li&gt;Our median is a production median on our traffic. Yours will differ.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The claim is narrow: interactive latency and production economics for voice agents, on commodity CPUs, with no accelerator in the serving path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;The free plan includes 60 minutes a month. Get an API key at &lt;a href="https://app.lokutor.com" rel="noopener noreferrer"&gt;app.lokutor.com&lt;/a&gt; and read the docs at &lt;a href="https://docs.lokutor.com" rel="noopener noreferrer"&gt;docs.lokutor.com&lt;/a&gt;. Benchmark cold and warm paths, then send us the numbers. We would rather compare real serving conditions than trade screenshots of ideal runs.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Try the voices in your browser on &lt;a href="https://huggingface.co/spaces/lokutor-ai/tts-on-cpu" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt;. Lokutor 2.0 launches on &lt;a href="https://www.producthunt.com/products/lokutor-2-0?launch=lokutor-2-0" rel="noopener noreferrer"&gt;Product Hunt&lt;/a&gt; October 13.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>voiceai</category>
      <category>voiceagents</category>
      <category>tts</category>
      <category>cpu</category>
    </item>
    <item>
      <title>Three real sessions on a CPU voice agent API</title>
      <dc:creator>Daniel Varela</dc:creator>
      <pubDate>Thu, 24 Sep 2026 23:08:46 +0000</pubDate>
      <link>https://dev.to/danivs10/three-real-sessions-on-a-cpu-voice-agent-api-d5o</link>
      <guid>https://dev.to/danivs10/three-real-sessions-on-a-cpu-voice-agent-api-d5o</guid>
      <description>&lt;p&gt;Three recorded sessions against the &lt;a href="https://lokutor.com/" rel="noopener noreferrer"&gt;Lokutor&lt;/a&gt; voice agent API. Speech-to-text, LLM, and voice all run on CPU. The millisecond figure on screen is measured server-side: first audio after end of turn detected.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Booking a dentist, in Spanish
&lt;/h2&gt;

&lt;p&gt;Two agents on the same API: a caller and a clinic receptionist. The caller asks for a cleaning this week, gets offered Thursday 5:30 PM or Friday 10 AM, picks Thursday, asks the price (45 euros), and books. Latency shown per turn: 207 ms, 206 ms, 144 ms. 40 seconds.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/NB2og_AatkM" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  2. One voice, four languages
&lt;/h2&gt;

&lt;p&gt;Same agent, same voice, switching live between Spanish, Catalan, Galician, and Basque. Latency on screen: 204 ms in Spanish, 144 ms in Catalan, 67 ms in Galician. The API also handles English, Portuguese, French, Italian, and German. 31 seconds.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/tMPdaZy7TlA" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Two agents debating
&lt;/h2&gt;

&lt;p&gt;Two Lokutor agents debate in Spanish whether AI will end humanity. Unscripted: each one hears the other and answers on its own. Humanity survives, for now. 36 seconds.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/x-r9woMbf0o" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;Free API key at &lt;a href="https://app.lokutor.com" rel="noopener noreferrer"&gt;app.lokutor.com&lt;/a&gt;, docs at &lt;a href="https://docs.lokutor.com" rel="noopener noreferrer"&gt;docs.lokutor.com&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Try the voices in your browser on &lt;a href="https://huggingface.co/spaces/lokutor-ai/tts-on-cpu" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt;. Lokutor 2.0 launches on &lt;a href="https://www.producthunt.com/products/lokutor-2-0?launch=lokutor-2-0" rel="noopener noreferrer"&gt;Product Hunt&lt;/a&gt; October 13.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>voiceai</category>
      <category>voiceagents</category>
      <category>ai</category>
      <category>tts</category>
    </item>
    <item>
      <title>Voice Agents for 2¢ a Minute, Everything Included</title>
      <dc:creator>Daniel Varela</dc:creator>
      <pubDate>Wed, 23 Sep 2026 03:02:46 +0000</pubDate>
      <link>https://dev.to/danivs10/voice-agents-for-2c-a-minute-everything-included-1n5j</link>
      <guid>https://dev.to/danivs10/voice-agents-for-2c-a-minute-everything-included-1n5j</guid>
      <description>&lt;p&gt;From today, a Lokutor voice agent costs about 2¢ a minute, with speech recognition, the language model and the voice all included. On the Business plan it is 1.4¢. There is no separate fee for the model, no separate fee for the voice, and a free plan with 60 minutes a month to try it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What 2¢ buys, next to what the category charges
&lt;/h2&gt;

&lt;p&gt;Most voice-agent platforms publish a per-minute price that covers only part of the call. Here is each company's own pay-as-you-go rate, read from their pricing pages on 23 September 2026:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Platform&lt;/th&gt;
&lt;th&gt;Published price per minute&lt;/th&gt;
&lt;th&gt;What that covers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Lokutor&lt;/td&gt;
&lt;td&gt;1.4 to 2¢&lt;/td&gt;
&lt;td&gt;Everything: recognition, language model, voice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cartesia&lt;/td&gt;
&lt;td&gt;6¢&lt;/td&gt;
&lt;td&gt;Call duration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deepgram Voice Agent&lt;/td&gt;
&lt;td&gt;7.5¢&lt;/td&gt;
&lt;td&gt;Standard tier, pay as you go&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ElevenLabs Agents&lt;/td&gt;
&lt;td&gt;8¢&lt;/td&gt;
&lt;td&gt;Per minute beyond a plan's allocation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retell AI&lt;/td&gt;
&lt;td&gt;5.5¢ + LLM + voice&lt;/td&gt;
&lt;td&gt;Voice engine; model and TTS billed on top&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vapi&lt;/td&gt;
&lt;td&gt;5¢ + providers&lt;/td&gt;
&lt;td&gt;Platform fee; recognition, model and TTS billed on top&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The whole Lokutor minute costs less than most platforms' fee before you have chosen a model.&lt;/p&gt;

&lt;h2&gt;
  
  
  The plans
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;th&gt;Minutes included&lt;/th&gt;
&lt;th&gt;Per minute&lt;/th&gt;
&lt;th&gt;Concurrent calls&lt;/th&gt;
&lt;th&gt;Beyond the allowance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Starter&lt;/td&gt;
&lt;td&gt;$19/mo&lt;/td&gt;
&lt;td&gt;1,000&lt;/td&gt;
&lt;td&gt;1.9¢&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;2¢&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Growth&lt;/td&gt;
&lt;td&gt;$99/mo&lt;/td&gt;
&lt;td&gt;6,000&lt;/td&gt;
&lt;td&gt;1.65¢&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;1.8¢&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Business&lt;/td&gt;
&lt;td&gt;$349/mo&lt;/td&gt;
&lt;td&gt;25,000&lt;/td&gt;
&lt;td&gt;1.4¢&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;1.5¢&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every plan also includes text-to-speech and speech-to-text API allowances. 12 million characters of synthesis on Growth works out at about $8 per million, where the premium voice APIs charge $30 to $100. Growth and Business add voice cloning and phone numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it can be this cheap
&lt;/h2&gt;

&lt;p&gt;Because the models are small. Our voice model, Versa, is a few tens of millions of parameters, not billions. Our turn-taking and noise-suppression models are smaller still. Most of the category serves models tens of times larger on scarce GPUs, and prices the minute to match. Ours are built to run on commodity hardware, which is also why the same stack can run on your own servers or on a device. This is not an introductory price that doubles next quarter; it is what the architecture costs to run, plus a business.&lt;/p&gt;

&lt;h2&gt;
  
  
  How you are billed, all of it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The minute is measured per second of connected session, and includes recognition, the language model and the voice. Unused minutes don't roll over.&lt;/li&gt;
&lt;li&gt;Phone calls count as two minutes per minute on the line, to cover the carrier: 3.3¢ on Growth, 2.8¢ on Business. Numbers are $5 a month.&lt;/li&gt;
&lt;li&gt;Web search is off unless you turn it on for an agent; each search counts as four minutes.&lt;/li&gt;
&lt;li&gt;At your limit, calls stop until next month, with no surprise bill. Or turn on overage and keep going at your plan's rate, up to one extra month's minutes, billed with your next invoice.&lt;/li&gt;
&lt;li&gt;Changing plan: upgrade any time and a new billing month starts that day, less credit for what's left of the old one. A smaller plan applies straight away.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That list is the entire bill. It is on the &lt;a href="https://lokutor.com/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt; too, with an estimator that shows the maths for your volume.&lt;/p&gt;

&lt;h2&gt;
  
  
  Built in Spain, native in the languages others skip
&lt;/h2&gt;

&lt;p&gt;Lokutor speaks nine languages: Spanish, Catalan, Galician, Basque, Portuguese, English, French, Italian and German. Catalan, Galician and Basque are first-class, not an afterthought: if you serve customers across Iberia, you don't have to choose a platform that only does Spanish. And if your audio can't leave your infrastructure, the whole pipeline runs on-premise in the EU.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://app.lokutor.com" rel="noopener noreferrer"&gt;Start free with 60 minutes, no card&lt;/a&gt;, or &lt;a href="https://calendly.com/lokutor/reunion-lokutor" rel="noopener noreferrer"&gt;talk to us&lt;/a&gt; about volume.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Try the voices in your browser on &lt;a href="https://huggingface.co/spaces/lokutor-ai/tts-on-cpu" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt;. Lokutor 2.0 launches on &lt;a href="https://www.producthunt.com/products/lokutor-2-0?launch=lokutor-2-0" rel="noopener noreferrer"&gt;Product Hunt&lt;/a&gt; October 13.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>voiceai</category>
      <category>voiceagents</category>
      <category>pricing</category>
    </item>
    <item>
      <title>Near-GPU TTS latency, zero GPU: what voice agents actually need in production</title>
      <dc:creator>Daniel Varela</dc:creator>
      <pubDate>Sun, 20 Sep 2026 16:08:42 +0000</pubDate>
      <link>https://dev.to/danivs10/near-gpu-tts-latency-zero-gpu-what-voice-agents-actually-need-in-production-23h9</link>
      <guid>https://dev.to/danivs10/near-gpu-tts-latency-zero-gpu-what-voice-agents-actually-need-in-production-23h9</guid>
      <description>&lt;p&gt;Versa 2.0 has produced first audio in 74 ms on one CPU thread.&lt;/p&gt;

&lt;p&gt;In our RTX 4090 test, Kokoro's first cold synthesis took 2.44 seconds, with 14.9 seconds of model load time.&lt;/p&gt;

&lt;p&gt;That is the attractive comparison. It is also incomplete.&lt;/p&gt;

&lt;p&gt;Once warm, Kokoro generated a short phrase in 47 ms. It was faster than Versa, as a good GPU model should be. But that 47 ms was full-utterance generation, not streaming time to first audio. Our Kokoro setup did not stream. Versa's 74 ms figure is its fastest streaming time to first audio, while its production median is 139 ms.&lt;/p&gt;

&lt;p&gt;These numbers are not an apples-to-apples model race. They expose the production decision teams actually face:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;optimize for the fastest warm benchmark on a GPU;&lt;/li&gt;
&lt;li&gt;or get interactive latency on ordinary CPU infrastructure, without making a GPU part of the serving path.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For many voice agents, the second option is more useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmark, with the conditions left in
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Versa 2.0&lt;/th&gt;
&lt;th&gt;Kokoro 82M&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hardware&lt;/td&gt;
&lt;td&gt;c8g.2xlarge, Graviton4 CPU&lt;/td&gt;
&lt;td&gt;RTX 4090 spot instance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Execution&lt;/td&gt;
&lt;td&gt;Single CPU thread&lt;/td&gt;
&lt;td&gt;GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model size&lt;/td&gt;
&lt;td&gt;~58M parameters&lt;/td&gt;
&lt;td&gt;82M parameters&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cold model load&lt;/td&gt;
&lt;td&gt;Already resident in production&lt;/td&gt;
&lt;td&gt;14.9 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cold first synthesis&lt;/td&gt;
&lt;td&gt;Streaming service path&lt;/td&gt;
&lt;td&gt;2.44 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fastest observed latency&lt;/td&gt;
&lt;td&gt;74 ms to first audio&lt;/td&gt;
&lt;td&gt;47 ms, warm short phrase&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Production / warm reference&lt;/td&gt;
&lt;td&gt;139 ms median TTFA, 211 ms p90&lt;/td&gt;
&lt;td&gt;47 ms short, 124 ms medium, 132 ms long&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Real-time factor&lt;/td&gt;
&lt;td&gt;0.26-0.36&lt;/td&gt;
&lt;td&gt;0.02 warm&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Streaming in tested setup&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Versa figures are from live customer-facing infrastructure, serving calls at six flow-matching steps. Kokoro was run through our test harness on a spot RTX 4090 supplied through &lt;a href="https://amics.ai" rel="noopener noreferrer"&gt;amics.ai&lt;/a&gt;, created by Alvaro Fragoso.&lt;/p&gt;

&lt;p&gt;The comparison needs two warnings:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Versa reports streaming time to first audio. Kokoro reports the time to finish generating the full utterance because the tested setup did not stream.&lt;/li&gt;
&lt;li&gt;The 74 ms and 2.44 s figures are deliberately the best Versa observation and the cold Kokoro observation. For steady-state throughput, warm Kokoro wins comfortably.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The point is not that a CPU beats a 4090. It does not. The point is that a voice product is more than its best warm inference number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cold starts are a product metric
&lt;/h2&gt;

&lt;p&gt;A benchmark often assumes the model is loaded, the accelerator is allocated, kernels are warm, and traffic is steady. Production traffic does not always behave that way.&lt;/p&gt;

&lt;p&gt;Voice agents are bursty. A campaign begins. A queue goes from zero to hundreds of calls. A worker is replaced. A region fails over. An autoscaler adds capacity. A low-traffic language has been idle.&lt;/p&gt;

&lt;p&gt;In those moments, model load and first-request latency become user experience.&lt;/p&gt;

&lt;p&gt;A 47 ms warm result can coexist with a 14.9 second load. Both are true. Only one appears in most benchmark headlines.&lt;/p&gt;

&lt;p&gt;This changes how teams should test TTS:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cold path = provision + model load + first synthesis
warm path = request arrival + first audio + generation
recovery path = replacement worker + readiness + first successful request
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Report all three. If a system scales to zero, cold latency belongs in the product SLO. If it never scales to zero, the cost of keeping it warm belongs in the infrastructure plan.&lt;/p&gt;

&lt;p&gt;CPU inference changes that tradeoff. CPU capacity is easier to find, easier to autoscale, and does not require a separate accelerator pool. A team can keep latency low without reserving a GPU for every serving unit or designing around GPU availability.&lt;/p&gt;

&lt;h2&gt;
  
  
  Time to first audio matters more than time to finish
&lt;/h2&gt;

&lt;p&gt;For a voice agent, users do not wait for the whole sentence to be synthesized. They hear the first chunk while the rest is still being produced.&lt;/p&gt;

&lt;p&gt;That makes streaming time to first audio, or TTFA, the useful number. Full-utterance latency measures something else: how quickly a system can generate a completed file.&lt;/p&gt;

&lt;p&gt;Both metrics matter, but for different products:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Voice agents: TTFA, p90/p99 TTFA, chunk cadence, interruption behavior&lt;/li&gt;
&lt;li&gt;Audiobooks and batch generation: full-utterance latency and total throughput&lt;/li&gt;
&lt;li&gt;High-volume outbound systems: concurrency per node and recovery time&lt;/li&gt;
&lt;li&gt;On-device or private deployments: hardware requirements and memory footprint&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A system that generates a sentence in 47 ms but cannot emit audio until the sentence is complete may still feel slower than a system that starts speaking in 139 ms and continues streaming.&lt;/p&gt;

&lt;p&gt;That is why we report Versa's production median and tail, not only its best run:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;74 ms fastest time to first audio&lt;/li&gt;
&lt;li&gt;139 ms median&lt;/li&gt;
&lt;li&gt;211 ms p90&lt;/li&gt;
&lt;li&gt;0.26-0.36 real-time factor&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At the complete voice-agent level, our measured caller-observed time to first audio is 385-442 ms in the best case and about 660 ms typically. The model number matters, but the caller experiences the whole pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  What latency does a voice agent actually need?
&lt;/h2&gt;

&lt;p&gt;Past a certain point, shaving another 20 or 30 ms from isolated TTS inference has less effect than fixing the rest of the turn.&lt;/p&gt;

&lt;p&gt;The practical target is not "the lowest TTS number possible." It is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Start audio quickly enough that the exchange feels responsive.&lt;/li&gt;
&lt;li&gt;Keep p90 and p99 under control when traffic changes.&lt;/li&gt;
&lt;li&gt;Stream continuously without audible gaps.&lt;/li&gt;
&lt;li&gt;Leave room in the latency budget for turn detection, language-model output, transport, and telephony.&lt;/li&gt;
&lt;li&gt;Recover without multi-second stalls when capacity changes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Teams should benchmark the system in the state users will hit, not only a warmed-up notebook.&lt;/p&gt;

&lt;p&gt;A useful test matrix looks like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;What to record&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;First request on a new worker&lt;/td&gt;
&lt;td&gt;Provisioning, load, TTFA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Warm short reply&lt;/td&gt;
&lt;td&gt;TTFA, full generation time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Warm long reply&lt;/td&gt;
&lt;td&gt;TTFA, chunk cadence, RTF&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Burst from idle&lt;/td&gt;
&lt;td&gt;p50, p90, p99, failed requests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sustained concurrency&lt;/td&gt;
&lt;td&gt;Calls per node, tail latency, CPU/GPU utilization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Worker replacement&lt;/td&gt;
&lt;td&gt;Time until healthy traffic resumes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Latency is not enough if speech quality breaks
&lt;/h2&gt;

&lt;p&gt;Fast speech still has to be intelligible.&lt;/p&gt;

&lt;p&gt;We use held-out text and transcribe synthesized output with the same recognizer used on genuine human recordings. This does not prove that synthetic speech is "better than humans." WER is a proxy for intelligibility, and a recognizer can have its own biases. It does give us a repeatable way to catch pronunciation regressions.&lt;/p&gt;

&lt;p&gt;From the previous Versa version to Versa 2.0:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;English WER fell from 6.51% to 3.20%.&lt;/li&gt;
&lt;li&gt;Mean WER across ten benchmarks fell from 14.03% to 11.21%.&lt;/li&gt;
&lt;li&gt;On the held-out German, Italian, Catalan, and Basque sets, synthesized speech had lower WER than the genuine human comparator when both were scored by the same recognizer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each language set contained roughly 60 to 200 samples. We also measure an English UTMOS score of 3.14. Versa supports zero-shot voice cloning from an eight-second reference.&lt;/p&gt;

&lt;p&gt;The important part is not one quality number. It is that latency work should be gated by repeatable quality checks. Otherwise a faster checkpoint may simply be speaking less clearly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The infrastructure question behind the benchmark
&lt;/h2&gt;

&lt;p&gt;GPU TTS can be extremely fast. Kokoro's 0.02 warm RTF on the 4090 makes that clear.&lt;/p&gt;

&lt;p&gt;The tradeoff is operational dependency:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;accelerator availability by region;&lt;/li&gt;
&lt;li&gt;warm capacity during quiet periods;&lt;/li&gt;
&lt;li&gt;GPU-aware scheduling and autoscaling;&lt;/li&gt;
&lt;li&gt;recovery when a worker disappears;&lt;/li&gt;
&lt;li&gt;separate deployment paths for cloud, private VPC, edge, or on-prem environments.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Versa runs on x86 and Arm CPUs. In a measured AMD Genoa test, one node handled about 19 concurrent calls, equal to roughly $22 of infrastructure per concurrent stream per month. That test is not our current production deployment, so we treat it as a capacity result, not a production guarantee.&lt;/p&gt;

&lt;p&gt;The hosted API follows the same idea. Public voice-agent rates range from $0.035 to $0.116 per minute depending on plan. For directional context, current public list rates for ElevenLabs and Deepgram are around $0.08 per minute. Feature bundles and billing units differ, so this is not a like-for-like price benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  When CPU-native TTS is the better choice
&lt;/h2&gt;

&lt;p&gt;Use a GPU model when maximum warm throughput is the main constraint and you already operate GPU infrastructure well.&lt;/p&gt;

&lt;p&gt;CPU-native TTS is worth testing when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;you want one deployment model across cloud and private infrastructure;&lt;/li&gt;
&lt;li&gt;traffic is bursty and cold capacity matters;&lt;/li&gt;
&lt;li&gt;GPU availability or regional coverage is a constraint;&lt;/li&gt;
&lt;li&gt;you need low-latency streaming without maintaining an accelerator fleet;&lt;/li&gt;
&lt;li&gt;voice is part of the product, but GPU operations should not become part of the company.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is the problem we built Versa 2.0 to solve: near-GPU interactive latency, on a single CPU thread, in a production streaming API.&lt;/p&gt;

&lt;p&gt;Try the live demo at &lt;a href="https://lokutor.com" rel="noopener noreferrer"&gt;lokutor.com&lt;/a&gt;, or get a free API key at &lt;a href="https://app.lokutor.com" rel="noopener noreferrer"&gt;app.lokutor.com&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you benchmark it, test both cold and warm paths. Send us the numbers. We would rather compare real serving conditions than trade screenshots of ideal runs.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Try the voices in your browser on &lt;a href="https://huggingface.co/spaces/lokutor-ai/tts-on-cpu" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt;. Lokutor 2.0 launches on &lt;a href="https://www.producthunt.com/products/lokutor-2-0?launch=lokutor-2-0" rel="noopener noreferrer"&gt;Product Hunt&lt;/a&gt; October 13.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>voiceai</category>
      <category>tts</category>
      <category>cpu</category>
      <category>buildinpublic</category>
    </item>
    <item>
      <title>Edge Voice AI: Why Voice Interfaces Are Moving Off the Cloud</title>
      <dc:creator>Daniel Varela</dc:creator>
      <pubDate>Fri, 18 Sep 2026 16:23:09 +0000</pubDate>
      <link>https://dev.to/danivs10/edge-voice-ai-why-voice-interfaces-are-moving-off-the-cloud-4l0e</link>
      <guid>https://dev.to/danivs10/edge-voice-ai-why-voice-interfaces-are-moving-off-the-cloud-4l0e</guid>
      <description>&lt;p&gt;Every voice interface shipped in the last decade has made the same architectural bet: send audio to the cloud, run inference on a GPU cluster, send audio back. That bet made sense when voice AI lived inside a phone app or a smart speaker with a permanent Wi-Fi connection and a company willing to subsidize the inference cost. It stops making sense the moment voice becomes the interface for a robot, a wearable, or an industrial device sold once and used for years.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lifetime cost problem
&lt;/h2&gt;

&lt;p&gt;A cloud-dependent voice stack doesn't just cost money to build. It costs money for as long as the product exists. Sell a voice-enabled device once, and GPU-based cloud inference means paying a recurring bill for every unit, indefinitely. That math works for a subscription software product. It works poorly for hardware, where the sale happens once but the voice interface needs to keep working for years.&lt;/p&gt;

&lt;p&gt;This is the core reason robotics and wearable companies are increasingly asking whether voice AI can run on the device, or on infrastructure they control, instead of a third party's GPU cloud.&lt;/p&gt;

&lt;h2&gt;
  
  
  Latency and connectivity aren't optional
&lt;/h2&gt;

&lt;p&gt;Real-time conversation has a biological rhythm to it. Human turn-taking gaps average around 200 milliseconds. A round-trip to a cloud GPU cluster, especially over an unreliable connection such as a warehouse robot, a car, or a wearable outdoors, fights against that rhythm directly. Voice interfaces that depend on connectivity to function at all are also voice interfaces that stop working exactly when connectivity is worst, which for a physical device is often the moment it matters most.&lt;/p&gt;

&lt;p&gt;Edge and on-device inference removes the network hop by design. It doesn't guarantee low latency on its own because the model still has to be fast, but it removes the one latency source a company has zero control over.&lt;/p&gt;

&lt;h2&gt;
  
  
  Privacy is becoming a distribution advantage, not a compliance checkbox
&lt;/h2&gt;

&lt;p&gt;For consumer devices, "does my voice data leave this device" is an increasingly direct purchasing question. For enterprise and regulated buyers in healthcare, finance, and government, it's a procurement requirement, not a preference. The GDPR and the EU AI Act both push toward systems where data processing is auditable and, ideally, contained within the customer's own environment.&lt;/p&gt;

&lt;p&gt;A voice AI stack that can run entirely inside a customer's perimeter, rather than routing audio through an external GPU cloud, turns compliance from a limitation into a sales advantage. That's the case for regulated buyers specifically: an in-perimeter, auditable voice stack shortens procurement instead of complicating it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What has to be true for this to work
&lt;/h2&gt;

&lt;p&gt;Edge voice AI isn't just "take the cloud model and shrink it." Autoregressive transformer models, the dominant architecture behind most cloud TTS and a lot of cloud STT, are expensive to run well without GPU acceleration. That's precisely why the cloud GPU pattern became the default. Making voice AI genuinely edge-viable means designing models differently from the start:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Non-autoregressive generation where possible, so inference is parallelizable instead of sequential&lt;/li&gt;
&lt;li&gt;Small parameter counts by design, not as an afterthought quantization pass&lt;/li&gt;
&lt;li&gt;Architectures like ConvNeXt blocks instead of full transformer attention stacks, trading some of the attention mechanism's overhead for efficiency at a fraction of the compute&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is why "runs on CPU" and "runs on the edge" tend to be the same underlying engineering problem: both require models built to be small and efficient from the ground up, not GPU-scale models with the GPU removed after the fact.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this goes next
&lt;/h2&gt;

&lt;p&gt;The near-term path for edge voice AI looks like three overlapping waves: a developer API and platform running on CPU infrastructure today, private enterprise deployments that run entirely inside a customer's own environment, and eventually on-device SDKs for robots, wearables, and appliances that need voice without a permanent cloud dependency. Each step removes another layer of GPU dependency and another reason voice AI has to live in someone else's data center.&lt;/p&gt;

&lt;p&gt;The industry's default architecture assumed infinite, cheap GPU capacity would always be available to whoever needed it. For a wearable that has to work for years on a fixed hardware budget, that assumption was never going to hold. The interesting engineering problem isn't "how do we get a bigger GPU," it's "how do we need one less."&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Try the voices in your browser on &lt;a href="https://huggingface.co/spaces/lokutor-ai/tts-on-cpu" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt;. Lokutor 2.0 launches on &lt;a href="https://www.producthunt.com/products/lokutor-2-0?launch=lokutor-2-0" rel="noopener noreferrer"&gt;Product Hunt&lt;/a&gt; October 13.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>voiceai</category>
      <category>edgecomputing</category>
      <category>cpu</category>
    </item>
  </channel>
</rss>
