<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Smallest AI</title>
    <description>The latest articles on DEV Community by Smallest AI (@smallestai).</description>
    <link>https://dev.to/smallestai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3854927%2F8dabc078-5e46-402d-ad31-71a95a7a510b.png</url>
      <title>DEV Community: Smallest AI</title>
      <link>https://dev.to/smallestai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/smallestai"/>
    <language>en</language>
    <item>
      <title>Why Your AI Voice Assistant Gets Lost When Users Switch Languages</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Fri, 25 Sep 2026 07:19:26 +0000</pubDate>
      <link>https://dev.to/smallestai/why-your-ai-voice-assistant-gets-lost-when-users-switch-languages-4nk2</link>
      <guid>https://dev.to/smallestai/why-your-ai-voice-assistant-gets-lost-when-users-switch-languages-4nk2</guid>
      <description>&lt;p&gt;&lt;em&gt;Follow one mixed-language request through speech recognition, language routing, response generation, and playback.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A user asks a booking assistant, “Can you move my reservation para mañana?” The request begins in English and finishes in Spanish. Imagine that the reply arrives in English, or that it pronounces the Spanish phrase awkwardly. The user might describe the assistant as having a bad voice or not understanding the request.&lt;/p&gt;

&lt;p&gt;Either symptom could have several causes. Speech recognition might have lost the Spanish words, the application might have selected the wrong response language, or the speech model might not handle the mixed-language output well. Even a correct answer can feel disjointed if switching models adds a noticeable pause. Before replacing a voice, a developer needs to find out which handoff failed.&lt;/p&gt;

&lt;p&gt;That is the central problem in a multilingual voice assistant: speech-to-text (STT), a language model, text-to-speech (TTS), and the orchestration layer all need a consistent view of the language being used. Let's follow this one request through the pipeline and see what each component needs from the next.&lt;/p&gt;

&lt;h2&gt;
  
  
  A language choice travels through the entire pipeline
&lt;/h2&gt;

&lt;p&gt;The microphone supplies audio. STT converts it into a transcript. A language model or application logic interprets the request and decides what to say. TTS converts that answer into audio. Orchestration routes the work, preserves conversation state, and handles errors, retries, and playback.&lt;/p&gt;

&lt;p&gt;The booking request makes these familiar components harder to coordinate. The first few words point toward English. The ending is Spanish, and the intended response language may depend on the rest of the conversation. If the first component quietly commits to English, the other components may inherit that guess as though it were a fact.&lt;/p&gt;

&lt;p&gt;A useful design treats language as session context that can be revised, not a fixed setting inferred from the first word. The components can still have different capabilities, but the orchestration layer needs to know when a language decision is provisional and when the application has enough evidence to use it.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-a-multilingual-voice-assistant"&gt;Smallest AI homepage&lt;/a&gt; provides an overview of the speech products available for this kind of application. For a broader breakdown of the individual stages, Smallest AI's &lt;a href="https://smallest.ai/blog/designing-voice-assistants-stt-llm-tts-tools-and-latency-budget?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-a-multilingual-voice-assistant"&gt;guide to STT, LLM, TTS, tools, and latency budgets&lt;/a&gt; explains the underlying voice architecture. Here, the extra complication is what happens when those stages disagree about language.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first language guess is not the final answer
&lt;/h2&gt;

&lt;h3&gt;
  
  
  An early audio guess is provisional
&lt;/h3&gt;

&lt;p&gt;A system can identify language from incoming audio before it transcribes, or infer language after obtaining text. An early audio-level guess can help configure speech recognition, but it can also commit too soon. Transcript-based detection has more textual evidence, although waiting for it can add work to the response path. Neither approach is automatically more accurate for every language pair or audio condition.&lt;/p&gt;

&lt;p&gt;For “Can you move my reservation para mañana?”, an early detector might report English with high confidence before the user reaches the Spanish phrase. A hybrid approach uses that early result as a starting signal, then checks the transcript and revises the language context if later words change the picture. It must also represent uncertainty when a single utterance contains more than one language.&lt;/p&gt;

&lt;h3&gt;
  
  
  Code-switching changes the language decision
&lt;/h3&gt;

&lt;p&gt;Mixing languages within one utterance is code-switching, and it is not the same as a user choosing a new language for the entire session. A useful test is whether the system preserves both parts of the utterance without forcing them into one language. If it does not, downstream language routing has little chance of recovering the original request.&lt;/p&gt;

&lt;p&gt;The right decision may also depend on deployment details. A multilingual detection mode may cover only a particular language group, and the languages supported by a model's streaming interface may differ from its prerecorded interface. Check the current language matrix and regional availability for the exact STT mode you intend to deploy rather than assuming a generic “multilingual” label covers every combination.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the transcript before changing the response model
&lt;/h2&gt;

&lt;p&gt;Once the utterance has ended, inspect what STT actually produced. Did it preserve “para mañana”? Did it interpret the switch as speech rather than noise, and did it retain the booking intent? A fluent response cannot compensate for important information that never reached the application.&lt;/p&gt;

&lt;p&gt;Evaluate the recognition layer using recordings from the accents, dialects, environments, and language combinations your users actually bring. Word error rate (WER) can expose transcription differences between languages, but an aggregate WER can conceal poor performance on a smaller supported group. Measure each language separately and keep dedicated code-switching tests.&lt;/p&gt;

&lt;p&gt;Also separate live transcription from batch transcription. A prerecorded transcription result may look good while a streaming assistant finalizes the words too late for the conversation. Check the latency to partial output and to a usable final transcript under realistic concurrency. Early text may help later stages begin work, but an unstable partial transcript should not trigger an irreversible action such as changing a booking.&lt;/p&gt;

&lt;p&gt;Smallest AI's &lt;a href="https://smallest.ai/speech-to-text?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-a-multilingual-voice-assistant"&gt;Pulse speech-to-text&lt;/a&gt; offers multilingual transcription and real-time streaming. Its supported languages and detection modes vary by interface and, for some languages, by region. Treat that coverage as a configuration decision to verify for your deployment, not as proof that every language pair will perform equally well.&lt;/p&gt;

&lt;h2&gt;
  
  
  The language model needs a response policy, not just a language tag
&lt;/h2&gt;

&lt;p&gt;Suppose the transcript is correct. The system still needs to choose a response language. It could use one multilingual model for all supported languages or route each turn to a language-specific model. The first reduces the number of routing decisions to manage; the second can be useful when a particular language requires a specialized model. Both approaches need an explicit policy for mixed-language turns.&lt;/p&gt;

&lt;p&gt;For the booking example, a system instruction could tell the model to consider the user's most recent message and the established session language, and to ask for clarification when the desired response language is ambiguous. That is an application policy, not a guarantee that the model will always follow it. Keep tests for a user who changes languages mid-session and for a user who mixes them within one sentence.&lt;/p&gt;

&lt;p&gt;Language routing introduces a different failure mode. If the detection layer labels the request incorrectly, the router may send a perfectly transcribed sentence to a model that is poorly suited to it. Store the detected language, the selected model, and the reason for the routing choice together so that a wrong-language reply can be traced back to the decision that produced it.&lt;/p&gt;

&lt;p&gt;For an intent-oriented benchmark, the &lt;a href="https://arxiv.org/abs/2212.06346" rel="noopener noreferrer"&gt;MASSIVE dataset and MMNLU-22 workshop report&lt;/a&gt; cover multilingual intent classification and slot filling across 52 languages. They provide useful evaluation context, but a benchmark result does not replace testing on your application's mixed-language audio and tasks.&lt;/p&gt;

&lt;h2&gt;
  
  
  A correct answer can still sound wrong in the target language
&lt;/h2&gt;

&lt;p&gt;After the application determines the answer, TTS has to speak it. This is where teams sometimes mistake a language mismatch for a poor voice. If the text contains Spanish while the chosen voice or model is tuned mainly for English, pronunciation, rhythm, and stress can make a correct answer sound out of place.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fylu5l8mul6cgsg8uvdfo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fylu5l8mul6cgsg8uvdfo.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
Test candidate voices in every target language with native speakers where possible. Include numbers, local names, domain terms, and mixed-language phrases from real workflows. Check more than a generic naturalness score: the words must be intelligible, the intended language must be recognizable, and the persona should remain appropriate to the product.&lt;/p&gt;

&lt;p&gt;For voice identity across channels or locales, &lt;a href="https://smallest.ai/voice-cloning?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-a-multilingual-voice-assistant"&gt;Smallest AI voice cloning&lt;/a&gt; is another option to evaluate. Cloning a voice does not, by itself, establish native pronunciation in every language. Verify the model, language, and voice combination you plan to use. For a related but less interactive workflow, the &lt;a href="https://smallest.ai/blog/multilingual-voice-dubbing-for-product-videos-how-to-localize-audio-without-re-recording?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-a-multilingual-voice-assistant"&gt;multilingual dubbing guide&lt;/a&gt; explores what changes when speech is localized for recorded content.&lt;/p&gt;
&lt;h2&gt;
  
  
  The handoffs decide how long the user waits
&lt;/h2&gt;

&lt;p&gt;Return to the booking request. Even if every stage produces the correct content, an extra language-detection pass, a model switch, or loading a different TTS voice can delay the first audible reply. The orchestrator must account for those costs alongside recognition, reasoning, synthesis, transport, and client playback.&lt;/p&gt;
&lt;h3&gt;
  
  
  Keep language context flexible while work overlaps
&lt;/h3&gt;

&lt;p&gt;A few choices are worth evaluating in the actual application:&lt;/p&gt;

&lt;p&gt;• &lt;strong&gt;Reuse language context, but let it change.&lt;/strong&gt; Cache the established language for the session without treating it as immutable when the user switches.&lt;/p&gt;

&lt;p&gt;• &lt;strong&gt;Overlap work carefully.&lt;/strong&gt; Streaming STT may let the application start preparing for a response, but provisional words and language guesses need a correction path. Waiting for final input may be safer for consequential actions.&lt;/p&gt;

&lt;p&gt;• &lt;strong&gt;Keep frequently used voices ready when supported.&lt;/strong&gt; Preloading can avoid a cold-start penalty, but it consumes resources and is an application-level choice, not a universal API capability.&lt;/p&gt;

&lt;p&gt;• &lt;strong&gt;Measure separate timestamps.&lt;/strong&gt; Capture the end of user speech, usable transcript, response generation, first synthesized audio, and first client playback. A fast model does not establish a fast end-to-end turn.&lt;/p&gt;
&lt;h3&gt;
  
  
  Measure the wait at the point of playback
&lt;/h3&gt;

&lt;p&gt;Set budgets from observed behavior in your own channels and languages. A fixed stage target lifted from another deployment may hide transport, region, buffering, or playback delays. Compare monolingual turns with code-switched turns so that language-specific overhead becomes visible rather than disappearing into an average.&lt;/p&gt;
&lt;h2&gt;
  
  
  Create and store the API key
&lt;/h2&gt;

&lt;p&gt;When you are ready to test the speech layer with the &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-a-multilingual-voice-assistant"&gt;Smallest AI API&lt;/a&gt;, keep authentication on the server. Before running an authenticated request, &lt;a href="https://docs.smallest.ai/voice-agents/platform/account/api-keys?utm_source=dev.to&amp;amp;utm_medium=Vizup&amp;amp;utm_campaign=how-to-build-a-multilingual-voice-assistant"&gt;create a Smallest.ai API key in the dashboard&lt;/a&gt; and store it in the SMALLEST_API_KEY environment variable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key-here"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Authenticated requests use the Authorization header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer &amp;lt;SMALLEST_API_KEY value&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Never place the key in browser JavaScript, mobile app code, a public repository, query parameters, screenshots, or client-side logs. Use a server-side secrets manager for production. The source article is an architecture guide rather than a working API tutorial, so it does not supply an endpoint-specific request example to reproduce here.&lt;/p&gt;

&lt;h2&gt;
  
  
  The languages with less training data need a different test plan
&lt;/h2&gt;

&lt;p&gt;The booking example uses two widely supported languages. A rollout to languages with less available training material may expose different recognition and synthesis limitations. Do not assume that a provider's strongest language-specific result transfers to Swahili, Bengali, Tagalog, or a regional dialect.&lt;/p&gt;

&lt;p&gt;Transfer learning and domain-specific fine-tuning can help where the chosen model offers that capability and suitable data is available. But the quantity of audio needed, the improvement achieved, and whether voice cloning works well are model- and language-dependent. Evaluate on representative consented recordings instead of assuming that a small reference sample or a few hours of fine-tuning will close every gap.&lt;/p&gt;

&lt;p&gt;That makes rollout scope an engineering decision. Begin with the languages your users need most, establish a quality baseline for each, and expand only when the acceptance tests cover that language's real tasks and failure cases.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run the test that reveals which layer failed
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Replay the language boundaries
&lt;/h3&gt;

&lt;p&gt;Before launch, record monolingual turns in every supported language, turns with deliberate code-switching, accented speech, and a request in an unsupported language. Then replay the same cases after changing a model, prompt, or voice. WER and task-level correctness help find regressions, while listening with native speakers catches unnatural pronunciation and tone that those metrics may miss.&lt;/p&gt;

&lt;h3&gt;
  
  
  Trace one failed turn to its originating stage
&lt;/h3&gt;

&lt;p&gt;For the original booking request, inspect the transcript first, then the detected language and routing decision, then the generated answer, the selected TTS voice, and finally the audio heard by the user. If the transcript lost the Spanish phrase, replacing the voice will not repair recognition. If the words and answer were correct but the pronunciation failed, investigate synthesis. If everything was accurate but the reply arrived late, inspect the time between stages and the actual playback start.&lt;/p&gt;

&lt;p&gt;A multilingual assistant becomes dependable when those failures can be separated and tested rather than dismissed as the voice sounding wrong. You can &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-a-multilingual-voice-assistant"&gt;start building with the Smallest AI API&lt;/a&gt; and evaluate your own language combinations, audio, and end-to-end response time before rolling the experience out more broadly.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>speechrecognition</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Your Speech Recognition Demo Works. So Why Does It Fail in Production?</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Fri, 25 Sep 2026 07:08:12 +0000</pubDate>
      <link>https://dev.to/smallestai/your-speech-recognition-demo-works-so-why-does-it-fail-in-production-3j4c</link>
      <guid>https://dev.to/smallestai/your-speech-recognition-demo-works-so-why-does-it-fail-in-production-3j4c</guid>
      <description>&lt;p&gt;At first, everything looks fine. The microphone works, the browser returns text, and the transcript appears almost immediately. But when the user says something longer, the text changes halfway through the sentence. An identifier gets transcribed incorrectly, or the recording stops when the network connection drops.&lt;/p&gt;

&lt;p&gt;The problem is not necessarily the speech recognition model. Your application also has to manage audio capture, interim results, browser compatibility, and the path between transcription and the action the user wants to perform.&lt;/p&gt;

&lt;p&gt;A working transcript is only the beginning.&lt;/p&gt;

&lt;p&gt;To build this feature properly, we'll start with the browser's built-in recognition API, examine where it becomes limiting, and then look at a server-side streaming architecture. Along the way, we'll address authentication, audio formats, transcript state, and the failures that tend to appear outside controlled development environments.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, decide what the browser needs to recognize
&lt;/h2&gt;

&lt;p&gt;Suppose our application lets users search for orders by speaking rather than typing.&lt;/p&gt;

&lt;p&gt;The browser captures audio, speech recognition turns it into text, and the application uses that text to search its backend.&lt;/p&gt;

&lt;p&gt;The distinction between speech recognition and voice recognition matters here. Speech recognition determines what someone said. Voice recognition concerns who is speaking. Our search feature needs the words, not the speaker's identity.&lt;/p&gt;

&lt;p&gt;The underlying technology is automatic speech recognition (ASR). It processes incoming audio and estimates the spoken text. Depending on the implementation, recognition can run locally, through a browser-managed service, or through an API connected to your application.&lt;/p&gt;

&lt;p&gt;That gives us two practical starting points:&lt;/p&gt;

&lt;p&gt;• Browser-native recognition: Let the browser manage speech recognition and return the transcript through JavaScript events.&lt;/p&gt;

&lt;p&gt;• Server-side ASR: Capture audio in the browser, send it to your backend, and use a dedicated recognition service to produce the transcript.&lt;/p&gt;

&lt;p&gt;Neither approach eliminates the need to handle application state. The difference is how much control you have over the recognition pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the Web Speech API
&lt;/h2&gt;

&lt;p&gt;For our first implementation, assume the application is an internal tool and the development team controls the supported browser.&lt;/p&gt;

&lt;p&gt;The Web Speech API is a reasonable place to start. It provides SpeechRecognition for converting speech into text and SpeechSynthesis for generating spoken output. We only need recognition.&lt;/p&gt;

&lt;p&gt;Create a simple HTML interface:&lt;/p&gt;

&lt;p&gt;HTML&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;button&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"startBtn"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;Start speaking&lt;span class="nt"&gt;&amp;lt;/button&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;button&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"stopBtn"&lt;/span&gt; &lt;span class="na"&gt;disabled&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;Stop&lt;span class="nt"&gt;&amp;lt;/button&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;p&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"status"&lt;/span&gt; &lt;span class="na"&gt;role=&lt;/span&gt;&lt;span class="s"&gt;"status"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;Ready&lt;span class="nt"&gt;&amp;lt;/p&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;p&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;"output"&lt;/span&gt; &lt;span class="na"&gt;aria-live=&lt;/span&gt;&lt;span class="s"&gt;"polite"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&amp;lt;/p&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Next, connect the buttons to the recognition API. The important detail is that recognition can produce multiple results. Some are provisional, while others are marked final. We should not treat every update as a new permanent line of text.&lt;/p&gt;

&lt;p&gt;JAVASCRIPT&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;startBtn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getElementById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;startBtn&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;stopBtn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getElementById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;stopBtn&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getElementById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;output&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getElementById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;status&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;Recognition&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;SpeechRecognition&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;webkitSpeechRecognition&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;Recognition&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Speech recognition is unavailable. Please use text input.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="nx"&gt;startBtn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;disabled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;recognition&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Recognition&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="nx"&gt;recognition&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;lang&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;en-US&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;recognition&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;interimResults&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;recognition&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;maxAlternatives&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="nx"&gt;recognition&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onstart&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Listening...&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;startBtn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;disabled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;stopBtn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;disabled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;

  &lt;span class="nx"&gt;recognition&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onresult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;confirmed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;interim&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;results&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;isFinal&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;confirmed&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;interim&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;confirmed&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;interim&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;

  &lt;span class="nx"&gt;recognition&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onerror&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
      &lt;span class="s2"&gt;`Recognition error: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;

  &lt;span class="nx"&gt;recognition&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onend&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Listening...&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Stopped&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="nx"&gt;startBtn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;disabled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;stopBtn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;disabled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;

  &lt;span class="nx"&gt;startBtn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;addEventListener&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;click&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;recognition&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;start&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="nx"&gt;stopBtn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;addEventListener&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;click&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;recognition&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stop&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This corrects a subtle problem in many minimal examples: reading only one transcript without accounting for its finality.&lt;/p&gt;

&lt;p&gt;When the user speaks, an interim result may change as the recognizer receives more audio. The UI should show that progress without repeatedly appending the same words.&lt;/p&gt;

&lt;p&gt;Before treating the example as complete, test it in your target browsers. &lt;a href="https://developer.mozilla.org/en-US/docs/Web/API/SpeechRecognition" rel="noopener noreferrer"&gt;MDN's SpeechRecognition reference&lt;/a&gt; documents the API and its browser support limitations.&lt;/p&gt;

&lt;p&gt;Chrome's conventional speech recognition implementation sends audio to a remote recognition service. Browser-native therefore does not necessarily mean offline or private to the user's device. Some implementations also provide experimental on-device recognition features, but availability and language-pack requirements need separate verification.&lt;/p&gt;

&lt;p&gt;For a controlled prototype, those limitations may be acceptable. For a customer-facing feature that needs more predictable control over the audio pipeline, we need another architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move audio capture out of the recognition API
&lt;/h2&gt;

&lt;p&gt;Now imagine extending the order-search feature into a customer-facing application.&lt;/p&gt;

&lt;p&gt;Users arrive with different browsers, microphones, accents, and network conditions. You may also need to evaluate different recognition providers against your application's vocabulary.&lt;/p&gt;

&lt;p&gt;A useful architectural separation is to let the browser handle capture and the backend handle recognition.&lt;/p&gt;

&lt;p&gt;The flow becomes:&lt;/p&gt;

&lt;p&gt;Browser microphone → application WebSocket → backend audio processing → ASR service → transcript → browser UI&lt;/p&gt;

&lt;p&gt;The frontend no longer needs to know which ASR model your backend uses. This also creates an important security boundary. Your application can authenticate the user, manage the recognition session, and keep the speech provider's API credentials on the server.&lt;/p&gt;

&lt;p&gt;For example, &lt;a href="https://smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-integrate-speech-recognition-into-a-web-app"&gt;Smallest AI&lt;/a&gt; offers the &lt;a href="https://smallest.ai/speech-to-text?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-integrate-speech-recognition-into-a-web-app"&gt;Pulse speech-to-text service&lt;/a&gt;, which supports streaming transcription through WebSocket. It can serve as the recognition layer behind this architecture.&lt;/p&gt;

&lt;p&gt;The browser still needs to send compatible audio. It cannot assume that whatever MediaRecorder produces will be accepted directly by the ASR service.&lt;/p&gt;

&lt;p&gt;That is the next implementation decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Capture microphone audio without assuming every browser uses the same format
&lt;/h2&gt;

&lt;p&gt;The browser's getUserMedia() API provides access to the microphone after the user grants permission. MediaRecorder can then divide that audio into chunks for transmission.&lt;/p&gt;

&lt;p&gt;The following example implements the browser-capture portion of our application. It assumes that you have configured an application-owned WebSocket endpoint at wss://YOUR_APP_HOST/asr. That endpoint must be implemented on your backend. It is not a Smallest AI API URL.&lt;/p&gt;

&lt;p&gt;JAVASCRIPT&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;startCapture&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;mimeType&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;audio/webm;codecs=opus&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;MediaRecorder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isTypeSupported&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;mimeType&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;This browser does not support the selected audio format.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nb"&gt;navigator&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;mediaDevices&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getUserMedia&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;recorder&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;MediaRecorder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;mimeType&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;socket&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;WebSocket&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;wss://YOUR_APP_HOST/asr&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onopen&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;recorder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;start&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;250&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;

    &lt;span class="nx"&gt;recorder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ondataavailable&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;size&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt;
        &lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;readyState&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nx"&gt;WebSocket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;OPEN&lt;/span&gt;
      &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;

    &lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onmessage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;transcript&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getElementById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;output&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
            &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Invalid transcript message&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;

    &lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onerror&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getElementById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;status&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
        &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;The transcription connection failed.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;

    &lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onclose&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;recorder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;inactive&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;recorder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stop&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;

      &lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getTracks&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;forEach&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;track&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;track&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stop&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nf"&gt;stop&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;recorder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;inactive&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="nx"&gt;recorder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stop&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getTracks&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;forEach&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;track&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;track&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stop&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;

        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;readyState&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nx"&gt;WebSocket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;OPEN&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;getTracks&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;forEach&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;track&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;track&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stop&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The 250 ms recording interval is an illustrative configuration inherited from the capture pattern. It is not a guaranteed chunk duration or an optimal value for every ASR provider.&lt;/p&gt;

&lt;p&gt;There is also an audio-format distinction worth preserving. audio/webm;codecs=opus describes Opus audio in a WebM container. An ASR service that supports Opus or Ogg Opus does not automatically accept WebM-wrapped audio.&lt;/p&gt;

&lt;p&gt;For Pulse, the official streaming audio specifications document supported encodings, sample rates, and channel requirements.&lt;/p&gt;

&lt;p&gt;Your backend must deliver audio that matches those requirements. One documented option is 16 kHz, mono, 16-bit PCM using the linear16 encoding.&lt;/p&gt;

&lt;p&gt;If you choose that format, you need a compatible browser capture path or a backend conversion step. Changing the encoding parameter in the API URL does not convert the audio itself.&lt;/p&gt;

&lt;p&gt;The code above is the browser-facing transport layer, not a complete provider integration. It also needs an application-specific user-authentication mechanism and production reconnection policy.&lt;/p&gt;

&lt;p&gt;Before wiring these pieces together, it is worth testing the ASR connection independently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create and store the API key
&lt;/h2&gt;

&lt;p&gt;Keep the API key in an environment variable rather than hard-coding it into the application.&lt;/p&gt;

&lt;p&gt;Before running the snippet, &lt;a href="https://docs.smallest.ai/voice-agents/platform/account/api-keys?utm_source=dev.to&amp;amp;utm_medium=Vizup&amp;amp;utm_campaign=how-to-integrate-speech-recognition-into-a-web-app"&gt;create a Smallest.ai API key in the dashboard&lt;/a&gt; and store it in the SMALLEST_API_KEY environment variable.&lt;/p&gt;

&lt;p&gt;BASH&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key-here"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every authenticated Smallest AI request sends the value through the Authorization header:&lt;/p&gt;

&lt;p&gt;TEXT&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer &amp;lt;SMALLEST_API_KEY value&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the key on your server. Do not expose it in browser JavaScript, mobile application code, public repositories, screenshots, query parameters, or client-side logs.&lt;/p&gt;

&lt;p&gt;For production deployments, use a server-side secrets manager and restrict access to the application components that need the credential.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3eee7p9b9fl24l2dg6ok.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3eee7p9b9fl24l2dg6ok.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the streaming ASR connection independently
&lt;/h2&gt;

&lt;p&gt;Before connecting live browser audio to the backend, verify that your server can authenticate with the recognition service and receive transcript messages.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-integrate-speech-recognition-into-a-web-app"&gt;Smallest AI API&lt;/a&gt; provides access to Pulse. Its documented streaming endpoint is:&lt;/p&gt;

&lt;p&gt;TEXT&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;wss://api.smallest.ai/waves/v1/stt/live?model=pulse
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The following Node.js test uses a local file containing raw 16 kHz, mono, signed 16-bit PCM audio. It is an isolated provider test, not the complete browser relay.&lt;/p&gt;

&lt;p&gt;Install the WebSocket dependency:&lt;/p&gt;

&lt;p&gt;BASH&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install &lt;/span&gt;ws
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your source is a WAV file, you can prepare the raw PCM test input locally with FFmpeg:&lt;/p&gt;

&lt;p&gt;BASH&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; sample.wav &lt;span class="nt"&gt;-ar&lt;/span&gt; 16000 &lt;span class="nt"&gt;-ac&lt;/span&gt; 1 &lt;span class="nt"&gt;-f&lt;/span&gt; s16le sample-16k-mono-s16le.raw
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This creates headerless PCM. Sending the original WAV bytes while declaring linear16 would include container data that the raw streaming example does not expect.&lt;/p&gt;

&lt;p&gt;Before running the snippet, &lt;a href="https://docs.smallest.ai/voice-agents/platform/account/api-keys?utm_source=dev.to&amp;amp;utm_medium=Vizup&amp;amp;utm_campaign=how-to-integrate-speech-recognition-into-a-web-app"&gt;create a Smallest.ai API key in the dashboard&lt;/a&gt; and store it in the SMALLEST_API_KEY environment variable.&lt;/p&gt;

&lt;p&gt;JAVASCRIPT&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// save as test-pulse.cjs&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;fs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;node:fs&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;once&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;node:events&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;sleep&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;node:timers/promises&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;WebSocket&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ws&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;apiKey&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SMALLEST_API_KEY is not set&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;audioFile&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;sample-16k-mono-s16le.raw&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sampleRate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;16000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;URL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;wss://api.smallest.ai/waves/v1/stt/live?model=pulse&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;searchParams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;language&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;en&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;searchParams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;encoding&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;linear16&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;searchParams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;sample_rate&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sampleRate&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;socket&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;WebSocket&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;Authorization&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Bearer &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;message&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;

    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;transcript&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;is_final&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Final:&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Interim:&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;transcript&lt;/span&gt;
      &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;is_last&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Unexpected response format&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;error&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Speech recognition connection failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;once&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;open&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;fs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createReadStream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;audioFile&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;highWaterMark&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;4096&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="k"&gt;await &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;readyState&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="nx"&gt;WebSocket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;OPEN&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;WebSocket connection closed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="c1"&gt;// Approximate real-time pacing for 16-bit mono PCM.&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunkDurationMs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
      &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sampleRate&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunkDurationMs&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;close_stream&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="p"&gt;}));&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="k"&gt;catch&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;socket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;terminate&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;exitCode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it with:&lt;/p&gt;

&lt;p&gt;BASH&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node test-pulse.cjs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Check whether the service returns transcript messages and marks completed segments with is_final.&lt;/p&gt;

&lt;p&gt;The close_stream message signals that audio transmission has finished. The documented is_last field identifies the final session response.&lt;/p&gt;

&lt;p&gt;This test isolates authentication, audio formatting, and provider response handling. It does not prove that your browser capture and backend relay are correct.&lt;/p&gt;

&lt;p&gt;Once it passes, you can connect the application-owned WebSocket to Pulse, forwarding appropriately converted audio upstream and transcript messages downstream. That relay needs its own session lifecycle, user authorization, resource limits, and failure handling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep interim text separate from confirmed text
&lt;/h2&gt;

&lt;p&gt;Return to the user searching for an order number.&lt;/p&gt;

&lt;p&gt;If the interface replaces the entire field with every incoming partial transcript, the displayed text can appear to flicker. If it permanently appends each interim update, words can appear more than once.&lt;/p&gt;

&lt;p&gt;The application needs two distinct states: confirmed text and text that may still change.&lt;/p&gt;

&lt;p&gt;For a streaming service that emits separate transcript segments, the application can maintain:&lt;/p&gt;

&lt;p&gt;JAVASCRIPT&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;confirmed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;interim&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;handleTranscript&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;is_final&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;confirmed&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;transcript&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;interim&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;interim&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getElementById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;output&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
    &lt;span class="nx"&gt;confirmed&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;interim&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This assumes the service emits finalized segments rather than repeatedly sending the entire cumulative transcript. Confirm the exact response semantics before using it with another provider.&lt;/p&gt;

&lt;p&gt;The important engineering rule is independent of the provider: do not let provisional text trigger irreversible application behavior.&lt;/p&gt;

&lt;p&gt;For our order-search example, showing a partial transcript is fine. Submitting the search or changing an order based on an incomplete identifier is a different decision.&lt;/p&gt;

&lt;p&gt;A command-oriented application may use a confirmed segment or an explicit user confirmation before performing an action.&lt;/p&gt;

&lt;p&gt;If your provider offers word-level confidence or timestamps, those can help with review and alignment. They should not be assumed to exist in every response.&lt;/p&gt;

&lt;p&gt;For more background on ASR failure modes, Smallest AI's &lt;a href="https://smallest.ai/blog/ai-speech-recognition-challenges-accents-noise-asr?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-integrate-speech-recognition-into-a-web-app"&gt;guide to accents, noise, and speech recognition challenges&lt;/a&gt; provides additional evaluation context.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat microphone permission as part of the feature
&lt;/h2&gt;

&lt;p&gt;Our example begins with a button click for a reason.&lt;/p&gt;

&lt;p&gt;Microphone access through getUserMedia() requires a secure context. In a deployed web application, that normally means HTTPS. Localhost is treated as a trustworthy origin for development.&lt;/p&gt;

&lt;p&gt;Users should understand why the application is asking for microphone access. Request permission when they choose to record, rather than attempting to open the microphone during page initialization.&lt;/p&gt;

&lt;p&gt;Once recording begins, the interface should make that state visible and provide an accessible stop control.&lt;/p&gt;

&lt;p&gt;The application also needs to distinguish failures that have different remedies:&lt;/p&gt;

&lt;p&gt;• Permission denied: Explain that microphone access is blocked and provide a text-input alternative.&lt;/p&gt;

&lt;p&gt;• Unsupported browser: Detect unavailable APIs and offer another input method.&lt;/p&gt;

&lt;p&gt;• Network interruption: Stop or suspend capture safely and make the incomplete transcript visible.&lt;/p&gt;

&lt;p&gt;• Audio-format mismatch: Check the actual audio encoding, sample rate, and channel count before changing recognition settings.&lt;/p&gt;

&lt;p&gt;• No-speech timeout: Let the user restart without losing previously confirmed text.&lt;/p&gt;

&lt;p&gt;WebSocket reconnection deserves particular care. Reopening a socket does not guarantee that the previous recognition session can resume.&lt;/p&gt;

&lt;p&gt;If your backend starts a fresh session, preserve confirmed application state and explicitly decide how to handle audio that was not processed. Do not silently replay a command that might already have triggered an action.&lt;/p&gt;

&lt;p&gt;Privacy also extends beyond microphone permission. If you retain transcripts or audio, document what is stored, how long it remains available, and how users can request deletion.&lt;/p&gt;

&lt;h2&gt;
  
  
  A transcript is not yet an application action
&lt;/h2&gt;

&lt;p&gt;Once recognition works, our order-search feature still needs to interpret the text.&lt;/p&gt;

&lt;p&gt;The full application path is:&lt;/p&gt;

&lt;p&gt;Audio capture → ASR → transcript processing → intent or entity extraction → application action&lt;/p&gt;

&lt;p&gt;Not every application needs a language model after transcription. For a simple voice-search field, ordinary search logic may be sufficient. A command interface might use keyword or rule-based intent matching. A more conversational application may need additional language processing.&lt;/p&gt;

&lt;p&gt;The important boundary is between recognizing words and deciding what those words mean for your application.&lt;/p&gt;

&lt;p&gt;Post-processing can also include punctuation restoration, normalization, speaker labeling, or domain-specific vocabulary handling, depending on the recognition provider and the requirements of the task.&lt;/p&gt;

&lt;p&gt;An incorrect identifier should not be silently corrected into a plausible but different identifier. If the application needs an exact order number, confirmation may be safer than guessing.&lt;/p&gt;

&lt;p&gt;That is why end-to-end evaluation should include the action produced from the transcript, not only whether individual words were recognized correctly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test with the audio your users will actually produce
&lt;/h2&gt;

&lt;p&gt;The final test should reproduce the conditions that made our initial implementation unreliable.&lt;/p&gt;

&lt;p&gt;Record a small, permissioned evaluation set containing representative order IDs, names, commands, and phrases. Include the accents, background noise, microphones, and network conditions your product is expected to handle.&lt;/p&gt;

&lt;p&gt;Then measure the stages separately:&lt;/p&gt;

&lt;p&gt;• Time from speech capture to the first transcript update.&lt;/p&gt;

&lt;p&gt;• Time until the relevant transcript segment is finalized.&lt;/p&gt;

&lt;p&gt;• Recognition accuracy on application-specific vocabulary.&lt;/p&gt;

&lt;p&gt;• Frequency of incorrect or duplicated confirmed text.&lt;/p&gt;

&lt;p&gt;• Behavior during microphone denial, connection loss, and unsupported formats.&lt;/p&gt;

&lt;p&gt;• Whether downstream actions use the intended final transcript.&lt;/p&gt;

&lt;p&gt;Word error rate can help compare recognition outputs against reference transcripts, but generic benchmarks cannot tell you whether a particular order-search workflow is reliable.&lt;/p&gt;

&lt;p&gt;You also need to distinguish recognition delay from UI update delay and backend processing time. Improving one component does not automatically fix the complete user experience.&lt;/p&gt;

&lt;p&gt;The browser-native implementation may be enough if your supported environment is narrow and the feature is low-stakes. A dedicated streaming ASR service offers another integration path when you need more control over capture, processing, and provider selection.&lt;/p&gt;

&lt;p&gt;Either way, the architecture should make failures observable rather than hiding them behind a single error message.&lt;/p&gt;

&lt;p&gt;Return to the user speaking an order number. The goal is not merely to display text while the microphone is active. It is to preserve the correct identifier, present the result clearly, and let the user complete the search.&lt;/p&gt;

&lt;p&gt;If you're evaluating a server-side approach, you can &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-integrate-speech-recognition-into-a-web-app"&gt;start building with the Smallest AI API&lt;/a&gt; and test Pulse with your own audio before integrating it into your web application's capture and transcript-handling workflow.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>javascript</category>
      <category>speechrecognition</category>
    </item>
    <item>
      <title>Building Production-Ready AI Voice SDK Integrations in Python, Browser, and Mobile</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Wed, 16 Sep 2026 10:48:49 +0000</pubDate>
      <link>https://dev.to/smallestai/building-production-ready-ai-voice-sdk-integrations-in-python-browser-and-mobile-2jji</link>
      <guid>https://dev.to/smallestai/building-production-ready-ai-voice-sdk-integrations-in-python-browser-and-mobile-2jji</guid>
      <description>&lt;p&gt;Getting speech working in a demo is usually the easy part.&lt;/p&gt;

&lt;p&gt;The harder problems show up after the microphone is live, audio is moving continuously, the network drops for a moment, a mobile operating system interrupts playback, or a browser refuses to start audio without a user gesture.&lt;/p&gt;

&lt;p&gt;That is where the quality of a voice SDK starts to matter.&lt;/p&gt;

&lt;p&gt;An AI voice SDK gives developers a language or platform-specific layer over speech capabilities such as text-to-speech, speech-to-text, streaming audio, and voice cloning. The underlying API defines what the service can do. The SDK determines how comfortably your application can use it.&lt;/p&gt;

&lt;p&gt;Whether you are building a Python service, a browser application, or a native mobile experience, the important question is not simply whether an SDK exists. It is whether the integration can survive the environment in which you plan to run it.&lt;/p&gt;

&lt;p&gt;Platforms such as &lt;a href="https://smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-voice-sdks-for-python-browser-and-mobile-apps"&gt;Smallest AI&lt;/a&gt; expose speech capabilities through a developer platform, but the application architecture around those capabilities still determines security, responsiveness, and reliability.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the architecture, not the package manager
&lt;/h2&gt;

&lt;p&gt;It is tempting to evaluate voice SDKs by comparing installation commands, supported languages, and the number of methods in the client library.&lt;/p&gt;

&lt;p&gt;Those things matter, but they are rarely what causes trouble in production.&lt;/p&gt;

&lt;p&gt;Before choosing an SDK, ask what happens when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A streaming connection disappears midway through an utterance.&lt;/li&gt;
&lt;li&gt;Your application needs access to the raw audio buffer.&lt;/li&gt;
&lt;li&gt;The user interrupts synthesized speech.&lt;/li&gt;
&lt;li&gt;The microphone permission is denied or revoked.&lt;/li&gt;
&lt;li&gt;You replace a speech model later.&lt;/li&gt;
&lt;li&gt;The same feature has to work across backend, browser, and mobile environments.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three properties are especially important.&lt;/p&gt;

&lt;h3&gt;
  
  
  Streaming-first behavior
&lt;/h3&gt;

&lt;p&gt;Real-time speech should be treated as a continuous flow rather than a sequence of large synchronous requests.&lt;/p&gt;

&lt;p&gt;For speech recognition, incremental audio processing means the application can begin receiving transcription before the entire recording exists. For text-to-speech, streaming lets playback begin before the complete response has been synthesized.&lt;/p&gt;

&lt;p&gt;This is also why &lt;a href="https://smallest.ai/blog/why-streaming-architecture-is-non-negotiable-for-real-time-voice-agents?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-voice-sdks-for-python-browser-and-mobile-apps"&gt;streaming voice architectures&lt;/a&gt; matter so much in conversational applications. Waiting for each stage to finish before starting the next one adds latency throughout the pipeline.&lt;/p&gt;

&lt;h3&gt;
  
  
  Turn detection and segmentation
&lt;/h3&gt;

&lt;p&gt;Silence is not always the end of a turn.&lt;/p&gt;

&lt;p&gt;People pause while thinking, restart sentences, interrupt each other, and leave gaps between words. A conversational system needs enough control over segmentation and turn handling to avoid treating every short silence as the end of an utterance.&lt;/p&gt;

&lt;p&gt;If an SDK hides all of this behind a single high-level call, make sure it still exposes enough state for your application to react correctly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Portability without losing control
&lt;/h3&gt;

&lt;p&gt;A convenient abstraction is useful until you need something it does not expose.&lt;/p&gt;

&lt;p&gt;Look for an SDK that simplifies normal cases but still lets you control audio buffers, streaming state, errors, timeouts, and connection lifecycle when necessary.&lt;/p&gt;

&lt;p&gt;Portability also matters. If your Python backend, browser client, and mobile application use completely different integration models, feature drift becomes almost inevitable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Python: keep the speech loop asynchronous
&lt;/h2&gt;

&lt;p&gt;Python is still a natural starting point for many speech applications because the backend can safely own API credentials, queues, model calls, and application state.&lt;/p&gt;

&lt;p&gt;The basic TTS workflow is straightforward:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Send text and the relevant voice configuration.&lt;/li&gt;
&lt;li&gt;Receive synthesized audio.&lt;/li&gt;
&lt;li&gt;Stream that audio to the next component or persist it if the workflow is asynchronous.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Speech-to-text works in the opposite direction. Audio arrives from a microphone, uploaded recording, call stream, or another source, then gets sent to the transcription layer.&lt;/p&gt;

&lt;p&gt;For live applications, incremental processing is generally preferable to waiting for a complete recording. A &lt;a href="https://smallest.ai/speech-to-text?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-voice-sdks-for-python-browser-and-mobile-apps"&gt;speech-to-text API&lt;/a&gt; can sit behind the Python service while the application handles buffering, downstream events, and transcript state.&lt;/p&gt;

&lt;p&gt;A practical audio chunk size for real-time transcription is often somewhere around 20 ms to 100 ms, although the best value depends on the transport and speech service you are using.&lt;/p&gt;

&lt;p&gt;Smaller chunks reduce the amount of audio waiting in a buffer, but increase request or framing overhead. Larger chunks reduce overhead, but can add delay before downstream processing starts.&lt;/p&gt;

&lt;p&gt;Treat the chunk size as something to benchmark with your actual pipeline rather than a constant you copy from a tutorial.&lt;/p&gt;

&lt;p&gt;The same principle applies to TTS. If the application needs to speak while a response is still being produced, streaming audio is generally a better fit than waiting for one complete audio object.&lt;/p&gt;

&lt;p&gt;Your Python layer should also own:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Retry policy&lt;/li&gt;
&lt;li&gt;Request timeouts&lt;/li&gt;
&lt;li&gt;Backpressure&lt;/li&gt;
&lt;li&gt;Queueing&lt;/li&gt;
&lt;li&gt;Connection cleanup&lt;/li&gt;
&lt;li&gt;Error translation&lt;/li&gt;
&lt;li&gt;Observability&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That keeps transport behavior out of the rest of your application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create and store the API key
&lt;/h2&gt;

&lt;p&gt;For developers using Smallest AI as the server-side speech layer, the &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-voice-sdks-for-python-browser-and-mobile-apps"&gt;Smallest AI API&lt;/a&gt; provides the authenticated entry point for application integrations.&lt;/p&gt;

&lt;p&gt;Keep the API key in an environment variable rather than hard-coding it into the application.&lt;/p&gt;

&lt;p&gt;Before running the snippet, &lt;a href="https://docs.smallest.ai/models/api-reference/authentication?utm_source=dev.to&amp;amp;utm_medium=Vizup&amp;amp;utm_campaign=ai-voice-sdks-for-python-browser-and-mobile-apps"&gt;create a Smallest.ai API key in the dashboard&lt;/a&gt; and store it in the &lt;code&gt;SMALLEST_API_KEY&lt;/code&gt; environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key-here"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every authenticated request sends the value through the &lt;code&gt;Authorization&lt;/code&gt; header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer &amp;lt;SMALLEST_API_KEY value&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the key on your server.&lt;/p&gt;

&lt;p&gt;Do not expose it in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Browser JavaScript&lt;/li&gt;
&lt;li&gt;Client-side React code&lt;/li&gt;
&lt;li&gt;Native mobile application code&lt;/li&gt;
&lt;li&gt;Public repositories&lt;/li&gt;
&lt;li&gt;Screenshots&lt;/li&gt;
&lt;li&gt;Query parameters&lt;/li&gt;
&lt;li&gt;Client-side logs&lt;/li&gt;
&lt;li&gt;Error messages returned to users&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For production environments, store secrets in an appropriate server-side secrets manager rather than relying on values committed to application configuration.&lt;/p&gt;

&lt;p&gt;This separation becomes especially important once the same speech service is used by browser and mobile clients.&lt;/p&gt;

&lt;h2&gt;
  
  
  Browser: capture and playback on the client, credentials on the server
&lt;/h2&gt;

&lt;p&gt;Voice applications in the browser face a different set of constraints.&lt;/p&gt;

&lt;p&gt;You have capable primitives such as &lt;code&gt;MediaStream&lt;/code&gt;, the Web Audio API, and &lt;code&gt;AudioContext&lt;/code&gt;, but you also have permissions, autoplay policies, browser lifecycle behavior, and platform differences to deal with.&lt;/p&gt;

&lt;p&gt;A useful browser SDK should reduce those problems without preventing you from controlling the important parts of the audio pipeline.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prefer a real streaming transport
&lt;/h3&gt;

&lt;p&gt;Polling an HTTP endpoint repeatedly is usually a poor fit for real-time transcription.&lt;/p&gt;

&lt;p&gt;A persistent streaming connection, commonly a WebSocket, lets audio move continuously while partial results come back through the same session.&lt;/p&gt;

&lt;p&gt;The important part is not simply "does this SDK support WebSockets?" It is whether the SDK exposes connection state, reconnection behavior, partial output, cancellation, and errors clearly enough for your application.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl9f74ituvmk9j6m9s0d4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl9f74ituvmk9j6m9s0d4.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Check AudioWorklet support
&lt;/h3&gt;

&lt;p&gt;For low-latency audio processing, modern browser applications should generally use &lt;code&gt;AudioWorklet&lt;/code&gt; instead of relying on the older &lt;code&gt;ScriptProcessorNode&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;If an SDK handles microphone capture internally, inspect how that capture pipeline is implemented.&lt;/p&gt;

&lt;p&gt;You may eventually need direct control over resampling, buffering, channel layout, or custom voice-activity logic.&lt;/p&gt;

&lt;h3&gt;
  
  
  Account for autoplay policies
&lt;/h3&gt;

&lt;p&gt;Browsers often prevent audio playback until the user interacts with the page.&lt;/p&gt;

&lt;p&gt;Your application should establish the required audio context after an appropriate user gesture rather than assuming synthesized speech can begin automatically on page load.&lt;/p&gt;

&lt;h3&gt;
  
  
  Keep credentials out of the browser
&lt;/h3&gt;

&lt;p&gt;The browser should not hold a long-lived provider API key.&lt;/p&gt;

&lt;p&gt;A safer production architecture looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Browser
  |
  | microphone audio / application events
  v
Your backend
  |
  | authenticated speech requests
  v
Voice service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The browser handles capture, playback, and UI state.&lt;/p&gt;

&lt;p&gt;Your backend owns authentication and external API communication.&lt;/p&gt;

&lt;p&gt;This design has another benefit. If you replace the underlying speech provider or model later, the browser does not need to know.&lt;/p&gt;

&lt;h3&gt;
  
  
  Test mobile browsers separately
&lt;/h3&gt;

&lt;p&gt;Desktop Chrome working correctly tells you very little about how the same application will behave on iOS Safari.&lt;/p&gt;

&lt;p&gt;Microphone permissions, background behavior, playback rules, and audio session behavior vary enough that mobile browser testing should happen before the architecture is considered finished.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mobile: platform audio rules become part of the SDK contract
&lt;/h2&gt;

&lt;p&gt;Native mobile applications remove some browser constraints, but introduce a new set of operating-system rules.&lt;/p&gt;

&lt;p&gt;On iOS, &lt;code&gt;AVAudioSession&lt;/code&gt; configuration affects whether your application can record, play audio, coexist with other applications, and recover when the audio session is interrupted.&lt;/p&gt;

&lt;p&gt;A voice feature can appear stable during normal testing and then fail as soon as a phone call, assistant activation, route change, or another application takes control of the audio session.&lt;/p&gt;

&lt;p&gt;Android has a similar problem space.&lt;/p&gt;

&lt;p&gt;Your application needs to manage audio focus correctly and request the appropriate microphone permissions, including &lt;code&gt;RECORD_AUDIO&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Permission denial should also be treated as a normal application state.&lt;/p&gt;

&lt;p&gt;A robust SDK should surface a useful error that lets your interface explain what happened. It should not simply fail to start recording.&lt;/p&gt;

&lt;h3&gt;
  
  
  React Native and Flutter need special attention
&lt;/h3&gt;

&lt;p&gt;Cross-platform frameworks introduce a bridge between native audio code and JavaScript or Dart.&lt;/p&gt;

&lt;p&gt;Passing every small raw audio buffer across that bridge can become expensive.&lt;/p&gt;

&lt;p&gt;When possible, keep latency-sensitive audio processing inside the native layer. Pass higher-level information across the bridge, such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Transcript updates&lt;/li&gt;
&lt;li&gt;Playback commands&lt;/li&gt;
&lt;li&gt;Connection state&lt;/li&gt;
&lt;li&gt;Errors&lt;/li&gt;
&lt;li&gt;Turn events&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This keeps high-frequency audio work closer to the platform APIs that are designed to handle it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Voice cloning: separate voice identity from playback
&lt;/h2&gt;

&lt;p&gt;Voice cloning introduces another kind of state into the integration.&lt;/p&gt;

&lt;p&gt;At the application level, it is useful to think of the workflow as two stages.&lt;/p&gt;

&lt;p&gt;First, create or provision the voice resource. The service returns an identifier representing that voice.&lt;/p&gt;

&lt;p&gt;Second, reference that identifier during later synthesis requests.&lt;/p&gt;

&lt;p&gt;The exact API shape varies between providers. Some platforms perform voice creation through a single endpoint and return the resulting ID immediately or after processing.&lt;/p&gt;

&lt;p&gt;The application architecture is still similar:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Reference audio
      |
      v
Voice creation
      |
      v
Server-managed voice ID
      |
      v
TTS requests
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keeping the voice ID on the server instead of scattering it across browser or mobile state makes future changes easier.&lt;/p&gt;

&lt;p&gt;If a voice needs to be replaced, migrated, or updated, your backend can change the mapping without requiring every client to change.&lt;/p&gt;

&lt;p&gt;Smallest AI also exposes &lt;a href="https://smallest.ai/voice-cloning?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-voice-sdks-for-python-browser-and-mobile-apps"&gt;voice cloning&lt;/a&gt; as part of its speech stack.&lt;/p&gt;

&lt;p&gt;Reference-audio requirements differ between voice-cloning systems, so verify the current provider requirements instead of hard-coding assumptions about recording duration or format.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production scaling is mostly an orchestration problem
&lt;/h2&gt;

&lt;p&gt;A voice integration working for one developer on a laptop tells you very little about how it will behave under concurrent traffic.&lt;/p&gt;

&lt;p&gt;Production introduces constraints such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Concurrency limits&lt;/li&gt;
&lt;li&gt;Connection limits&lt;/li&gt;
&lt;li&gt;Cold starts&lt;/li&gt;
&lt;li&gt;Queue buildup&lt;/li&gt;
&lt;li&gt;Network failures&lt;/li&gt;
&lt;li&gt;Timeouts&lt;/li&gt;
&lt;li&gt;Retries&lt;/li&gt;
&lt;li&gt;Partial responses&lt;/li&gt;
&lt;li&gt;Usage-based cost&lt;/li&gt;
&lt;li&gt;Downstream backpressure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One common mistake is treating every speech operation as if it were an ordinary synchronous REST request.&lt;/p&gt;

&lt;p&gt;That model starts to break down once multiple users are producing and consuming audio continuously.&lt;/p&gt;

&lt;p&gt;For non-interactive workloads, queues and asynchronous workers can absorb traffic bursts.&lt;/p&gt;

&lt;p&gt;For real-time conversations, you usually cannot hide everything behind a long queue because the user is waiting. Instead, the system needs explicit limits, cancellation, backpressure, and a clear fallback path.&lt;/p&gt;

&lt;p&gt;Your SDK wrapper should also differentiate between failures.&lt;/p&gt;

&lt;p&gt;A timeout is different from invalid input. A transient network problem is different from an authentication failure. A rate limit is different from the speech service returning no usable result.&lt;/p&gt;

&lt;p&gt;Collapsing all of those into one generic exception makes production debugging unnecessarily difficult.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure time-to-first-audio, not just total generation time
&lt;/h2&gt;

&lt;p&gt;For text-to-speech, total synthesis duration is not the only latency metric that matters.&lt;/p&gt;

&lt;p&gt;In a conversational application, the user cares about when the response begins.&lt;/p&gt;

&lt;p&gt;Time-to-first-audio, or TTFA, measures the interval between sending a synthesis request and receiving the first playable audio.&lt;/p&gt;

&lt;p&gt;A system can generate a complete response quickly and still feel slow if it waits too long before playback begins.&lt;/p&gt;

&lt;p&gt;That is why streaming changes the perceived responsiveness of a voice interface. It allows useful output to begin before the entire operation has completed.&lt;/p&gt;

&lt;p&gt;The same idea applies throughout a voice pipeline.&lt;/p&gt;

&lt;p&gt;Do not optimize only the model call. Measure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;audio capture
    -&amp;gt; buffering
    -&amp;gt; transport
    -&amp;gt; speech recognition
    -&amp;gt; application logic
    -&amp;gt; synthesis
    -&amp;gt; playback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;End-to-end latency is the number the user actually experiences.&lt;/p&gt;

&lt;h2&gt;
  
  
  Own the integration boundary before SDK fragmentation owns you
&lt;/h2&gt;

&lt;p&gt;SDK fragmentation becomes painful once an application spans multiple environments.&lt;/p&gt;

&lt;p&gt;You may find one provider with an excellent Python integration, another with the browser behavior you want, and another with a feature that only exists on mobile.&lt;/p&gt;

&lt;p&gt;Now your team owns:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Multiple authentication systems&lt;/li&gt;
&lt;li&gt;Different error formats&lt;/li&gt;
&lt;li&gt;Multiple billing surfaces&lt;/li&gt;
&lt;li&gt;Different streaming protocols&lt;/li&gt;
&lt;li&gt;Separate monitoring&lt;/li&gt;
&lt;li&gt;Different release schedules&lt;/li&gt;
&lt;li&gt;Different client behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The strongest defense is to make your own application boundary stable.&lt;/p&gt;

&lt;p&gt;Your browser and mobile applications should not need to understand how a particular speech provider authenticates requests or formats every response.&lt;/p&gt;

&lt;p&gt;Instead, define the small set of operations your product actually needs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;start transcription
stop transcription
synthesize speech
cancel playback
create or resolve voice
report stream state
handle failure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then hide provider-specific details behind that layer.&lt;/p&gt;

&lt;p&gt;This also makes it much easier to evaluate a platform with your own workloads rather than designing the entire product around whichever SDK you tried first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The SDK decision is an architecture decision
&lt;/h2&gt;

&lt;p&gt;A production voice SDK should do more than shorten a few API calls.&lt;/p&gt;

&lt;p&gt;In Python, it needs to fit an asynchronous system that can stream data, handle failures, and scale beyond one request at a time.&lt;/p&gt;

&lt;p&gt;In the browser, it needs to coexist with microphone permissions, Web Audio, autoplay policies, and a server-side authentication boundary.&lt;/p&gt;

&lt;p&gt;On mobile, it has to respect native audio sessions, focus changes, permissions, interruptions, and cross-platform bridge costs.&lt;/p&gt;

&lt;p&gt;Voice cloning adds persistent voice identity. Production traffic adds backpressure, concurrency, retries, and observability.&lt;/p&gt;

&lt;p&gt;The goal is not to find the SDK with the shortest quickstart. It is to find an integration model that remains understandable after the prototype becomes a real product.&lt;/p&gt;

&lt;p&gt;If you are building one of these pipelines, &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-voice-sdks-for-python-browser-and-mobile-apps"&gt;start building with the Smallest AI API&lt;/a&gt; and test the architecture with the same audio, devices, networks, and concurrency patterns your application will actually face.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>webdev</category>
      <category>mobile</category>
    </item>
    <item>
      <title>Build a Real-Time Voice AI Agent with LiveKit, Pulse STT, and Lightning TTS</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Tue, 15 Sep 2026 07:47:27 +0000</pubDate>
      <link>https://dev.to/smallestai/build-a-real-time-voice-ai-agent-with-livekit-pulse-stt-and-lightning-tts-4ad7</link>
      <guid>https://dev.to/smallestai/build-a-real-time-voice-ai-agent-with-livekit-pulse-stt-and-lightning-tts-4ad7</guid>
      <description>&lt;p&gt;A real-time voice agent rarely feels fast because of one model alone. It feels fast because speech recognition, reasoning, and speech synthesis start doing useful work before the previous stage has completely finished.&lt;/p&gt;

&lt;p&gt;That is the important mental model when building with LiveKit.&lt;/p&gt;

&lt;p&gt;LiveKit handles the real-time media and agent plumbing. Your STT and TTS choices still determine how quickly speech becomes usable text and how quickly generated text becomes audible speech. In this guide, we will connect LiveKit Agents with &lt;a href="https://smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-with-livekit-stt-and-tts-integration-guide"&gt;Smallest AI&lt;/a&gt;, using Pulse for speech-to-text and Lightning for text-to-speech.&lt;/p&gt;

&lt;p&gt;The focus is not just getting a demo running. We will also look at endpointing, interruptions, failure recovery, state management, and the latency boundaries you should measure before shipping.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline is concurrent, not sequential
&lt;/h2&gt;

&lt;p&gt;The basic architecture looks simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User audio -&amp;gt; STT -&amp;gt; LLM -&amp;gt; TTS -&amp;gt; User audio
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But implementing those stages as a strictly sequential pipeline creates unnecessary waiting.&lt;/p&gt;

&lt;p&gt;In a streaming voice system:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;STT begins producing transcript updates while the user audio is still being processed.&lt;/li&gt;
&lt;li&gt;The LLM can begin generating a response once enough text is available.&lt;/li&gt;
&lt;li&gt;TTS can begin synthesizing useful text chunks before the full answer is complete.&lt;/li&gt;
&lt;li&gt;Audio playback can begin while later text is still being generated and synthesized.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The stages overlap.&lt;/p&gt;

&lt;p&gt;That is why first useful output matters so much in voice systems. Waiting for a complete transcript, then a complete LLM response, then a complete audio file makes each stage's latency accumulate.&lt;/p&gt;

&lt;p&gt;LiveKit Agents is designed around this streaming model. Its plugin system also lets you replace individual STT, LLM, and TTS components without rewriting the complete media pipeline.&lt;/p&gt;

&lt;p&gt;If you want the architecture argument in more detail, the guide on &lt;a href="https://smallest.ai/blog/why-streaming-architecture-is-non-negotiable-for-real-time-voice-agents?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-with-livekit-stt-and-tts-integration-guide"&gt;why streaming architecture matters for real-time voice agents&lt;/a&gt; explains why partial results and overlapping work matter so much for conversational latency.&lt;/p&gt;

&lt;p&gt;If LiveKit Agents is new to you, the &lt;a href="https://docs.livekit.io/agents/" rel="noopener noreferrer"&gt;LiveKit Agents documentation&lt;/a&gt; is a useful starting point before continuing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Set up the LiveKit and Smallest AI stack
&lt;/h2&gt;

&lt;p&gt;For a current Python setup, you need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A LiveKit Cloud project or your own LiveKit deployment&lt;/li&gt;
&lt;li&gt;A Smallest AI account and API key&lt;/li&gt;
&lt;li&gt;Credentials for the LLM provider you are using&lt;/li&gt;
&lt;li&gt;Python 3.11 for the setup shown here&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Create a virtual environment and install the required packages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3.11 &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv
&lt;span class="nb"&gt;source&lt;/span&gt; .venv/bin/activate

pip &lt;span class="nb"&gt;install &lt;/span&gt;livekit-plugins-smallestai livekit-plugins-openai livekit-plugins-silero python-dotenv
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Windows, activate the environment with the equivalent &lt;code&gt;.venv\Scripts\activate&lt;/code&gt; command.&lt;/p&gt;

&lt;p&gt;The important dependency here is &lt;code&gt;livekit-plugins-smallestai&lt;/code&gt;. It exposes both the Pulse STT and Lightning TTS integrations, so you do not need separate provider packages for each side of the speech pipeline.&lt;/p&gt;

&lt;p&gt;For parameters, supported models, and integration-specific configuration, keep &lt;a href="https://smallest.ai/voice-ai-apps/livekit?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-with-livekit-stt-and-tts-integration-guide"&gt;Smallest AI's LiveKit integration&lt;/a&gt; alongside the LiveKit documentation while implementing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create and store the API key
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-with-livekit-stt-and-tts-integration-guide"&gt;Smallest AI API&lt;/a&gt; uses authenticated access to the speech services used by this integration.&lt;/p&gt;

&lt;p&gt;Keep the API key in an environment variable rather than hard-coding it into the application.&lt;/p&gt;

&lt;p&gt;Before running the snippet, &lt;a href="https://docs.smallest.ai/voice-agents/platform/account/api-keys?utm_source=dev.to&amp;amp;utm_medium=Vizup&amp;amp;utm_campaign=real-time-voice-ai-with-livekit-stt-and-tts-integration-guide"&gt;create a Smallest.ai API key in the dashboard&lt;/a&gt; and store it in the &lt;code&gt;SMALLEST_API_KEY&lt;/code&gt; environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key-here"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every authenticated request sends the value through the Authorization header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer &amp;lt;SMALLEST_API_KEY value&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the key on your server. Do not expose it in browser JavaScript, mobile application code, public repositories, screenshots, query parameters, client-side logs, or error messages returned to users.&lt;/p&gt;

&lt;p&gt;For production deployments, store it in your infrastructure's server-side secrets manager rather than committing it to source control.&lt;/p&gt;

&lt;p&gt;You will also need your LiveKit credentials and whichever LLM-provider credentials your application uses. For the example below, that means configuring:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LIVEKIT_URL
LIVEKIT_API_KEY
LIVEKIT_API_SECRET
OPENAI_API_KEY
SMALLEST_API_KEY
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Wire Pulse STT and Lightning TTS into LiveKit
&lt;/h2&gt;

&lt;p&gt;Current LiveKit Agents integrations use &lt;code&gt;AgentSession&lt;/code&gt;, with each model supplied as a component of the session.&lt;/p&gt;

&lt;p&gt;Before running the snippet, &lt;a href="https://docs.smallest.ai/voice-agents/platform/account/api-keys?utm_source=dev.to&amp;amp;utm_medium=Vizup&amp;amp;utm_campaign=real-time-voice-ai-with-livekit-stt-and-tts-integration-guide"&gt;create a Smallest.ai API key in the dashboard&lt;/a&gt; and store it in the &lt;code&gt;SMALLEST_API_KEY&lt;/code&gt; environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dotenv&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;load_dotenv&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;livekit.agents&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;AgentSession&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;JobContext&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;JobProcess&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;WorkerOptions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;cli&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;livekit.plugins&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;silero&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;smallestai&lt;/span&gt;


&lt;span class="nf"&gt;load_dotenv&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;require_env&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;RuntimeError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; is not set&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;VoiceAgent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;super&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;instructions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are a concise, helpful voice assistant.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;prewarm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;proc&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;JobProcess&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;proc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;userdata&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;vad&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;silero&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;VAD&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;


&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;entrypoint&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;JobContext&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;smallest_api_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require_env&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="nf"&gt;require_env&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LIVEKIT_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;require_env&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LIVEKIT_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;require_env&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LIVEKIT_API_SECRET&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;require_env&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;session&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;AgentSession&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;vad&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;proc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;userdata&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;vad&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;stt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;smallestai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;STT&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;smallest_api_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pulse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;en&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;LLM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-4o-mini&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;tts&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;smallestai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;TTS&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;smallest_api_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lightning_v3.1_pro&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;voice_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;meher&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;en&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;start&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;VoiceAgent&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="n"&gt;room&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;room&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;cli&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run_app&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="nc"&gt;WorkerOptions&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;entrypoint_fnc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;entrypoint&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;prewarm_fnc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;prewarm&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the agent in development mode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 agent.py dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key architectural point is inside &lt;code&gt;AgentSession&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;smallestai.STT(...)&lt;/code&gt; handles the incoming speech layer. The LLM generates the response. &lt;code&gt;smallestai.TTS(...)&lt;/code&gt; converts that response back into streaming speech.&lt;/p&gt;

&lt;p&gt;LiveKit coordinates those stages as part of one session rather than requiring you to manually shuttle audio and text between independent services.&lt;/p&gt;

&lt;p&gt;Model names, voice IDs, language options, and plugin behavior can change as releases evolve, so verify those values against the current documentation before pinning a production configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Pulse contributes to the pipeline
&lt;/h2&gt;

&lt;p&gt;Pulse is the speech-to-text side of the integration.&lt;/p&gt;

&lt;p&gt;For a real-time agent, the STT layer is not simply a function that returns a transcript. It participates in turn timing.&lt;/p&gt;

&lt;p&gt;The application needs to answer questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;When has the user actually finished speaking?&lt;/li&gt;
&lt;li&gt;Do partial transcripts arrive early enough to help downstream processing?&lt;/li&gt;
&lt;li&gt;How does background noise affect end-of-turn detection?&lt;/li&gt;
&lt;li&gt;What happens when a user pauses in the middle of a sentence?&lt;/li&gt;
&lt;li&gt;Are language settings correct for the traffic you actually receive?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;LiveKit's STT integration can consume Pulse as a streaming provider, while the rest of the agent continues through the normal LiveKit session abstraction.&lt;/p&gt;

&lt;p&gt;That separation is useful because you can evaluate speech recognition independently without redesigning your LLM and TTS layers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Finls97bp0jqa9n9ua3ka.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Finls97bp0jqa9n9ua3ka.png" alt=" " width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the TTS stage affects perceived speed so much
&lt;/h2&gt;

&lt;p&gt;A voice agent can have accurate transcription and a fast LLM and still feel slow.&lt;/p&gt;

&lt;p&gt;The user does not hear either of those components directly. They hear the first audio produced by the TTS layer.&lt;/p&gt;

&lt;p&gt;If the application waits for the entire LLM response before starting synthesis, you lose much of the benefit of streaming upstream.&lt;/p&gt;

&lt;p&gt;Lightning fits the streaming TTS side of the LiveKit pipeline by turning generated text into audio as the interaction progresses rather than requiring a completed long-form response first.&lt;/p&gt;

&lt;p&gt;This means your latency budget should not only measure total speech-generation time. Measure when the first useful audio reaches the listener.&lt;/p&gt;

&lt;p&gt;For voice agents, that first audible response strongly affects whether the interaction feels continuous or delayed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Think about each stage separately
&lt;/h2&gt;

&lt;p&gt;A useful debugging model is to break the pipeline into boundaries:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;What to watch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;STT&lt;/td&gt;
&lt;td&gt;User audio&lt;/td&gt;
&lt;td&gt;Partial and final text&lt;/td&gt;
&lt;td&gt;Turn timing, recognition delay, transcript stability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM&lt;/td&gt;
&lt;td&gt;Conversation text&lt;/td&gt;
&lt;td&gt;Generated response text&lt;/td&gt;
&lt;td&gt;Time to first useful text, tool latency, context size&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TTS&lt;/td&gt;
&lt;td&gt;Generated text&lt;/td&gt;
&lt;td&gt;Audio chunks&lt;/td&gt;
&lt;td&gt;Time to first audio, buffering, cancellation behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Playback&lt;/td&gt;
&lt;td&gt;Audio chunks&lt;/td&gt;
&lt;td&gt;What the user hears&lt;/td&gt;
&lt;td&gt;Queueing, interruption handling, transport delay&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This makes performance problems easier to diagnose.&lt;/p&gt;

&lt;p&gt;For example, a slow conversation does not automatically mean the TTS model is slow. The delay could come from endpointing, an LLM tool call, audio buffering, or the network path between your worker and a provider.&lt;/p&gt;

&lt;p&gt;Measure boundaries independently before optimizing the wrong component.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production concerns most demos do not expose
&lt;/h2&gt;

&lt;p&gt;A successful demo proves that the pipeline is connected.&lt;/p&gt;

&lt;p&gt;Production asks harder questions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Endpointing and turn detection
&lt;/h3&gt;

&lt;p&gt;The system must decide when the user's turn is complete.&lt;/p&gt;

&lt;p&gt;Real users hesitate, restart sentences, speak over background noise, and pause before important words. An endpointing configuration that feels responsive for short test phrases can become frustrating with real conversation.&lt;/p&gt;

&lt;p&gt;If your turn detector fires too early, the agent may cut the user off.&lt;/p&gt;

&lt;p&gt;If it waits too long, every response begins with an awkward silence.&lt;/p&gt;

&lt;p&gt;Recent LiveKit APIs are consolidating turn behavior under newer turn-handling configuration, so check the version-specific documentation instead of copying older &lt;code&gt;min_endpointing_delay&lt;/code&gt; or &lt;code&gt;max_endpointing_delay&lt;/code&gt; examples blindly.&lt;/p&gt;

&lt;p&gt;There is another subtle issue: provider-side end-of-utterance logic and LiveKit turn detection can both contribute delay. Test the combined pipeline rather than tuning each component in isolation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Barge-in and interruption handling
&lt;/h3&gt;

&lt;p&gt;People interrupt each other naturally.&lt;/p&gt;

&lt;p&gt;Your agent should expect the same behavior.&lt;/p&gt;

&lt;p&gt;When the user starts speaking while the agent is talking, the system needs to stop or suppress the current playback, return attention to incoming speech, and continue the new turn without leaving stale audio queued behind.&lt;/p&gt;

&lt;p&gt;LiveKit supports interruptible voice-agent flows, but you should still test the exact combinations your application will encounter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Very short interruptions&lt;/li&gt;
&lt;li&gt;Users correcting themselves&lt;/li&gt;
&lt;li&gt;Interruptions during long TTS responses&lt;/li&gt;
&lt;li&gt;Background speech that should not cancel playback&lt;/li&gt;
&lt;li&gt;Repeated barge-in during a single session&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not evaluate interruption handling with one happy-path sentence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Network failures and provider errors
&lt;/h3&gt;

&lt;p&gt;STT and TTS calls depend on network connections.&lt;/p&gt;

&lt;p&gt;Transient errors will happen.&lt;/p&gt;

&lt;p&gt;Your application should define what happens when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The STT stream disconnects&lt;/li&gt;
&lt;li&gt;TTS synthesis fails after text has already been generated&lt;/li&gt;
&lt;li&gt;A request times out&lt;/li&gt;
&lt;li&gt;The user disconnects during synthesis&lt;/li&gt;
&lt;li&gt;A retry succeeds after the conversation has already moved on&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use bounded retries and exponential backoff where appropriate.&lt;/p&gt;

&lt;p&gt;More importantly, make retries conversation-aware. Replaying a stale TTS response after the user has already started another turn is worse than dropping the response.&lt;/p&gt;

&lt;h3&gt;
  
  
  First-request connection overhead
&lt;/h3&gt;

&lt;p&gt;Persistent streaming connections often behave differently on the first request than after a worker has warmed up.&lt;/p&gt;

&lt;p&gt;If your deployment creates workers on demand, test cold-start behavior separately from steady-state performance.&lt;/p&gt;

&lt;p&gt;For production voice AI, averages can hide a frustrating first interaction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi-turn context is a separate design problem
&lt;/h2&gt;

&lt;p&gt;Once latency is acceptable, long conversations introduce another issue: state.&lt;/p&gt;

&lt;p&gt;A single-turn assistant can pass one transcript to the LLM and return one answer.&lt;/p&gt;

&lt;p&gt;A multi-turn agent needs to decide how much previous conversation should remain in context.&lt;/p&gt;

&lt;p&gt;Two common patterns are useful.&lt;/p&gt;

&lt;h3&gt;
  
  
  Sliding conversation window
&lt;/h3&gt;

&lt;p&gt;Keep only the most recent N turns.&lt;/p&gt;

&lt;p&gt;This is straightforward and works well when recent context matters more than information from much earlier in the session.&lt;/p&gt;

&lt;p&gt;The tradeoff is obvious: once older turns fall out of the window, the model can no longer use them directly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Periodic summarization
&lt;/h3&gt;

&lt;p&gt;Instead of retaining every raw turn, periodically summarize older conversation state and keep that summary alongside the most recent exchanges.&lt;/p&gt;

&lt;p&gt;This reduces context growth while preserving important information from earlier in the session.&lt;/p&gt;

&lt;p&gt;The difficult part is deciding what the summary is allowed to discard.&lt;/p&gt;

&lt;p&gt;For support, workflow, or transactional agents, preserving an account choice, requested action, or unresolved issue may matter much more than preserving exact wording.&lt;/p&gt;

&lt;p&gt;Treat context management as application logic, not simply an LLM setting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure the pipeline you actually plan to ship
&lt;/h2&gt;

&lt;p&gt;A provider benchmark can tell you something useful about one model.&lt;/p&gt;

&lt;p&gt;It cannot tell you the latency of your full application.&lt;/p&gt;

&lt;p&gt;Measure at least these boundaries in your own deployment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;End of user speech to usable transcript&lt;/li&gt;
&lt;li&gt;Usable transcript to first generated LLM text&lt;/li&gt;
&lt;li&gt;First generated text to first TTS audio&lt;/li&gt;
&lt;li&gt;End of user speech to first audible response&lt;/li&gt;
&lt;li&gt;Time required to cancel playback after interruption&lt;/li&gt;
&lt;li&gt;Error and reconnect rates during longer sessions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Run these tests with realistic network paths, audio devices, languages, sentence lengths, and concurrency.&lt;/p&gt;

&lt;p&gt;A local development machine on fast Wi-Fi is not the production environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Smallest AI officially supported by LiveKit?
&lt;/h3&gt;

&lt;p&gt;Smallest AI appears in LiveKit's STT and TTS provider documentation, and the current integration uses the &lt;code&gt;smallestai&lt;/code&gt; plugin package with &lt;code&gt;AgentSession&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That is different from maintaining your own custom STT or TTS adapter outside LiveKit's plugin ecosystem.&lt;/p&gt;

&lt;h3&gt;
  
  
  What end-to-end latency should I expect?
&lt;/h3&gt;

&lt;p&gt;There is no single number that applies to every deployment.&lt;/p&gt;

&lt;p&gt;End-to-end latency includes turn detection, speech recognition, network transport, LLM generation, tool calls, speech synthesis, buffering, and playback.&lt;/p&gt;

&lt;p&gt;Provider-level latency figures are useful for component evaluation, but measure the complete pipeline using your own deployment architecture.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I use a cloned voice?
&lt;/h3&gt;

&lt;p&gt;Lightning supports voice-cloning workflows, but model and voice compatibility matters.&lt;/p&gt;

&lt;p&gt;Verify that the Lightning model you choose supports the cloned voice you intend to use, then provide the corresponding &lt;code&gt;voice_id&lt;/code&gt; to the TTS integration.&lt;/p&gt;

&lt;p&gt;Do not assume that every voice is available across every Lightning model or language configuration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is LiveKit Agents free to use?
&lt;/h3&gt;

&lt;p&gt;LiveKit Agents is open source.&lt;/p&gt;

&lt;p&gt;Running an agent can still create infrastructure, bandwidth, model-provider, and other service costs depending on your deployment.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which languages does Pulse support?
&lt;/h3&gt;

&lt;p&gt;Pulse supports multilingual speech recognition, but supported languages and regional options can differ between real-time and pre-recorded transcription.&lt;/p&gt;

&lt;p&gt;Check the current language list before deploying rather than hard-coding assumptions from an older model release.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next steps
&lt;/h2&gt;

&lt;p&gt;The most important lesson is not that one STT or TTS model makes an agent real-time.&lt;/p&gt;

&lt;p&gt;The system feels conversational when the entire pipeline is designed to stream, overlap work, cancel obsolete work, and recover cleanly when real users behave unpredictably.&lt;/p&gt;

&lt;p&gt;Start with the simple STT-LLM-TTS pipeline. Then measure each boundary independently. Tune endpointing and interruptions with real conversations. Finally, decide how your application will handle context, retries, cold starts, and long-running sessions.&lt;/p&gt;

&lt;p&gt;When you are ready to test the speech layer with your own LiveKit application, &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-with-livekit-stt-and-tts-integration-guide"&gt;start building with the Smallest AI API&lt;/a&gt; and measure the complete pipeline under the conditions your users will actually encounter.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>livekit</category>
      <category>viceai</category>
      <category>python</category>
    </item>
    <item>
      <title>Streaming TTS for Developers: Latency, Buffering, and Real-Time Voice Architecture</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Thu, 10 Sep 2026 10:01:52 +0000</pubDate>
      <link>https://dev.to/smallestai/streaming-tts-for-developers-latency-buffering-and-real-time-voice-architecture-1po0</link>
      <guid>https://dev.to/smallestai/streaming-tts-for-developers-latency-buffering-and-real-time-voice-architecture-1po0</guid>
      <description>&lt;p&gt;Most developers first encounter text-to-speech as a simple request-response operation:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Send text.&lt;/li&gt;
&lt;li&gt;Wait for synthesis.&lt;/li&gt;
&lt;li&gt;Receive an audio file.&lt;/li&gt;
&lt;li&gt;Play it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That model is perfectly reasonable for short prompts, narration, notifications, or anything that can be generated ahead of time.&lt;/p&gt;

&lt;p&gt;It becomes much more noticeable when you build conversational AI, real-time IVR, voice assistants, or other systems where the output is created dynamically.&lt;/p&gt;

&lt;p&gt;In those applications, the user is not waiting for "audio generation." They are experiencing silence.&lt;/p&gt;

&lt;p&gt;That distinction is why streaming TTS matters.&lt;/p&gt;

&lt;p&gt;Instead of waiting for the entire response to be synthesized, a streaming pipeline starts delivering audio while synthesis is still happening. Playback can begin before the final part of the response exists.&lt;/p&gt;

&lt;p&gt;For developers, the interesting part is not simply that streaming is faster. It changes where latency appears, how you buffer audio, how you connect an LLM to the speech layer, how you handle concurrency, and when caching still beats streaming.&lt;/p&gt;

&lt;p&gt;If you are evaluating real-time voice infrastructure, &lt;a href="https://smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=streaming-tts-explained-for-developers-ux-latency-and-cost-guide"&gt;Smallest AI&lt;/a&gt; is one example of a platform exposing TTS for these kinds of workloads.&lt;/p&gt;

&lt;p&gt;The global text-to-speech market was valued at USD 4.8 billion in 2025 and projected to reach USD 5.7 billion in 2026, according to &lt;a href="https://www.gminsights.com/industry-analysis/text-to-speech-market" rel="noopener noreferrer"&gt;Global Market Insights&lt;/a&gt;. More applications are adding generated speech, but real-time systems have very different architectural requirements from offline voice generation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What streaming text-to-speech actually means
&lt;/h2&gt;

&lt;p&gt;Traditional batch TTS works as a complete request-response cycle.&lt;/p&gt;

&lt;p&gt;You send the full text to the synthesis service. The service generates the complete audio output. Only then does the application receive something it can play.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Text
  |
  v
TTS synthesis
  |
  v
Complete audio
  |
  v
Playback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Streaming changes the delivery model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Text
  |
  v
TTS synthesis
  |
  +--&amp;gt; Audio chunk 1 --&amp;gt; Playback starts
  |
  +--&amp;gt; Audio chunk 2
  |
  +--&amp;gt; Audio chunk 3
  |
  +--&amp;gt; ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The client no longer waits for synthesis to finish before doing useful work.&lt;/p&gt;

&lt;p&gt;Depending on the API, streaming can be delivered using mechanisms such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Server-Sent Events&lt;/li&gt;
&lt;li&gt;WebSockets&lt;/li&gt;
&lt;li&gt;Streaming HTTP responses&lt;/li&gt;
&lt;li&gt;Provider-specific real-time protocols&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There are also two different problems that are often grouped under "streaming TTS."&lt;/p&gt;

&lt;p&gt;The first is streaming audio output. You already have the complete text, but the TTS service begins returning generated audio before synthesis finishes.&lt;/p&gt;

&lt;p&gt;The second is streaming text input and audio output. The input itself is still being generated, often by an LLM, and the TTS pipeline begins synthesizing before the complete LLM response exists.&lt;/p&gt;

&lt;p&gt;The second case is especially important for conversational AI.&lt;/p&gt;

&lt;h2&gt;
  
  
  Streaming TTS is not the same as low-latency TTS
&lt;/h2&gt;

&lt;p&gt;This distinction is easy to miss.&lt;/p&gt;

&lt;p&gt;A low-latency batch TTS service might synthesize a short response quickly and return the complete audio file.&lt;/p&gt;

&lt;p&gt;A streaming service might take longer to finish the entire synthesis but begin returning playable audio much earlier.&lt;/p&gt;

&lt;p&gt;For an interactive application, those are different performance characteristics.&lt;/p&gt;

&lt;p&gt;You should usually measure at least:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Time to first audio, how long the user waits before hearing anything&lt;/li&gt;
&lt;li&gt;Total synthesis latency, how long the complete generation takes&lt;/li&gt;
&lt;li&gt;Real-time factor, whether audio is generated faster than it is consumed&lt;/li&gt;
&lt;li&gt;Playback underruns, how often playback catches up with generation&lt;/li&gt;
&lt;li&gt;Tail latency, what happens to slower requests rather than only the average&lt;/li&gt;
&lt;li&gt;Latency under concurrency, whether performance changes when many sessions run at once&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For conversational systems, time to first audio often matters more to perceived responsiveness than the time required to generate the final byte.&lt;/p&gt;

&lt;p&gt;That is one reason &lt;a href="https://smallest.ai/text-to-speech?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=streaming-tts-explained-for-developers-ux-latency-and-cost-guide"&gt;Smallest AI's text-to-speech product&lt;/a&gt; focuses on real-time speech generation alongside conventional TTS workloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why perceived latency is the real UX problem
&lt;/h2&gt;

&lt;p&gt;Human conversations leave very little dead space between turns.&lt;/p&gt;

&lt;p&gt;A few hundred milliseconds of pause can feel completely natural. Once additional processing layers start stacking up, however, the experience becomes noticeably less conversational.&lt;/p&gt;

&lt;p&gt;A typical voice pipeline may already contain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User finishes speaking
        |
        v
Speech recognition finalizes
        |
        v
Application / LLM generates response
        |
        v
TTS begins synthesis
        |
        v
Network delivery
        |
        v
Client playback begins
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;TTS is only one component of the latency budget.&lt;/p&gt;

&lt;p&gt;If every stage waits for the previous stage to finish completely, delays accumulate.&lt;/p&gt;

&lt;p&gt;Streaming lets you overlap work.&lt;/p&gt;

&lt;p&gt;For example, an LLM can continue producing text while TTS synthesizes an earlier sentence. The client can play that sentence while the next audio chunk is still being generated.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LLM completes
      |
      v
TTS completes
      |
      v
Playback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you get something closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LLM tokens -------&amp;gt;

       sentence 1
            |
            v
       TTS chunk 1 --------&amp;gt; playback

              sentence 2
                   |
                   v
              TTS chunk 2 --------&amp;gt;

                     sentence 3
                          |
                          v
                     TTS chunk 3 --------&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That overlap is where much of the perceived latency improvement comes from.&lt;/p&gt;

&lt;h2&gt;
  
  
  Audio format matters, but streaming is not a codec feature
&lt;/h2&gt;

&lt;p&gt;Streaming is sometimes explained as if only particular audio formats can be streamed.&lt;/p&gt;

&lt;p&gt;The actual situation is more nuanced.&lt;/p&gt;

&lt;p&gt;Raw PCM is attractive for latency-sensitive pipelines because it has very little decoding or container overhead. Opus is also commonly used for real-time communication because it provides efficient compression and is designed for interactive audio.&lt;/p&gt;

&lt;p&gt;MP3 and AAC can also be delivered progressively in appropriate streaming configurations. The tradeoff is that container framing, decoding support, buffering behavior, and browser compatibility can make them less convenient for some ultra-low-latency pipelines.&lt;/p&gt;

&lt;p&gt;The correct question is not simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Does this format support streaming?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead, ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How quickly can the decoder consume the first received bytes?&lt;/li&gt;
&lt;li&gt;Does each chunk contain enough framing information?&lt;/li&gt;
&lt;li&gt;Does your playback environment support the codec?&lt;/li&gt;
&lt;li&gt;How much CPU does decoding require?&lt;/li&gt;
&lt;li&gt;How much bandwidth does the format save?&lt;/li&gt;
&lt;li&gt;Are you targeting browsers, telephony, mobile devices, or native applications?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a voice agent running over telephony, an 8 kHz telephony format may be more useful than high-fidelity audio.&lt;/p&gt;

&lt;p&gt;For a browser application, codec and playback API support may become the limiting factor.&lt;/p&gt;

&lt;p&gt;Architecture should follow the final listening environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cost equation is not just API pricing
&lt;/h2&gt;

&lt;p&gt;Streaming and batch synthesis are often priced similarly at the model layer.&lt;/p&gt;

&lt;p&gt;If an API charges according to characters or generated usage, streaming the response does not necessarily make synthesis itself cheaper.&lt;/p&gt;

&lt;p&gt;The infrastructure differences appear elsewhere.&lt;/p&gt;

&lt;h3&gt;
  
  
  Buffering
&lt;/h3&gt;

&lt;p&gt;With a batch implementation, your service may receive a complete audio payload before forwarding it.&lt;/p&gt;

&lt;p&gt;That means larger temporary buffers, especially when outputs are long.&lt;/p&gt;

&lt;p&gt;A streaming implementation can forward smaller pieces as they arrive.&lt;/p&gt;

&lt;p&gt;At high concurrency, reducing the amount of audio held in application memory can matter.&lt;/p&gt;

&lt;h3&gt;
  
  
  Connection management
&lt;/h3&gt;

&lt;p&gt;Streaming does not magically eliminate connection costs.&lt;/p&gt;

&lt;p&gt;In fact, streaming systems often keep HTTP or WebSocket connections active while audio is being generated or consumed.&lt;/p&gt;

&lt;p&gt;That shifts part of your capacity planning toward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Concurrent connections&lt;/li&gt;
&lt;li&gt;Socket limits&lt;/li&gt;
&lt;li&gt;Connection lifetime&lt;/li&gt;
&lt;li&gt;Proxy behavior&lt;/li&gt;
&lt;li&gt;Load balancer timeouts&lt;/li&gt;
&lt;li&gt;Backpressure&lt;/li&gt;
&lt;li&gt;Retry behavior&lt;/li&gt;
&lt;li&gt;Regional network latency&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Caching
&lt;/h3&gt;

&lt;p&gt;This is where batch synthesis can win decisively.&lt;/p&gt;

&lt;p&gt;Suppose your application repeatedly says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Your payment was successful.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is little reason to synthesize that sentence for every request.&lt;/p&gt;

&lt;p&gt;Generate it once, cache it, and serve the existing audio.&lt;/p&gt;

&lt;p&gt;Streaming is most valuable when the output is dynamic enough that caching provides little benefit.&lt;/p&gt;

&lt;p&gt;If you want a broader benchmark-oriented view of latency and production TTS evaluation, Smallest AI's &lt;a href="https://smallest.ai/blog/top-fastest-text-to-speech-apis-in-2026?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=streaming-tts-explained-for-developers-ux-latency-and-cost-guide"&gt;guide to fast text-to-speech APIs&lt;/a&gt; covers the metrics developers should compare.&lt;/p&gt;

&lt;h2&gt;
  
  
  When streaming TTS is the right architecture
&lt;/h2&gt;

&lt;p&gt;Streaming is a strong fit when the text does not exist until runtime and the user is waiting for the response.&lt;/p&gt;

&lt;p&gt;Typical examples include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Conversational AI agents&lt;/li&gt;
&lt;li&gt;LLM-powered voice assistants&lt;/li&gt;
&lt;li&gt;Open-ended IVR systems&lt;/li&gt;
&lt;li&gt;Real-time customer support automation&lt;/li&gt;
&lt;li&gt;Interactive accessibility applications&lt;/li&gt;
&lt;li&gt;Dynamic navigation prompts&lt;/li&gt;
&lt;li&gt;Live narration&lt;/li&gt;
&lt;li&gt;Generated commentary&lt;/li&gt;
&lt;li&gt;Applications where speech changes based on user context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The common pattern is interactivity.&lt;/p&gt;

&lt;p&gt;The faster useful output begins, the more responsive the system feels.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F174kscn40cdgnfhbyg69.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F174kscn40cdgnfhbyg69.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  When batch TTS is still better
&lt;/h2&gt;

&lt;p&gt;Streaming is not a universal replacement for batch synthesis.&lt;/p&gt;

&lt;p&gt;Batch is often the simpler architecture when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Content is static&lt;/li&gt;
&lt;li&gt;Audio can be generated ahead of time&lt;/li&gt;
&lt;li&gt;The same output will be played repeatedly&lt;/li&gt;
&lt;li&gt;Total rendering time matters more than time to first audio&lt;/li&gt;
&lt;li&gt;You are creating podcasts or audiobooks offline&lt;/li&gt;
&lt;li&gt;Persistent streaming connections add unnecessary complexity&lt;/li&gt;
&lt;li&gt;Aggressive caching can eliminate repeated synthesis costs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, a set of 50 fixed IVR menu prompts should probably not be synthesized live for every caller.&lt;/p&gt;

&lt;p&gt;Dynamic answers inside that same IVR may benefit from streaming.&lt;/p&gt;

&lt;p&gt;Production systems frequently use both.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to connect an LLM to streaming TTS
&lt;/h2&gt;

&lt;p&gt;A common real-time pipeline looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User speech
    |
    v
Speech recognition
    |
    v
LLM token stream
    |
    v
Text buffering
    |
    v
Sentence / phrase boundary
    |
    v
Streaming TTS
    |
    v
Audio queue
    |
    v
Playback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important step is the buffer between the LLM and TTS.&lt;/p&gt;

&lt;p&gt;Sending every individual token directly to synthesis is usually not ideal.&lt;/p&gt;

&lt;p&gt;Imagine an LLM producing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The
weather
tomorrow
will
be
warmer
than
today.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Synthesizing each token independently would destroy natural phrasing.&lt;/p&gt;

&lt;p&gt;Instead, accumulate enough text to create a meaningful speech unit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The weather tomorrow will be warmer than today.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then send that unit into TTS while the LLM continues generating the next part of the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sentence boundaries are a latency-quality tradeoff
&lt;/h2&gt;

&lt;p&gt;Waiting for a complete paragraph gives the TTS model more linguistic context, but increases latency.&lt;/p&gt;

&lt;p&gt;Sending tiny fragments reduces waiting time, but can damage prosody.&lt;/p&gt;

&lt;p&gt;The ideal chunk size depends on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Model behavior&lt;/li&gt;
&lt;li&gt;Language&lt;/li&gt;
&lt;li&gt;Punctuation&lt;/li&gt;
&lt;li&gt;Response length&lt;/li&gt;
&lt;li&gt;Network latency&lt;/li&gt;
&lt;li&gt;Conversation style&lt;/li&gt;
&lt;li&gt;Whether the TTS API maintains context across chunks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A simple implementation might flush the text buffer when it encounters:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A period&lt;/li&gt;
&lt;li&gt;A question mark&lt;/li&gt;
&lt;li&gt;An exclamation mark&lt;/li&gt;
&lt;li&gt;A sufficiently long comma-separated phrase&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But punctuation alone is not perfect.&lt;/p&gt;

&lt;p&gt;Consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Dr. Patel will arrive at 4:30 p.m. tomorrow.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Naive splitting can create terrible boundaries.&lt;/p&gt;

&lt;p&gt;Production systems often need sentence segmentation that understands abbreviations, numbers, domains, and the language being synthesized.&lt;/p&gt;

&lt;p&gt;The key is to measure both latency and listening quality rather than optimizing one in isolation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Streaming full text and streaming LLM output are different cases
&lt;/h2&gt;

&lt;p&gt;For developers using the Smallest AI speech layer, the current streaming interface distinguishes between these patterns.&lt;/p&gt;

&lt;p&gt;When you already have the complete text, an SSE-based request can return audio chunks while synthesis is happening.&lt;/p&gt;

&lt;p&gt;When the text itself arrives incrementally, such as an LLM token stream, a persistent WebSocket is the more natural architecture.&lt;/p&gt;

&lt;p&gt;That distinction matters because the application controls different forms of backpressure in each case.&lt;/p&gt;

&lt;p&gt;With complete text:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Full text -&amp;gt; TTS -&amp;gt; streamed audio
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With generated text:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LLM -&amp;gt; text chunks -&amp;gt; TTS connection -&amp;gt; audio chunks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the second design, your application must decide when a text fragment is ready to synthesize.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create and store the API key
&lt;/h2&gt;

&lt;p&gt;Keep the API key in an environment variable rather than hard-coding it into the application.&lt;/p&gt;

&lt;p&gt;Before running the snippet, &lt;a href="https://docs.smallest.ai/voice-agents/platform/account/api-keys?utm_source=dev.to&amp;amp;utm_medium=Vizup&amp;amp;utm_campaign=streaming-tts-explained-for-developers-ux-latency-and-cost-guide"&gt;create a Smallest.ai API key in the dashboard&lt;/a&gt; and store it in the SMALLEST_API_KEY environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key-here"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every authenticated request sends the value through the Authorization header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer &amp;lt;SMALLEST_API_KEY value&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the key on your server.&lt;/p&gt;

&lt;p&gt;Do not expose it in browser JavaScript, mobile application code, public repositories, screenshots, query parameters, client-side logs, or error messages returned to users.&lt;/p&gt;

&lt;p&gt;For production systems, store the secret using your hosting provider's secret-management system or another appropriate server-side secrets manager rather than checking it into configuration files.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing streamed TTS with cURL
&lt;/h2&gt;

&lt;p&gt;The following example sends complete text and receives streamed audio events from the TTS endpoint.&lt;/p&gt;

&lt;p&gt;Before running the snippet, &lt;a href="https://docs.smallest.ai/voice-agents/platform/account/api-keys?utm_source=dev.to&amp;amp;utm_medium=Vizup&amp;amp;utm_campaign=streaming-tts-explained-for-developers-ux-latency-and-cost-guide"&gt;Create aSmallest.ai API key in the dashboard&lt;/a&gt; and store it in the SMALLEST_API_KEY environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;--fail-with-body&lt;/span&gt; &lt;span class="nt"&gt;--show-error&lt;/span&gt; &lt;span class="nt"&gt;-N&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"https://api.smallest.ai/waves/v1/tts/live"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$SMALLEST_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "text": "Streaming this paragraph chunk by chunk so playback can start sooner.",
    "voice_id": "magnus",
    "sample_rate": 24000,
    "output_format": "pcm"
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part from an architecture perspective is not the cURL command itself.&lt;/p&gt;

&lt;p&gt;It is that the response can be consumed incrementally rather than waiting for a completed audio file.&lt;/p&gt;

&lt;p&gt;For an actual application, your server would parse the stream, decode the audio payload, and forward the appropriate audio data to the client or telephony layer.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=streaming-tts-explained-for-developers-ux-latency-and-cost-guide"&gt;Smallest AI developer platform&lt;/a&gt; is the place to create credentials and begin testing the speech layer with your own workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  Client-side playback needs its own buffer
&lt;/h2&gt;

&lt;p&gt;Receiving audio quickly does not guarantee smooth playback.&lt;/p&gt;

&lt;p&gt;Networks are inconsistent.&lt;/p&gt;

&lt;p&gt;Imagine the server generates chunks at these times:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;chunk 1   100 ms
chunk 2   190 ms
chunk 3   270 ms
chunk 4   610 ms
chunk 5   690 ms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the client plays everything the moment it arrives, the delay before chunk 4 could create an audible gap.&lt;/p&gt;

&lt;p&gt;That is why real-time playback usually maintains a small jitter or playback buffer.&lt;/p&gt;

&lt;p&gt;The buffer gives the system enough headroom to absorb short network stalls.&lt;/p&gt;

&lt;p&gt;Too small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;lower latency
higher underrun risk
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Too large:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;higher stability
higher perceived latency
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no universal correct size.&lt;/p&gt;

&lt;p&gt;Measure it under the network conditions your users actually experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Browser playback considerations
&lt;/h2&gt;

&lt;p&gt;The browser gives you several possible audio paths.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://developer.mozilla.org/en-US/docs/Web/API/Web_Audio_API" rel="noopener noreferrer"&gt;Web Audio API&lt;/a&gt; gives applications detailed control over audio processing and scheduling.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://developer.mozilla.org/en-US/docs/Web/API/Media_Source_Extensions_API" rel="noopener noreferrer"&gt;Media Source API&lt;/a&gt; is useful when your chosen media format and browser support fit the MediaSource model.&lt;/p&gt;

&lt;p&gt;For raw PCM, many applications maintain their own queue and schedule decoded audio buffers through Web Audio.&lt;/p&gt;

&lt;p&gt;Your client logic needs to handle:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Chunk ordering&lt;/li&gt;
&lt;li&gt;Decoding&lt;/li&gt;
&lt;li&gt;Playback scheduling&lt;/li&gt;
&lt;li&gt;Buffer underruns&lt;/li&gt;
&lt;li&gt;Connection loss&lt;/li&gt;
&lt;li&gt;Stream completion&lt;/li&gt;
&lt;li&gt;User interruption&lt;/li&gt;
&lt;li&gt;Reconnection&lt;/li&gt;
&lt;li&gt;Cancellation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For browser applications, authenticated TTS requests should still be made by your server. Do not ship your provider API key to the browser simply because the playback component lives there.&lt;/p&gt;

&lt;p&gt;A safer flow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Browser
   |
   | authenticated application request
   v
Your server
   |
   | provider API credential
   v
TTS API
   |
   | streamed audio
   v
Your server
   |
   v
Browser playback queue
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Backpressure matters once everything is streaming
&lt;/h2&gt;

&lt;p&gt;Streaming makes pipelines faster partly because stages can operate concurrently.&lt;/p&gt;

&lt;p&gt;It also introduces a new problem.&lt;/p&gt;

&lt;p&gt;What happens when one stage is faster than the next?&lt;/p&gt;

&lt;p&gt;Suppose the LLM generates text much faster than TTS can synthesize it.&lt;/p&gt;

&lt;p&gt;Your queue grows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LLM
 |
 v
[text][text][text][text][text][text]
                         |
                         v
                        TTS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now you have increased memory usage and potentially created several seconds of speech that the user has not heard yet.&lt;/p&gt;

&lt;p&gt;That makes interruption difficult.&lt;/p&gt;

&lt;p&gt;If the user changes direction while five sentences are already queued for speech, the system may need to throw that work away.&lt;/p&gt;

&lt;p&gt;A production streaming architecture should define:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Maximum queued text&lt;/li&gt;
&lt;li&gt;Maximum queued audio&lt;/li&gt;
&lt;li&gt;Cancellation semantics&lt;/li&gt;
&lt;li&gt;Barge-in behavior&lt;/li&gt;
&lt;li&gt;Timeouts&lt;/li&gt;
&lt;li&gt;Retry rules&lt;/li&gt;
&lt;li&gt;Whether generated but unplayed audio should be discarded&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Streaming is not only about moving bytes sooner. It is also about controlling work that is happening concurrently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Barge-in changes the design again
&lt;/h2&gt;

&lt;p&gt;Conversational voice applications often allow the user to interrupt the assistant.&lt;/p&gt;

&lt;p&gt;That means your system may have to cancel:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Audio currently playing&lt;/li&gt;
&lt;li&gt;Audio already generated but not played&lt;/li&gt;
&lt;li&gt;TTS requests still synthesizing&lt;/li&gt;
&lt;li&gt;LLM tokens still being generated&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Without coordinated cancellation, the application may stop playback locally while servers continue generating content nobody will use.&lt;/p&gt;

&lt;p&gt;The cancellation path should be treated as part of the normal architecture, not an edge case.&lt;/p&gt;

&lt;p&gt;A useful mental model is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User interruption
      |
      +--&amp;gt; stop playback
      |
      +--&amp;gt; clear audio queue
      |
      +--&amp;gt; cancel TTS generation
      |
      +--&amp;gt; cancel or redirect LLM generation
      |
      +--&amp;gt; begin listening again
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is one reason WebSocket-oriented pipelines are common in highly interactive voice applications.&lt;/p&gt;

&lt;h2&gt;
  
  
  SSML becomes more complicated when text is chunked
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://www.w3.org/TR/speech-synthesis11/" rel="noopener noreferrer"&gt;Speech Synthesis Markup Language specification&lt;/a&gt; provides controls for things such as pronunciation, pauses, rate, pitch, and emphasis.&lt;/p&gt;

&lt;p&gt;In batch synthesis, the TTS system can inspect the complete SSML document before generating speech.&lt;/p&gt;

&lt;p&gt;Streaming makes that harder.&lt;/p&gt;

&lt;p&gt;A chunk might contain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;prosody&lt;/span&gt; &lt;span class="na"&gt;rate=&lt;/span&gt;&lt;span class="s"&gt;"slow"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;while the closing tag arrives later.&lt;/p&gt;

&lt;p&gt;Whether this works depends on the implementation.&lt;/p&gt;

&lt;p&gt;Do not assume a provider's batch SSML behavior is identical to its streaming behavior.&lt;/p&gt;

&lt;p&gt;If SSML is important to your application, test:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tags crossing chunk boundaries&lt;/li&gt;
&lt;li&gt;Pronunciation dictionaries&lt;/li&gt;
&lt;li&gt;Nested elements&lt;/li&gt;
&lt;li&gt;Pauses&lt;/li&gt;
&lt;li&gt;Prosody controls&lt;/li&gt;
&lt;li&gt;Invalid partial markup&lt;/li&gt;
&lt;li&gt;Cancellation in the middle of marked-up text&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Voice consistency across chunks needs testing
&lt;/h2&gt;

&lt;p&gt;Neural speech generation depends heavily on context.&lt;/p&gt;

&lt;p&gt;A sentence synthesized in isolation may not sound exactly the same as that sentence synthesized as part of a paragraph.&lt;/p&gt;

&lt;p&gt;Aggressively splitting responses can therefore introduce:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pitch changes&lt;/li&gt;
&lt;li&gt;Pacing changes&lt;/li&gt;
&lt;li&gt;Unnatural pauses&lt;/li&gt;
&lt;li&gt;Repeated intonation patterns&lt;/li&gt;
&lt;li&gt;Different emotional delivery&lt;/li&gt;
&lt;li&gt;Audible boundaries between chunks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is another reason not to treat the smallest possible text chunk as the best possible text chunk.&lt;/p&gt;

&lt;p&gt;You are optimizing a multi-dimensional system:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;latency
quality
stability
cost
interruptibility
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The best production setting is usually a compromise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compliance and data handling still apply to streams
&lt;/h2&gt;

&lt;p&gt;Streaming does not reduce your responsibility for the data moving through the system.&lt;/p&gt;

&lt;p&gt;For applications handling sensitive information, document:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Where input text originates&lt;/li&gt;
&lt;li&gt;Which service receives it&lt;/li&gt;
&lt;li&gt;Whether requests are logged&lt;/li&gt;
&lt;li&gt;Whether audio is retained&lt;/li&gt;
&lt;li&gt;How traffic is encrypted&lt;/li&gt;
&lt;li&gt;How long temporary buffers live&lt;/li&gt;
&lt;li&gt;Which regions process data&lt;/li&gt;
&lt;li&gt;What happens during retries&lt;/li&gt;
&lt;li&gt;Which systems can access generated audio&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For healthcare workloads containing electronic protected health information, the &lt;a href="https://www.hhs.gov/hipaa/for-professionals/security/index.html" rel="noopener noreferrer"&gt;HIPAA Security Rule&lt;/a&gt; requires appropriate safeguards.&lt;/p&gt;

&lt;p&gt;Streaming can make the data-flow diagram more complicated because text and audio may pass through several services simultaneously.&lt;/p&gt;

&lt;p&gt;That makes explicit architecture documentation even more important.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to benchmark before choosing a streaming TTS architecture
&lt;/h2&gt;

&lt;p&gt;Do not benchmark a TTS provider using one short sentence from your laptop and call the evaluation finished.&lt;/p&gt;

&lt;p&gt;Use your actual workload.&lt;/p&gt;

&lt;p&gt;Measure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Time to first audio&lt;/li&gt;
&lt;li&gt;Median latency&lt;/li&gt;
&lt;li&gt;P95 and P99 latency&lt;/li&gt;
&lt;li&gt;Audio generation speed&lt;/li&gt;
&lt;li&gt;Performance under concurrency&lt;/li&gt;
&lt;li&gt;Chunk arrival variance&lt;/li&gt;
&lt;li&gt;Playback underruns&lt;/li&gt;
&lt;li&gt;Connection failures&lt;/li&gt;
&lt;li&gt;Retry behavior&lt;/li&gt;
&lt;li&gt;Long-response consistency&lt;/li&gt;
&lt;li&gt;Sentence-boundary quality&lt;/li&gt;
&lt;li&gt;Telephony quality if applicable&lt;/li&gt;
&lt;li&gt;Performance from your deployment region&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Also test the failure path.&lt;/p&gt;

&lt;p&gt;What happens when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The TTS connection closes halfway through a sentence?&lt;/li&gt;
&lt;li&gt;The LLM stops producing tokens?&lt;/li&gt;
&lt;li&gt;A client disappears?&lt;/li&gt;
&lt;li&gt;The user interrupts?&lt;/li&gt;
&lt;li&gt;Your playback queue grows too large?&lt;/li&gt;
&lt;li&gt;An upstream proxy buffers a supposedly streamed response?&lt;/li&gt;
&lt;li&gt;The network changes from Wi-Fi to mobile data?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A streaming architecture is only fast if every component preserves streaming behavior.&lt;/p&gt;

&lt;p&gt;One buffering reverse proxy can quietly turn your stream back into a batch response.&lt;/p&gt;

&lt;h2&gt;
  
  
  Streaming or batch: a practical decision framework
&lt;/h2&gt;

&lt;p&gt;Choose streaming TTS when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Output is generated dynamically&lt;/li&gt;
&lt;li&gt;Users are waiting interactively&lt;/li&gt;
&lt;li&gt;Time to first audio matters&lt;/li&gt;
&lt;li&gt;Responses cannot be cached effectively&lt;/li&gt;
&lt;li&gt;LLM output is incremental&lt;/li&gt;
&lt;li&gt;The application supports connection and buffer management&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Choose batch TTS when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Output is static&lt;/li&gt;
&lt;li&gt;Content can be pre-generated&lt;/li&gt;
&lt;li&gt;Audio will be reused&lt;/li&gt;
&lt;li&gt;Total render time matters more than first audio&lt;/li&gt;
&lt;li&gt;Simpler infrastructure is valuable&lt;/li&gt;
&lt;li&gt;Caching materially reduces synthesis volume&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use both when the application contains a mixture of static prompts and dynamic responses.&lt;/p&gt;

&lt;p&gt;That hybrid design is common and often more economical than forcing every piece of speech through the same path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key takeaways
&lt;/h2&gt;

&lt;p&gt;Streaming TTS is not automatically better than batch synthesis.&lt;/p&gt;

&lt;p&gt;It solves a specific problem: getting useful audio to the listener before the complete synthesis job has finished.&lt;/p&gt;

&lt;p&gt;For developers building interactive voice systems, the main lessons are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Optimize time to first audio separately from total synthesis time.&lt;/li&gt;
&lt;li&gt;Overlap LLM generation, speech synthesis, and playback when possible.&lt;/li&gt;
&lt;li&gt;Do not send individual LLM tokens blindly into TTS.&lt;/li&gt;
&lt;li&gt;Treat sentence and phrase segmentation as a quality-versus-latency decision.&lt;/li&gt;
&lt;li&gt;Use buffering to absorb network jitter, but keep the buffer small enough to preserve responsiveness.&lt;/li&gt;
&lt;li&gt;Plan for backpressure, cancellation, and user interruption.&lt;/li&gt;
&lt;li&gt;Keep authenticated speech API calls on the server.&lt;/li&gt;
&lt;li&gt;Test SSML and voice consistency specifically in streaming mode.&lt;/li&gt;
&lt;li&gt;Use batch synthesis and caching when content is predictable.&lt;/li&gt;
&lt;li&gt;Benchmark the complete production pipeline rather than the model in isolation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The difference between a voice application that feels responsive and one that feels sluggish is rarely controlled by a single latency number.&lt;/p&gt;

&lt;p&gt;It comes from how the entire pipeline overlaps work.&lt;/p&gt;

&lt;p&gt;If you want to test that architecture with your own prompts and audio pipeline, &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=streaming-tts-explained-for-developers-ux-latency-and-cost-guide"&gt;create an API key and start building with the Smallest AI API&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>voiceai</category>
      <category>webdev</category>
      <category>texttospeech</category>
    </item>
    <item>
      <title>Building Real-Time Voice AI: How STT, LLM, TTS, and Telephony Work Together</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Mon, 07 Sep 2026 09:16:29 +0000</pubDate>
      <link>https://dev.to/smallestai/building-real-time-voice-ai-how-stt-llm-tts-and-telephony-work-together-5fl</link>
      <guid>https://dev.to/smallestai/building-real-time-voice-ai-how-stt-llm-tts-and-telephony-work-together-5fl</guid>
      <description>&lt;p&gt;A voice agent can have excellent speech recognition, a capable language model, and natural text-to-speech, then still feel broken.&lt;/p&gt;

&lt;p&gt;The reason is simple: users experience the system as one conversation, not as four separate services.&lt;/p&gt;

&lt;p&gt;They notice the pause after they stop speaking. They notice when the agent talks over them. They notice when a transcription error sends the conversation in the wrong direction. They notice when high-quality synthetic speech arrives too late.&lt;/p&gt;

&lt;p&gt;That makes real-time voice AI an architecture problem as much as a model problem.&lt;/p&gt;

&lt;p&gt;A typical system has four core layers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Speech-to-text (STT)&lt;/li&gt;
&lt;li&gt;A large language model (LLM)&lt;/li&gt;
&lt;li&gt;Text-to-speech (TTS)&lt;/li&gt;
&lt;li&gt;Telephony or another real-time audio transport&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The interesting engineering happens in the boundaries between them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture in one sentence
&lt;/h2&gt;

&lt;p&gt;At its simplest, the pipeline looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User speech
 ↓
Streaming STT
 ↓
LLM
 ↓
Streaming TTS
 ↓
Telephony / WebRTC
 ↓
User hears response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That diagram is useful, but it can also be misleading.&lt;/p&gt;

&lt;p&gt;A production voice system should not behave like a serial batch-processing pipeline where STT finishes completely, then the LLM starts, then TTS starts, then audio is finally delivered.&lt;/p&gt;

&lt;p&gt;If every stage waits for the previous one to finish, latency compounds.&lt;/p&gt;

&lt;p&gt;Responsive systems stream and overlap work wherever possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. STT turns audio into usable state
&lt;/h2&gt;

&lt;p&gt;Speech-to-text, also called automatic speech recognition or ASR, converts incoming audio into text that the rest of the system can process. Developers evaluating this layer can compare the architecture against a production &lt;a href="https://smallest.ai/speech-to-text?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-architecture-stt-llm-tts-telephony"&gt;speech-to-text system&lt;/a&gt; rather than treating recognition as an isolated offline task.&lt;/p&gt;

&lt;p&gt;For offline transcription, waiting for a complete recording may be fine. Real-time conversations do not have that luxury.&lt;/p&gt;

&lt;p&gt;A streaming STT engine emits partial transcription hypotheses while the user is still speaking. That gives downstream components something to work with before the utterance has completely finished.&lt;/p&gt;

&lt;p&gt;There is a tradeoff.&lt;/p&gt;

&lt;p&gt;Partial transcripts can change as more audio arrives. If the system commits too early, one incorrectly recognized word can steer the LLM toward the wrong intent.&lt;/p&gt;

&lt;p&gt;This is why STT quality is only part of the problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  VAD can make a fast model feel slow
&lt;/h3&gt;

&lt;p&gt;Voice Activity Detection, or VAD, determines whether the user is currently speaking.&lt;/p&gt;

&lt;p&gt;It also helps answer a critical question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When has the user actually finished their turn?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the endpointer is conservative, it waits longer before deciding that speech has ended.&lt;/p&gt;

&lt;p&gt;To the infrastructure, that might be a few hundred milliseconds of uncertainty.&lt;/p&gt;

&lt;p&gt;To the user, it feels like the AI is thinking too slowly.&lt;/p&gt;

&lt;p&gt;A surprising number of apparent model-latency problems are actually endpointing problems. Smallest AI's guide to &lt;a href="https://smallest.ai/blog/voice-activity-detection-for-real-time-voice-apps-latency-false-triggers-and-production-tuning?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-architecture-stt-llm-tts-telephony"&gt;Voice Activity Detection for real-time voice apps&lt;/a&gt; goes deeper into latency, false triggers, and production tuning around this layer.&lt;/p&gt;

&lt;p&gt;Aggressive VAD settings create the opposite failure mode. The system may decide that the user has finished during a natural pause and start generating a response too early.&lt;/p&gt;

&lt;p&gt;Real-time speech systems therefore have to balance responsiveness with turn-detection accuracy.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The LLM is the reasoning layer, not the whole voice system
&lt;/h2&gt;

&lt;p&gt;Once enough transcript is available, the LLM determines what the system should say next.&lt;/p&gt;

&lt;p&gt;For many voice architectures, this is the largest individual source of compute latency.&lt;/p&gt;

&lt;p&gt;A large general-purpose model may provide excellent reasoning, but that capability comes with an inference cost. In a voice application, every extra delay becomes visible because the user is waiting for speech to resume.&lt;/p&gt;

&lt;p&gt;Two architectural choices help.&lt;/p&gt;

&lt;p&gt;First, structured workflows do not always require the largest available model. Customer support, appointment scheduling, qualification, routing, and similar flows may benefit from smaller task-focused models when their capabilities fit the workflow.&lt;/p&gt;

&lt;p&gt;Second, the LLM should stream output.&lt;/p&gt;

&lt;p&gt;If the application waits until the complete response has been generated, the TTS layer sits idle.&lt;/p&gt;

&lt;p&gt;With token streaming, the system can send usable text fragments downstream while the remainder of the response is still being generated.&lt;/p&gt;

&lt;p&gt;The LLM is important, but optimizing it while ignoring the rest of the stack is a mistake. Users judge the complete conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. TTS determines when the response becomes real
&lt;/h2&gt;

&lt;p&gt;Text-to-speech converts the LLM response back into audio. A productionctext-to-speech layer has to be evaluated by how quickly it can begin returning usable audio, not only by the quality of a completed file.&lt;/p&gt;

&lt;p&gt;For a real-time system, the critical question is not simply:&lt;/p&gt;

&lt;p&gt;"How long does synthesis take?"&lt;/p&gt;

&lt;p&gt;A more useful question is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How quickly can the system produce the first playable audio?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is commonly measured as Time to First Audio Byte, or TTFAB.&lt;/p&gt;

&lt;p&gt;A streaming TTS system begins emitting audio chunks before the entire response is available. That allows playback to start while synthesis continues.&lt;/p&gt;

&lt;p&gt;Without streaming, even a fast LLM can be followed by a noticeable pause while the speech engine waits for the complete sentence and synthesizes it.&lt;/p&gt;

&lt;p&gt;This is another reason component benchmarks cannot be evaluated in isolation. What matters is how quickly usable output moves across the entire chain.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Telephony is part of the architecture, not plumbing
&lt;/h2&gt;

&lt;p&gt;The final layer carries audio between your application and the user.&lt;/p&gt;

&lt;p&gt;Depending on the product, that might involve:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The public telephone network&lt;/li&gt;
&lt;li&gt;SIP&lt;/li&gt;
&lt;li&gt;VoIP infrastructure&lt;/li&gt;
&lt;li&gt;A PBX or cloud phone system&lt;/li&gt;
&lt;li&gt;WebRTC in a browser or application&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This layer creates its own failure modes.&lt;/p&gt;

&lt;p&gt;Networks introduce jitter and variable delay. Codecs compress audio. Packet loss affects intelligibility. Transcoding can alter the audio reaching the STT system.&lt;/p&gt;

&lt;p&gt;Those effects often appear only under production traffic.&lt;/p&gt;

&lt;p&gt;A speech recognition model that performs well on clean microphone recordings may behave differently after audio has passed through a telephone codec.&lt;/p&gt;

&lt;p&gt;Similarly, perfect TTS output generated in the backend is irrelevant if the delivery path degrades it before the caller hears it.&lt;/p&gt;

&lt;p&gt;Telephony therefore belongs inside the performance budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  Streaming changes the shape of the pipeline
&lt;/h2&gt;

&lt;p&gt;The most important architectural idea in real-time voice AI is overlap.&lt;/p&gt;

&lt;p&gt;Instead of this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;STT completes
 ↓
LLM completes
 ↓
TTS completes
 ↓
Audio plays
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you want behavior closer to this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;STT ===============
LLM =============
TTS =============
Playback ===========
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The stages remain logically separate, but execution overlaps.&lt;/p&gt;

&lt;p&gt;STT can emit partial text while the user is speaking.&lt;/p&gt;

&lt;p&gt;The LLM can begin once enough stable context exists.&lt;/p&gt;

&lt;p&gt;TTS can begin when it receives a usable text fragment.&lt;/p&gt;

&lt;p&gt;Playback can begin as soon as the first synthesized audio reaches the transport layer.&lt;/p&gt;

&lt;p&gt;This architecture turns latency from a simple sum into a coordination problem.&lt;/p&gt;

&lt;p&gt;For a deeper treatment of this pattern, the Smallest AI guide to &lt;a href="https://smallest.ai/blog/why-streaming-architecture-is-non-negotiable-for-real-time-voice-agents?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-architecture-stt-llm-tts-telephony"&gt;streaming architecture for real-time voice agents&lt;/a&gt; explains why streaming has to extend across the pipeline rather than being added to only one component.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the latency budget goes
&lt;/h2&gt;

&lt;p&gt;An illustrative latency budget might look like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Typical contribution&lt;/th&gt;
&lt;th&gt;Primary optimization lever&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Streaming STT&lt;/td&gt;
&lt;td&gt;50-100 ms&lt;/td&gt;
&lt;td&gt;Streaming transcription and VAD tuning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM inference&lt;/td&gt;
&lt;td&gt;150-300 ms&lt;/td&gt;
&lt;td&gt;Smaller models, token streaming, caching&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TTS synthesis&lt;/td&gt;
&lt;td&gt;50-150 ms to first audio&lt;/td&gt;
&lt;td&gt;Streaming TTS and low-latency models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Telephony / network&lt;/td&gt;
&lt;td&gt;20-80 ms&lt;/td&gt;
&lt;td&gt;Deployment location, codec choice, WebRTC&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These values should be treated as architectural estimates, not universal guarantees. Real performance varies with model size, infrastructure, geography, audio conditions, and the transport path.&lt;/p&gt;

&lt;p&gt;The important pattern is the distribution.&lt;/p&gt;

&lt;p&gt;The LLM is commonly the largest individual contributor, but optimizing only the LLM does not guarantee a responsive conversation.&lt;/p&gt;

&lt;p&gt;You might reduce inference time and still lose the improvement because VAD waits too long.&lt;/p&gt;

&lt;p&gt;You might deploy faster TTS and then add network latency between regions.&lt;/p&gt;

&lt;p&gt;You might improve STT accuracy while choosing a model that processes audio too slowly for the desired interaction.&lt;/p&gt;

&lt;p&gt;Voice latency is an end-to-end property.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frqjxnwkn0k3qzy79wxfs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frqjxnwkn0k3qzy79wxfs.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What a production deployment actually looks like
&lt;/h2&gt;

&lt;p&gt;Consider a contact-center workflow.&lt;/p&gt;

&lt;p&gt;A caller enters through SIP into an existing PBX or cloud telephony environment.&lt;/p&gt;

&lt;p&gt;Streaming STT begins processing audio as it arrives.&lt;/p&gt;

&lt;p&gt;The LLM receives the transcript and determines the next response. When the workflow requires external information, the application can consult a knowledge base, CRM, scheduling system, or another business tool through application logic or function calling.&lt;/p&gt;

&lt;p&gt;The response is streamed toward TTS.&lt;/p&gt;

&lt;p&gt;Synthesized audio is then returned through the telephony layer.&lt;/p&gt;

&lt;p&gt;If the conversation moves outside the system's permitted scope or requires human intervention, the call can be escalated according to the application's fallback logic.&lt;/p&gt;

&lt;p&gt;The same architectural pattern can be used in healthcare scheduling, support, commerce, sales, and other conversational applications. The requirements change, but the underlying coordination problem remains similar.&lt;/p&gt;

&lt;p&gt;For teams that prefer an integrated implementation path rather than assembling every layer independently, the &lt;a href="https://smallest.ai/voice-agents?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-architecture-stt-llm-tts-telephony"&gt;Smallest AI voice-agent platform&lt;/a&gt; brings agent configuration, telephony, integrations, and deployment workflows into one environment. The related guide to &lt;a href="https://smallest.ai/blog/mastering-voice-bot-architecture-a-deep-dive-with-smallest-ai-s-atoms-sdk?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-architecture-stt-llm-tts-telephony"&gt;voice bot architecture&lt;/a&gt; shows how the speech and orchestration layers can be connected inside a production SDK.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three architecture misconceptions worth avoiding
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Misconception 1: Better accuracy automatically creates a better real-time system
&lt;/h3&gt;

&lt;p&gt;Accuracy and latency have to be evaluated together.&lt;/p&gt;

&lt;p&gt;A larger speech model may improve recognition but require more processing per audio chunk.&lt;/p&gt;

&lt;p&gt;For transcription workloads where accuracy dominates, that trade can make sense.&lt;/p&gt;

&lt;p&gt;For a live customer conversation, a modest accuracy improvement may not justify a delay that repeatedly disrupts turn-taking.&lt;/p&gt;

&lt;p&gt;There is no single correct tradeoff. It depends on the application.&lt;/p&gt;

&lt;p&gt;The important point is that an offline accuracy benchmark cannot tell you whether a model is appropriate for a real-time conversation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Misconception 2: The LLM is the voice AI
&lt;/h3&gt;

&lt;p&gt;It is not.&lt;/p&gt;

&lt;p&gt;A powerful LLM connected to poor STT can reason about the wrong transcript.&lt;/p&gt;

&lt;p&gt;A powerful LLM connected to slow TTS still feels slow.&lt;/p&gt;

&lt;p&gt;A powerful LLM behind badly configured VAD may constantly interrupt users or leave awkward pauses.&lt;/p&gt;

&lt;p&gt;A powerful LLM sent through an unreliable telephony layer still produces an unreliable product.&lt;/p&gt;

&lt;p&gt;The voice agent is the system formed by all of these components.&lt;/p&gt;

&lt;h3&gt;
  
  
  Misconception 3: Speech-to-speech eliminates the modular pipeline everywhere
&lt;/h3&gt;

&lt;p&gt;Speech-to-speech models can accept audio and produce audio without exposing explicit text stages in the same way as a traditional STT + LLM + TTS architecture.&lt;/p&gt;

&lt;p&gt;That can be useful in latency-sensitive applications.&lt;/p&gt;

&lt;p&gt;It does not automatically make modular architectures obsolete.&lt;/p&gt;

&lt;p&gt;Many production systems still require clear insertion points for business logic, predictable workflows, auditability, tool execution, and explicit control over individual stages.&lt;/p&gt;

&lt;p&gt;For those systems, keeping STT, reasoning, and TTS as identifiable components can be valuable.&lt;/p&gt;

&lt;p&gt;Speech-to-speech is therefore another architectural option, not a universal replacement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production problems start where the happy path ends
&lt;/h2&gt;

&lt;p&gt;Choosing models is only the beginning.&lt;/p&gt;

&lt;p&gt;Real users interrupt, hesitate, change their minds, speak over the system, call from noisy environments, and ask questions the application was never designed to answer.&lt;/p&gt;

&lt;p&gt;Several decisions become critical once a system reaches production. Smallest AI's article on &lt;a href="https://smallest.ai/blog/ai-voice-agents-architecture-voice-models-use-cases-and-safety-guardrails?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-architecture-stt-llm-tts-telephony"&gt;designing AI voice agents&lt;/a&gt; provides additional context on architecture, use cases, and safety guardrails beyond the basic STT + LLM + TTS pipeline.&lt;/p&gt;

&lt;h3&gt;
  
  
  Interruption handling
&lt;/h3&gt;

&lt;p&gt;Users will talk while the AI is speaking.&lt;/p&gt;

&lt;p&gt;The system needs to detect the new speech, cancel or stop the active TTS output, update the conversational state, and process the interruption.&lt;/p&gt;

&lt;p&gt;If it cannot, the agent talks over people.&lt;/p&gt;

&lt;p&gt;That immediately makes the experience feel mechanical.&lt;/p&gt;

&lt;h3&gt;
  
  
  Context management
&lt;/h3&gt;

&lt;p&gt;Conversation history grows on every turn.&lt;/p&gt;

&lt;p&gt;Sending an indefinitely growing context back to the LLM can increase latency and processing requirements.&lt;/p&gt;

&lt;p&gt;Common approaches include summarizing older turns or maintaining a sliding context window while preserving important application state.&lt;/p&gt;

&lt;p&gt;The right approach depends on how much historical context the workflow genuinely needs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Fallback and escalation
&lt;/h3&gt;

&lt;p&gt;A production system needs defined behavior for situations such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Low-confidence transcription&lt;/li&gt;
&lt;li&gt;Unsupported requests&lt;/li&gt;
&lt;li&gt;Missing business data&lt;/li&gt;
&lt;li&gt;Tool failures&lt;/li&gt;
&lt;li&gt;Requests that require human handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fallback logic should be designed as part of the architecture, not added after deployment.&lt;/p&gt;

&lt;h3&gt;
  
  
  Edge versus cloud placement
&lt;/h3&gt;

&lt;p&gt;Every network hop matters.&lt;/p&gt;

&lt;p&gt;Moving audio between regions can consume part of the latency budget before the models have done any work.&lt;/p&gt;

&lt;p&gt;Some architectures move latency-sensitive processing closer to the user while keeping other components in centralized cloud infrastructure.&lt;/p&gt;

&lt;p&gt;The correct boundary depends on infrastructure, model requirements, geography, and operational complexity.&lt;/p&gt;

&lt;h3&gt;
  
  
  Codec and audio quality
&lt;/h3&gt;

&lt;p&gt;Phone audio is not the same as clean studio audio.&lt;/p&gt;

&lt;p&gt;Codecs such as G.711 and Opus affect what reaches the speech recognition system and what the caller ultimately hears.&lt;/p&gt;

&lt;p&gt;When production accuracy suddenly differs from testing, inspect the audio path before assuming the model itself has regressed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prototyping the stack with Smallest AI
&lt;/h2&gt;

&lt;p&gt;The current &lt;a href="https://smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-architecture-stt-llm-tts-telephony"&gt;Smallest AI voice platform&lt;/a&gt; spans speech recognition, speech generation, speech-to-speech, and voice-agent workflows. For developers evaluating this architecture programmatically, the &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-architecture-stt-llm-tts-telephony"&gt;Smallest AI API&lt;/a&gt; provides the developer entry point for working with the relevant voice models and services.&lt;/p&gt;

&lt;p&gt;The useful way to test a real-time architecture is not to evaluate only isolated model output.&lt;/p&gt;

&lt;p&gt;Run the system using representative audio, your actual deployment geography, the codecs your application will use, realistic conversation lengths, and real interruption behavior.&lt;/p&gt;

&lt;p&gt;Measure at least:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Time from user speech ending to the first LLM output&lt;/li&gt;
&lt;li&gt;Time to first synthesized audio&lt;/li&gt;
&lt;li&gt;End-to-end turn latency&lt;/li&gt;
&lt;li&gt;STT errors under real audio conditions&lt;/li&gt;
&lt;li&gt;False VAD triggers&lt;/li&gt;
&lt;li&gt;Missed end-of-turn events&lt;/li&gt;
&lt;li&gt;Barge-in behavior&lt;/li&gt;
&lt;li&gt;Network and telephony delay&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If authenticated API requests are added to your implementation, keep API credentials on the server. Do not expose them in browser JavaScript, mobile application code, public repositories, screenshots, URLs, or client-side logs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture is the product
&lt;/h2&gt;

&lt;p&gt;The failure mode in real-time voice AI is rarely one component completely breaking.&lt;/p&gt;

&lt;p&gt;More often, the experience degrades through accumulation.&lt;/p&gt;

&lt;p&gt;A little VAD delay.&lt;/p&gt;

&lt;p&gt;A slow first LLM token.&lt;/p&gt;

&lt;p&gt;A TTS buffer that waits too long.&lt;/p&gt;

&lt;p&gt;An unnecessary network hop.&lt;/p&gt;

&lt;p&gt;A codec mismatch.&lt;/p&gt;

&lt;p&gt;A barge-in handler that does not cancel playback quickly enough.&lt;/p&gt;

&lt;p&gt;Each problem may look small in isolation. Together, they determine whether the conversation feels natural.&lt;/p&gt;

&lt;p&gt;That is the central architectural lesson: optimize the handoffs, not just the parts.&lt;/p&gt;

&lt;p&gt;STT has to stream.&lt;/p&gt;

&lt;p&gt;The reasoning layer has to produce output early enough for downstream synthesis.&lt;/p&gt;

&lt;p&gt;TTS has to return playable audio incrementally.&lt;/p&gt;

&lt;p&gt;The delivery layer has to preserve both timing and audio quality.&lt;/p&gt;

&lt;p&gt;And the entire system has to survive interruptions, imperfect networks, growing context, and unpredictable users.&lt;/p&gt;

&lt;p&gt;If you are building this pipeline, the best test is your own application under realistic conditions. &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=real-time-voice-ai-architecture-stt-llm-tts-telephony"&gt;Create an API key and prototype the voice workflow with Smallest AI&lt;/a&gt;, then measure the complete path from live audio input to the first audio returned to the user.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>voiceai</category>
      <category>machinelearning</category>
      <category>webdev</category>
    </item>
    <item>
      <title>How to Choose a Voice Agent API: Architecture, Latency, Streaming, and Stack Trade-Offs</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Wed, 26 Aug 2026 07:54:28 +0000</pubDate>
      <link>https://dev.to/smallestai/how-to-choose-a-voice-agent-api-architecture-latency-streaming-and-stack-trade-offs-3mhp</link>
      <guid>https://dev.to/smallestai/how-to-choose-a-voice-agent-api-architecture-latency-streaming-and-stack-trade-offs-3mhp</guid>
      <description>&lt;p&gt;A voice interface used to mean a phone tree: press 1 for billing, press 2 for support, and wait for the next prompt.&lt;/p&gt;

&lt;p&gt;Modern voice applications have a much higher bar. Users expect quick turn-taking, context that survives across multiple turns, and responses that start before the silence feels like a failure.&lt;/p&gt;

&lt;p&gt;That changes voice from a simple speech feature into a real-time systems problem.&lt;/p&gt;

&lt;p&gt;A Voice Agent API sits at the infrastructure layer of that problem. Instead of exposing transcription or speech generation as isolated utilities, it helps coordinate the conversational loop: listen, understand, reason, and respond.&lt;/p&gt;

&lt;p&gt;For developers, the difficult part is rarely getting each individual component to work. The harder problem is getting the entire pipeline to behave like one responsive system.&lt;/p&gt;

&lt;p&gt;This guide breaks down that architecture, where latency enters the stack, why streaming matters, and what to evaluate before committing to a production setup.&lt;/p&gt;

&lt;p&gt;For a broader architectural view, Smallest AI's &lt;a href="https://smallest.ai/blog/ai-voice-agents-architecture-voice-models-use-cases-and-safety-guardrails?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-an-ai-voice-agent-using-atoms-api"&gt;AI voice agent architecture guide&lt;/a&gt; covers the underlying voice models, common use cases, and deployment considerations.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is a Voice Agent API?
&lt;/h2&gt;

&lt;p&gt;A Voice Agent API is a programmatic interface for building conversational systems that can listen to speech, reason about what was said, and respond with synthesized audio inside a coordinated pipeline.&lt;/p&gt;

&lt;p&gt;The distinction from standalone speech APIs matters.&lt;/p&gt;

&lt;p&gt;A text-to-speech API handles:&lt;/p&gt;

&lt;p&gt;text → audio&lt;/p&gt;

&lt;p&gt;A speech-to-text API handles:&lt;/p&gt;

&lt;p&gt;audio → text&lt;/p&gt;

&lt;p&gt;A typical voice-agent pipeline handles:&lt;/p&gt;

&lt;p&gt;speech → transcription → reasoning → response text → synthesized speech&lt;/p&gt;

&lt;p&gt;It also needs to preserve conversation state between turns and handle interaction patterns that one-shot speech APIs do not need to solve.&lt;/p&gt;

&lt;p&gt;Examples include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;A caller correcting something they said earlier.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The user interrupting the agent while it is speaking.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;A later question depending on information from an earlier turn.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;An ambiguous request requiring clarification.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;A tool or backend action taking long enough that the conversation still needs to feel responsive.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important idea is that the "agent" is not simply STT plus TTS. It is the orchestration of speech recognition, reasoning, state, and speech generation into one conversational system.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stack behind a voice agent
&lt;/h2&gt;

&lt;p&gt;The most common architecture is a cascading pipeline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Audio
↓
Speech-to-Text
↓
Language Model
↓
Text-to-Speech
↓
Audio
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The concept is straightforward.&lt;/p&gt;

&lt;p&gt;The latency behavior is not.&lt;/p&gt;

&lt;p&gt;Every boundary introduces work: network transport, buffering, model inference, serialization, queueing, and coordination between services.&lt;/p&gt;

&lt;p&gt;If every stage waits for the previous stage to finish completely, those delays accumulate.&lt;/p&gt;

&lt;p&gt;That is why production voice-agent architecture is usually less about optimizing one model in isolation and more about optimizing the entire path from the end of the user's turn to the beginning of the agent's response.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why streaming changes the latency budget
&lt;/h2&gt;

&lt;p&gt;Consider two implementations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Batch pipeline
&lt;/h3&gt;

&lt;p&gt;A batch-oriented system might behave like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Wait for the user to finish speaking.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Upload or finalize the complete audio segment.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Wait for the complete transcript.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Send the transcript to the language model.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Wait for the complete model response.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Send the complete response to TTS.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Wait for enough synthesized audio.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Start playback.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each stage blocks the next.&lt;/p&gt;

&lt;p&gt;That architecture can work for offline transcription or generated narration. It is poorly suited to natural turn-taking.&lt;/p&gt;

&lt;h3&gt;
  
  
  Streaming pipeline
&lt;/h3&gt;

&lt;p&gt;A streaming architecture allows useful partial results to move downstream as soon as they become available.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User speech
↓ partial audio
Streaming STT
↓ incremental/final transcript
Reasoning layer
↓ streamed response tokens
Streaming TTS
↓ audio chunks
Playback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The major difference is overlap.&lt;/p&gt;

&lt;p&gt;The reasoning layer can begin processing as soon as sufficient speech context exists. TTS can begin receiving generated text without waiting for an entire paragraph. Playback can start while later audio is still being synthesized.&lt;/p&gt;

&lt;p&gt;Streaming does not remove latency from individual models.&lt;/p&gt;

&lt;p&gt;It stops the system from unnecessarily serializing every unit of work.&lt;/p&gt;

&lt;p&gt;For real-time voice applications, that distinction is fundamental.&lt;/p&gt;

&lt;h2&gt;
  
  
  Speech-to-text: the listening layer
&lt;/h2&gt;

&lt;p&gt;The STT system converts incoming audio into machine-readable text.&lt;/p&gt;

&lt;p&gt;Its quality limits everything downstream.&lt;/p&gt;

&lt;p&gt;If the recognizer repeatedly mishears names, numbers, accented speech, or domain-specific terminology, the reasoning layer starts from corrupted input.&lt;/p&gt;

&lt;p&gt;The result can be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Incorrect answers.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Wrong tool calls.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Extra clarification turns.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Failed task completion.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Increased conversation length.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A common transcription metric is Word Error Rate, or WER. Smallest AI has a separate &lt;a href="https://smallest.ai/blog/word-error-rate-explained-why-it-matters-for-voice-agent-quality?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-an-ai-voice-agent-using-atoms-api"&gt;guide to Word Error Rate for voice agents&lt;/a&gt; if you want to go deeper into evaluation.&lt;/p&gt;

&lt;p&gt;For voice agents, however, aggregate transcription accuracy is not the only consideration.&lt;/p&gt;

&lt;p&gt;You also need to evaluate how the STT system behaves while audio is still arriving.&lt;/p&gt;

&lt;p&gt;Questions worth testing include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;How quickly do partial transcripts arrive?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How stable are those partial transcripts?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How does endpoint detection behave?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;What happens with background noise?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How are interruptions handled?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;What happens when the network briefly degrades?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Smallest AI's current model stack uses Pulse as its streaming STT component.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reasoning layer
&lt;/h2&gt;

&lt;p&gt;The transcript then moves into a conversational model.&lt;/p&gt;

&lt;p&gt;That model has several responsibilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Understand what the user means.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Maintain conversation context.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Decide what should happen next.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Potentially call tools or external services.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Produce the text that will become speech.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For text applications, users may tolerate visible generation.&lt;/p&gt;

&lt;p&gt;Voice is less forgiving.&lt;/p&gt;

&lt;p&gt;A delay that looks normal in a chat interface becomes dead air in a phone call.&lt;/p&gt;

&lt;p&gt;That makes time-to-first-token important, but optimizing only that number is not enough. A voice-agent latency budget can also include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;End-of-user-turn detection
+ STT finalization
+ orchestration
+ model time-to-first-token
+ tool execution when required
+ TTS startup
+ network and playback buffering
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The right question is therefore not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Which LLM is fastest?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How much time passes between the user's conversational turn and the first useful audio response?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Smaller conversational models can be useful when reducing reasoning startup time is more important than maximizing general-purpose generation capability.&lt;/p&gt;

&lt;p&gt;Smallest AI uses Electron as the conversational model in its own Pulse → Electron → Lightning pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Text-to-speech: the voice layer
&lt;/h2&gt;

&lt;p&gt;TTS turns the generated response into audible speech.&lt;/p&gt;

&lt;p&gt;This is where the assistant becomes a voice experience rather than a text system with audio attached.&lt;/p&gt;

&lt;p&gt;Two properties matter immediately.&lt;/p&gt;

&lt;p&gt;First, the synthesized output needs to sound appropriate for the application.&lt;/p&gt;

&lt;p&gt;Second, audio needs to become available quickly enough to preserve conversational rhythm.&lt;/p&gt;

&lt;p&gt;Streaming TTS is valuable because the synthesizer can begin producing audio from partial generated text instead of waiting for the language model to finish an entire response.&lt;/p&gt;

&lt;p&gt;That creates another opportunity for overlap:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LLM: "I can help you..."
↓
TTS: begins synthesis
LLM: "...reschedule that appointment..."
↓
TTS: continues synthesis
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Smallest AI uses Lightning for streaming TTS in its cascading voice-agent stack. Its documentation also describes Hydra as a speech-to-speech model within the broader speech stack.&lt;/p&gt;

&lt;p&gt;For more detail on this part of the latency budget, see the &lt;a href="https://smallest.ai/blog/neural-tts-latency-explained-how-to-build-faster-ai-voice-agents?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-an-ai-voice-agent-using-atoms-api"&gt;neural TTS latency guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Voice Agent API architectures are not all the same
&lt;/h2&gt;

&lt;p&gt;Voice APIs generally fall into a few architectural patterns.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Approach&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Setup effort&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Control&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Latency implications&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Good fit&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Managed voice-agent platform&lt;/td&gt;
&lt;td&gt;Lower&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Orchestration handled by platform&lt;/td&gt;
&lt;td&gt;Teams that want to deploy complete agents quickly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Modular STT + LLM + TTS&lt;/td&gt;
&lt;td&gt;Higher&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;More inter-service boundaries to manage&lt;/td&gt;
&lt;td&gt;Teams that need deep control over individual components&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speech-to-speech&lt;/td&gt;
&lt;td&gt;Varies&lt;/td&gt;
&lt;td&gt;Depends on implementation&lt;/td&gt;
&lt;td&gt;Can reduce intermediate text-stage overhead&lt;/td&gt;
&lt;td&gt;Workloads suited to direct audio-to-audio interaction&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;None is universally correct.&lt;/p&gt;

&lt;p&gt;A managed system reduces orchestration work, but gives the platform more responsibility for implementation details.&lt;/p&gt;

&lt;p&gt;A modular stack gives developers more freedom to choose each component independently, but they also inherit retries, network hops, observability, synchronization, interruption handling, and failure recovery between services.&lt;/p&gt;

&lt;p&gt;Speech-to-speech can compress portions of the traditional cascade, but the available control, tooling, and state-management model depends heavily on the implementation.&lt;/p&gt;

&lt;p&gt;The choice should follow the application's actual requirements rather than the architecture that looks simplest in a demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the Smallest AI API fits
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fieuo595sc0khh9rf3kt6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fieuo595sc0khh9rf3kt6.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
For a concrete example of a vertically integrated stack, &lt;a href="https://smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-an-ai-voice-agent-using-atoms-api"&gt;Smallest AI&lt;/a&gt; exposes speech models and voice-agent infrastructure within the same ecosystem.&lt;/p&gt;

&lt;p&gt;The current stack includes Pulse for STT, Electron for conversational reasoning, and Lightning for TTS, while Atoms provides the higher-level voice-agent platform.&lt;/p&gt;

&lt;p&gt;Developers who want to evaluate the stack programmatically can use the &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-an-ai-voice-agent-using-atoms-api"&gt;Smallest AI API&lt;/a&gt;, while &lt;a href="https://smallest.ai/voice-agents?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-an-ai-voice-agent-using-atoms-api"&gt;Smallest AI Voice Agents&lt;/a&gt; is the relevant product layer for building and deploying complete agents.&lt;/p&gt;

&lt;p&gt;If you want an implementation-oriented follow-up, the &lt;a href="https://smallest.ai/blog/how-to-build-an-ai-voice-agent-using-atoms-api?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-an-ai-voice-agent-using-atoms-api"&gt;Atoms API voice-agent tutorial&lt;/a&gt; walks through the production-minded setup in more detail.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where voice-agent APIs are being used
&lt;/h2&gt;

&lt;p&gt;The architecture applies anywhere a machine needs to carry on a real-time spoken conversation rather than simply transcribe or narrate.&lt;/p&gt;
&lt;h3&gt;
  
  
  Customer support
&lt;/h3&gt;

&lt;p&gt;An inbound voice agent might:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Answer common questions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Retrieve account information.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Perform a structured workflow.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Route a complex case to a human.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The interesting engineering challenge is not just answering correctly. The system also needs interruption handling, escalation logic, reliable tool execution, and predictable response timing.&lt;/p&gt;
&lt;h3&gt;
  
  
  Healthcare scheduling and intake
&lt;/h3&gt;

&lt;p&gt;Structured tasks such as appointment confirmation, rescheduling, and intake can map naturally to conversational workflows.&lt;/p&gt;

&lt;p&gt;These deployments also raise stronger requirements around data handling, security, escalation, and failure behavior.&lt;/p&gt;
&lt;h3&gt;
  
  
  Sales development
&lt;/h3&gt;

&lt;p&gt;Outbound agents can handle structured qualification conversations, respond to common objections, collect information, and schedule the next step.&lt;/p&gt;

&lt;p&gt;The underlying architecture is still the same: listen, understand, decide, act, and respond without introducing unnatural pauses between stages.&lt;/p&gt;
&lt;h3&gt;
  
  
  Accessibility
&lt;/h3&gt;

&lt;p&gt;Voice interfaces can also provide hands-free interaction for people who cannot efficiently use conventional input methods.&lt;/p&gt;

&lt;p&gt;In that context, responsiveness is more than polish. Latency and predictable turn-taking can directly affect usability.&lt;/p&gt;
&lt;h2&gt;
  
  
  Three voice-agent mistakes developers make
&lt;/h2&gt;
&lt;h3&gt;
  
  
  1. Treating latency as a TTS-only problem
&lt;/h3&gt;

&lt;p&gt;TTS is visible because it is the final step before the user hears something.&lt;/p&gt;

&lt;p&gt;That makes it an easy component to blame.&lt;/p&gt;

&lt;p&gt;But the total delay can come from several places:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Endpoint detection
→ STT finalization
→ orchestration
→ model startup
→ tools
→ TTS
→ buffering
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Optimizing a fast synthesizer will not rescue a pipeline that spends most of its latency budget somewhere upstream.&lt;/p&gt;

&lt;p&gt;Measure the full turn.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Assuming a strong LLM automatically creates a strong voice agent
&lt;/h3&gt;

&lt;p&gt;A voice agent is an end-to-end system.&lt;/p&gt;

&lt;p&gt;An excellent reasoning model cannot fully compensate for poor transcription, unreliable endpointing, slow tool calls, unstable networking, or unnatural synthesis.&lt;/p&gt;

&lt;p&gt;For production systems, evaluate the interaction rather than ranking components independently.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Treating voice identity as an afterthought
&lt;/h3&gt;

&lt;p&gt;For brand-facing voice applications, the selected voice affects how users perceive the experience.&lt;/p&gt;

&lt;p&gt;Developers should evaluate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Pronunciation consistency.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Prosody.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Speaking rate.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Stability across longer responses.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Voice customization requirements.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Whether cloned or custom voices are actually needed.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Voice identity also creates trust and security considerations.&lt;/p&gt;

&lt;p&gt;The W3C's &lt;a href="https://www.w3.org/2025/10/smartagents-workshop/report.html" rel="noopener noreferrer"&gt;Smart Voice Agents workshop report&lt;/a&gt; highlights areas such as privacy-preserving authentication, user identification, accessibility, real-time interaction, and interoperability as important issues for voice-agent systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to evaluate before choosing a Voice Agent API
&lt;/h2&gt;

&lt;p&gt;A polished demo can hide architecture problems that become obvious under production traffic.&lt;/p&gt;

&lt;p&gt;Before selecting a stack, test the following areas.&lt;/p&gt;

&lt;h3&gt;
  
  
  End-to-end response latency
&lt;/h3&gt;

&lt;p&gt;Do not evaluate only model inference numbers.&lt;/p&gt;

&lt;p&gt;Measure the complete conversation path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;user finishes speaking
↓
agent detects turn boundary
↓
transcription completes
↓
reasoning begins
↓
response starts
↓
speech synthesis starts
↓
first audio reaches user
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Test under realistic network conditions and expected concurrency.&lt;/p&gt;

&lt;h3&gt;
  
  
  Streaming at every relevant stage
&lt;/h3&gt;

&lt;p&gt;A system that streams TTS but batches STT still has a major blocking stage.&lt;/p&gt;

&lt;p&gt;Verify how streaming works for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Incoming audio.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Partial transcripts.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Reasoning output.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Tool execution where relevant.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Synthesized audio.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Playback.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Also check what happens when a stream disconnects or a caller interrupts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Interruption and turn-taking behavior
&lt;/h3&gt;

&lt;p&gt;Real users do not wait politely for an agent to finish.&lt;/p&gt;

&lt;p&gt;They pause.&lt;/p&gt;

&lt;p&gt;They restart sentences.&lt;/p&gt;

&lt;p&gt;They interrupt.&lt;/p&gt;

&lt;p&gt;They say "actually, never mind."&lt;/p&gt;

&lt;p&gt;A production voice agent needs explicit behavior for those cases.&lt;/p&gt;

&lt;p&gt;Test:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Barge-in.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;False endpoint detection.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Long pauses inside a sentence.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Consecutive short utterances.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Double-talk.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Network jitter.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Voice customization
&lt;/h3&gt;

&lt;p&gt;If voice consistency matters to the product, determine what can actually be controlled.&lt;/p&gt;

&lt;p&gt;Possible considerations include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Voice selection.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Voice cloning.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Prosody.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Speaking rate.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Pronunciation behavior.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Language support.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not assume every API exposes the same controls.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cost under realistic conversations
&lt;/h3&gt;

&lt;p&gt;Pricing can be based on different units: minutes, characters, requests, individual model usage, hosting, telephony, or combinations of these.&lt;/p&gt;

&lt;p&gt;The useful calculation is not simply the advertised unit price.&lt;/p&gt;

&lt;p&gt;Model the workload you expect to run.&lt;/p&gt;

&lt;p&gt;Estimate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;daily conversations
× average conversation duration
× turns per conversation
× speech/model usage
× infrastructure and telephony costs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Smallest AI guide to &lt;a href="https://smallest.ai/blog/what-are-the-true-costs-associated-with-operating-a-voice-agent-at-scale?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-an-ai-voice-agent-using-atoms-api"&gt;voice-agent operating costs at scale&lt;/a&gt; goes deeper into that planning problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the stack itself matters
&lt;/h2&gt;

&lt;p&gt;A common voice-agent prototype starts with individually strong components.&lt;/p&gt;

&lt;p&gt;The team chooses an STT provider, an LLM, and a TTS service. Each performs well on its own.&lt;/p&gt;

&lt;p&gt;Then the components are connected.&lt;/p&gt;

&lt;p&gt;That is when the invisible overhead becomes visible.&lt;/p&gt;

&lt;p&gt;Suppose one stage needs to wait for an endpointing decision. Another service introduces network latency. The language model takes time before generating its first useful output. TTS then needs additional time before audio playback can begin.&lt;/p&gt;

&lt;p&gt;No individual service has to be dramatically slow for the combined interaction to feel sluggish.&lt;/p&gt;

&lt;p&gt;That is the "latency tax" of orchestration.&lt;/p&gt;

&lt;p&gt;A unified architecture can reduce some of those boundaries because the components are designed to work together. A modular architecture can still perform extremely well, but developers have to engineer those boundaries themselves.&lt;/p&gt;

&lt;p&gt;Smallest AI's current stack takes the integrated approach: Pulse, Electron, and Lightning can be used as one streaming pipeline, while Atoms adds the voice-agent orchestration layer.&lt;/p&gt;

&lt;p&gt;The relevant question when comparing this against a modular stack is not whether one architecture is theoretically superior.&lt;/p&gt;

&lt;p&gt;It is whether you want to own:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Streaming coordination.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Turn detection.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Inter-service communication.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Retry behavior.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Observability.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;State management.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Interruption handling.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Latency tuning at every boundary.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That decision often matters more than choosing between two models with similar benchmark numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical stack-selection checklist
&lt;/h2&gt;

&lt;p&gt;Before locking in a provider, answer these questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Can every latency-critical stage stream?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;What is the end-to-end time to first useful audio?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How does the system handle barge-in?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Can you observe latency by component?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;What happens when STT, reasoning, or TTS fails?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Can conversations recover from network interruptions?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How is conversation state preserved?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Which components can be replaced later?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;What voice controls are actually exposed?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How does pricing behave for your expected conversation length?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Are security and compliance requirements compatible with your deployment?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Can you reproduce realistic production conditions during testing?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You should be able to answer those questions before committing significant application logic to the stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is the difference between a Voice Agent API and a TTS API?
&lt;/h3&gt;

&lt;p&gt;A TTS API converts text into synthesized audio.&lt;/p&gt;

&lt;p&gt;A Voice Agent API coordinates a larger conversational loop that includes listening, reasoning, dialogue state, and speech output.&lt;/p&gt;

&lt;p&gt;TTS is one component of that architecture.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does a voice agent require separate STT, LLM, and TTS providers?
&lt;/h3&gt;

&lt;p&gt;No.&lt;/p&gt;

&lt;p&gt;You can assemble the components independently or use a platform that coordinates several stages for you.&lt;/p&gt;

&lt;p&gt;A modular architecture provides more component-level control. A unified platform can reduce integration work and the number of service boundaries you have to manage.&lt;/p&gt;

&lt;h3&gt;
  
  
  How much latency is acceptable?
&lt;/h3&gt;

&lt;p&gt;There is no single number that fits every application.&lt;/p&gt;

&lt;p&gt;Instead of optimizing toward an isolated benchmark, measure the complete turn from the end of the user's speech to the first useful agent audio.&lt;/p&gt;

&lt;p&gt;For conversational applications, lower and more predictable latency generally produces better turn-taking than a pipeline with long or inconsistent pauses.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is streaming TTS enough?
&lt;/h3&gt;

&lt;p&gt;Usually not.&lt;/p&gt;

&lt;p&gt;If STT or the reasoning layer still waits for complete inputs and outputs, those stages remain sequential bottlenecks.&lt;/p&gt;

&lt;p&gt;The biggest benefit comes from designing the entire latency-sensitive path around incremental processing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Should I use speech-to-speech instead?
&lt;/h3&gt;

&lt;p&gt;It depends on the application.&lt;/p&gt;

&lt;p&gt;Speech-to-speech architectures can reduce or reorganize parts of the traditional STT → LLM → TTS cascade, but they may expose a different set of controls, debugging surfaces, state-management mechanisms, and integration options.&lt;/p&gt;

&lt;p&gt;Evaluate the architecture against your workflow rather than latency alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final takeaway
&lt;/h2&gt;

&lt;p&gt;A Voice Agent API is not simply another speech endpoint.&lt;/p&gt;

&lt;p&gt;It is the infrastructure around a real-time feedback loop.&lt;/p&gt;

&lt;p&gt;The strongest production architecture is usually the one that treats latency, streaming, turn-taking, state, error recovery, and speech quality as one system rather than separate model-selection problems.&lt;/p&gt;

&lt;p&gt;If you are building that loop yourself, measure every handoff.&lt;/p&gt;

&lt;p&gt;If you would rather start with an integrated stack, you can &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=how-to-build-an-ai-voice-agent-using-atoms-api"&gt;start building with the Smallest AI API&lt;/a&gt; and evaluate the pipeline with your own conversational workloads.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>voiceai</category>
      <category>speechrecognition</category>
      <category>texttospeech</category>
    </item>
    <item>
      <title>How to Build Reliable Streaming Speech-to-Text in Production</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Wed, 26 Aug 2026 07:37:35 +0000</pubDate>
      <link>https://dev.to/smallestai/how-to-build-reliable-streaming-speech-to-text-in-production-419g</link>
      <guid>https://dev.to/smallestai/how-to-build-reliable-streaming-speech-to-text-in-production-419g</guid>
      <description>&lt;p&gt;A streaming speech-to-text demo is usually the easy part.&lt;/p&gt;

&lt;p&gt;You connect a microphone, send audio frames, receive partial transcripts, and everything feels instant.&lt;/p&gt;

&lt;p&gt;Production is where things get complicated.&lt;/p&gt;

&lt;p&gt;Real users do not have perfect networks. WebSocket connections drop. Audio packets arrive late. Partial transcripts appear out of order. Reconnecting can create duplicate text. And downstream systems need a way to decide whether a transcript is trustworthy.&lt;/p&gt;

&lt;p&gt;A production-grade streaming transcription system needs to handle these failure modes intentionally.&lt;/p&gt;

&lt;p&gt;This guide covers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;How streaming speech-to-text systems process audio&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How to recover from network dropouts&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How to reconnect without corrupting transcripts&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How to remove duplicate segments&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How to evaluate transcript quality in real time&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How streaming speech-to-text works
&lt;/h2&gt;

&lt;p&gt;Streaming transcription is not a simple request-response workflow.&lt;/p&gt;

&lt;p&gt;Instead, it is a continuous bidirectional connection where:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Audio frames are sent to the recognition engine.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The engine processes the incoming stream.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Partial and final transcript segments are returned asynchronously.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most production systems use persistent connections such as WebSockets because they avoid repeated connection overhead.&lt;/p&gt;

&lt;p&gt;A typical pipeline looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Microphone
    ↓
Audio frames
    ↓
Streaming connection
    ↓
Speech recognition engine
    ↓
Interim + final transcript segments
    ↓
Application logic
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Streaming systems usually return two types of results:&lt;/p&gt;

&lt;h3&gt;
  
  
  Interim results
&lt;/h3&gt;

&lt;p&gt;Interim results are temporary predictions.&lt;/p&gt;

&lt;p&gt;They are useful for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Live captions&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Real-time interfaces&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Voice assistants&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, they can change as more audio arrives.&lt;/p&gt;

&lt;h3&gt;
  
  
  Final results
&lt;/h3&gt;

&lt;p&gt;Final results are committed transcript segments.&lt;/p&gt;

&lt;p&gt;They should be used for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Storage&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Search indexing&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Compliance workflows&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Analytics pipelines&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Treating interim results as final is one of the most common causes of unreliable transcript experiences.&lt;/p&gt;

&lt;h2&gt;
  
  
  Handling dropout events
&lt;/h2&gt;

&lt;p&gt;A dropout happens whenever the continuous audio stream is interrupted.&lt;/p&gt;

&lt;p&gt;Common causes include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;WebSocket disconnections&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Packet loss&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Server timeouts&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Microphone permission changes&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Device switching&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Mobile application backgrounding&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first thing to measure is not only whether a disconnect happened, but how long the interruption lasted.&lt;/p&gt;

&lt;p&gt;A practical approach:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Short gaps can often be recovered with buffered audio.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Medium gaps may require context rebuilding.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Longer interruptions should usually trigger a fresh session.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Build dropout-aware buffering
&lt;/h2&gt;

&lt;p&gt;A client-side audio buffer helps recover from temporary interruptions.&lt;/p&gt;

&lt;p&gt;A production implementation should:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Keep a rolling audio buffer.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Track disconnect start and end times.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Store transcript segments with timing metadata.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Emit connection-state events to the application layer.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example metadata:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"session_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"session_123"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"segment_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timestamp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1710000000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"final"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is not only reconnecting the network connection.&lt;/p&gt;

&lt;p&gt;The goal is preserving transcript continuity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reconnect logic without transcript corruption
&lt;/h2&gt;

&lt;p&gt;A common mistake is treating reconnecting as:&lt;/p&gt;

&lt;p&gt;disconnect → reconnect → continue sending audio&lt;/p&gt;

&lt;p&gt;The connection may recover, but transcript consistency may not.&lt;/p&gt;

&lt;p&gt;After reconnecting, you now have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;A previous session&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;A new session&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Potentially overlapping audio&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without tracking session boundaries, your transcript assembler cannot know whether a segment is new or duplicated.&lt;/p&gt;

&lt;p&gt;A better approach is to attach:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Session ID&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Segment sequence number&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Absolute timestamp&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then assemble transcripts using timestamps rather than arrival order.&lt;/p&gt;

&lt;p&gt;Network delays can cause older segments to arrive after newer ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  Removing duplicate transcript segments
&lt;/h2&gt;

&lt;p&gt;Duplicate text usually comes from two situations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Interim-to-final promotion
&lt;/h3&gt;

&lt;p&gt;Example:&lt;/p&gt;

&lt;p&gt;Interim:&lt;/p&gt;

&lt;p&gt;"The meeting will start"&lt;/p&gt;

&lt;p&gt;Final:&lt;/p&gt;

&lt;p&gt;"The meeting will start at three"&lt;/p&gt;

&lt;p&gt;Appending both creates:&lt;/p&gt;

&lt;p&gt;The meeting will start The meeting will start at three&lt;/p&gt;

&lt;p&gt;The solution:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Track committed final positions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Replace interim text instead of appending it.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Replayed audio after reconnect
&lt;/h3&gt;

&lt;p&gt;When buffered audio is replayed after a dropout, the speech engine may transcribe the same audio again.&lt;/p&gt;

&lt;p&gt;Exact string matching is unreliable because transcripts may differ slightly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Capitalization&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Punctuation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Minor wording changes&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A better method is token-overlap comparison.&lt;/p&gt;

&lt;p&gt;If two segments:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Have overlapping timestamps&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Share a high percentage of tokens&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Represent the same spoken content&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keep the higher-confidence version.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxxk1o2v7szw99t952dj6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxxk1o2v7szw99t952dj6.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Measuring transcript quality
&lt;/h2&gt;

&lt;p&gt;Word Error Rate (WER) is the standard metric for evaluating speech recognition accuracy.&lt;/p&gt;

&lt;p&gt;WER measures:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Substitutions&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Insertions&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Deletions&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, WER requires a reference transcript, so it is not useful during a live conversation.&lt;/p&gt;

&lt;p&gt;For real-time systems, confidence scores are more practical.&lt;/p&gt;

&lt;p&gt;A production pipeline can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Calculate segment confidence.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Set thresholds.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Route uncertain segments for review.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;High confidence
→ Display immediately

Low confidence
→ Delay or review before downstream processing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For deeper evaluation methods, developers can also explore guides on &lt;a href="https://smallest.ai/blog/how-to-evaluate-asr-in-2026?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=streaming-speech-to-text-production"&gt;evaluating ASR systems&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Handling context loss after reconnects
&lt;/h2&gt;

&lt;p&gt;A reconnect does more than restore a network connection.&lt;/p&gt;

&lt;p&gt;It can also reset:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Acoustic context&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Language model context&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Conversation history&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This matters when users discuss:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Technical terminology&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Medical terms&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Legal vocabulary&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Industry-specific language&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Passing vocabulary hints or context information when starting a new session can improve recovery quality.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building production-ready streaming STT systems
&lt;/h2&gt;

&lt;p&gt;A reliable streaming speech-to-text implementation should include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;WebSocket event monitoring&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Audio buffering&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Session tracking&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Timestamp-based ordering&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Duplicate detection&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Confidence-based routing&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Context restoration strategies&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The recognition API is only one part of the system.&lt;/p&gt;

&lt;p&gt;The application layer determines whether the final experience feels reliable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building with Smallest AI
&lt;/h2&gt;

&lt;p&gt;Developers building real-time transcription workflows can use the &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=streaming-speech-to-text-production"&gt;Smallest AI speech-to-text API&lt;/a&gt; to integrate speech recognition into production applications.&lt;/p&gt;

&lt;p&gt;For developers designing complete voice pipelines, the &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=streaming-speech-to-text-production"&gt;Smallest AI API&lt;/a&gt; provides programmatic access for building speech-based applications.&lt;/p&gt;

&lt;p&gt;You can also explore the broader &lt;a href="https://smallest.ai/blog/top-10-speech-to-text-transcription-software-picks-for-2026?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=streaming-speech-to-text-production"&gt;speech-to-text transcription software landscape&lt;/a&gt; and compare different approaches for production deployments.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;Streaming speech-to-text becomes challenging when real-world conditions appear.&lt;/p&gt;

&lt;p&gt;Networks fail. Sessions restart. Audio overlaps. Confidence varies.&lt;/p&gt;

&lt;p&gt;The difference between a demo and a production system is having the engineering discipline to handle those cases.&lt;/p&gt;

&lt;p&gt;Design around failures from the beginning, and your transcription pipeline will remain reliable even when users and networks are unpredictable.&lt;/p&gt;

&lt;p&gt;Start building your own voice workflow with the &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=streaming-speech-to-text-production"&gt;Smallest AI API&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>websockets</category>
      <category>realtime</category>
      <category>speechrecognition</category>
    </item>
    <item>
      <title>How to Build a Low-Latency Voice Bot with Streaming STT and TTS</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Wed, 26 Aug 2026 07:20:48 +0000</pubDate>
      <link>https://dev.to/smallestai/how-to-build-a-low-latency-voice-bot-with-streaming-stt-and-tts-524j</link>
      <guid>https://dev.to/smallestai/how-to-build-a-low-latency-voice-bot-with-streaming-stt-and-tts-524j</guid>
      <description>&lt;p&gt;A voice bot can be described in three steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Listen to the user.&lt;/li&gt;
&lt;li&gt;Decide what to say.&lt;/li&gt;
&lt;li&gt;Speak the response.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That sounds straightforward until you try to make the loop happen in real time.&lt;/p&gt;

&lt;p&gt;Speech recognition has latency. Your reasoning layer has latency. Speech synthesis has latency. Network transit adds more. Audio arrives continuously rather than as neat request-response messages, and users expect to interrupt, hesitate, change direction, and speak over the system.&lt;/p&gt;

&lt;p&gt;That is why building a usable voice bot is less about connecting three APIs and more about designing the entire audio pipeline around streaming, turn-taking, cancellation, and latency.&lt;/p&gt;

&lt;p&gt;This article walks through the architecture behind that pipeline, using the same STT → reasoning → TTS pattern that powers many production voice applications. &lt;a href="https://smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;Smallest AI&lt;/a&gt; provides these layers through Pulse, Electron, Lightning, and its voice-agent platform, but the architectural principles apply regardless of which components you choose.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a voice bot actually does
&lt;/h2&gt;

&lt;p&gt;At the highest level, the loop looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User audio
    ↓
Speech-to-Text
    ↓
Transcript
    ↓
LLM or deterministic logic
    ↓
Response text
    ↓
Text-to-Speech
    ↓
Audio response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The architecture is easy to understand.&lt;/p&gt;

&lt;p&gt;The difficulty is that every boundary adds delay.&lt;/p&gt;

&lt;p&gt;If you treat the system as a conventional sequence of synchronous API calls, you can easily end up with this behavior:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;record full utterance
→ wait for transcription
→ wait for complete model response
→ wait for complete speech generation
→ begin playback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every stage blocks the next.&lt;/p&gt;

&lt;p&gt;For offline processing, that may be acceptable. For a live conversation, it feels slow.&lt;/p&gt;

&lt;p&gt;A real-time voice bot should instead try to overlap work wherever possible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;audio ───────────────►
       STT ──────────►
            LLM ─────►
                 TTS ─────────► audio
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Streaming is what turns a serial pipeline into an overlapping one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose the stack before writing the orchestration layer
&lt;/h2&gt;

&lt;p&gt;Three components determine most of the behavior of a traditional voice bot:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Speech-to-text for incoming audio&lt;/li&gt;
&lt;li&gt;A reasoning or decision layer&lt;/li&gt;
&lt;li&gt;Text-to-speech for outgoing audio&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Choosing each component independently gives you flexibility, but it also creates more interfaces to manage.&lt;/p&gt;

&lt;p&gt;Authentication, connection lifecycle, audio formats, retries, versioning, latency measurement, observability, and billing can all become integration concerns.&lt;/p&gt;

&lt;p&gt;For some applications, that control is worth it. For others, a unified platform reduces the amount of orchestration code your team has to own.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pick STT for conversations, not offline transcription
&lt;/h3&gt;

&lt;p&gt;A good transcription model for prerecorded files is not automatically a good STT engine for a voice bot.&lt;/p&gt;

&lt;p&gt;For real-time use, you should care about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Streaming transcription&lt;/li&gt;
&lt;li&gt;Stable partial transcripts&lt;/li&gt;
&lt;li&gt;Endpoint detection&lt;/li&gt;
&lt;li&gt;Word timestamps when required&lt;/li&gt;
&lt;li&gt;Language coverage&lt;/li&gt;
&lt;li&gt;Performance on your actual acoustic environment&lt;/li&gt;
&lt;li&gt;Whether processing can keep up with live audio&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One useful metric is real-time factor, or RTF.&lt;/p&gt;

&lt;p&gt;If a recognizer takes one second to process one second of speech, its RTF is 1.0. For a live system, you generally want processing comfortably below that level so work does not accumulate behind the incoming audio stream.&lt;/p&gt;

&lt;p&gt;Smallest AI’s Pulse is designed for real-time and prerecorded speech recognition. For a voice bot, the important architectural capability is that transcription can happen while speech is still arriving rather than after an entire recording has been uploaded.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pick TTS for time to first audio
&lt;/h3&gt;

&lt;p&gt;For a live voice application, speech quality is only one part of TTS performance.&lt;/p&gt;

&lt;p&gt;You also need to know how quickly playback can begin.&lt;/p&gt;

&lt;p&gt;A system that eventually produces excellent audio but leaves a long silence before the first syllable still feels broken in conversation.&lt;/p&gt;

&lt;p&gt;Streaming TTS addresses this by returning audio incrementally instead of waiting for synthesis of the entire response.&lt;/p&gt;

&lt;p&gt;Smallest AI’s Lightning is designed around this streaming model and can be used as the speech-generation layer of a real-time pipeline.&lt;/p&gt;

&lt;p&gt;Other practical considerations include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Voice consistency&lt;/li&gt;
&lt;li&gt;Language support&lt;/li&gt;
&lt;li&gt;Audio encoding&lt;/li&gt;
&lt;li&gt;Streaming behavior&lt;/li&gt;
&lt;li&gt;Prosody and pacing&lt;/li&gt;
&lt;li&gt;Voice cloning when your application requires a consistent custom voice&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Decide whether you actually need an LLM
&lt;/h3&gt;

&lt;p&gt;Not every voice bot needs a general-purpose language model.&lt;/p&gt;

&lt;p&gt;Consider a phone router that only needs to recognize requests such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;billing
technical support
cancel subscription
check order
speak to an agent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A deterministic intent layer may be faster, cheaper, and easier to audit than a large model.&lt;/p&gt;

&lt;p&gt;An LLM becomes more useful when users can phrase requests unpredictably or when the system needs multi-turn reasoning, tool calls, or context-sensitive responses.&lt;/p&gt;

&lt;p&gt;Smallest AI’s Electron can act as the conversational reasoning layer in this architecture. You can also use another compatible reasoning system if your application requires it.&lt;/p&gt;

&lt;p&gt;If you do not want to assemble the orchestration and agent infrastructure yourself, the &lt;a href="https://smallest.ai/voice-agents?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;Smallest AI voice-agent platform&lt;/a&gt; provides a managed path for building and deploying voice agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  Architecture blueprint: make every stage stream
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F52a54qysb66zma6wj3g5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F52a54qysb66zma6wj3g5.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
A practical voice bot architecture looks roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                      ┌─────────────────┐
Microphone / Call ───►│ Streaming STT   │
                      └────────┬────────┘
                               │
                        partial + final
                          transcripts
                               │
                               ▼
                      ┌─────────────────┐
                      │ Orchestrator    │
                      │ VAD / state /   │
                      │ cancellation    │
                      └────────┬────────┘
                               │
                               ▼
                      ┌─────────────────┐
                      │ LLM or rules    │
                      └────────┬────────┘
                               │
                         streamed text
                               │
                               ▼
                      ┌─────────────────┐
                      │ Streaming TTS   │
                      └────────┬────────┘
                               │
                               ▼
                         Audio playback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;WebSockets are a natural transport for this kind of architecture because the connection stays open while audio and events move continuously.&lt;/p&gt;

&lt;p&gt;The important idea is not WebSockets specifically. It is avoiding a workflow where every stage waits for a complete payload from the previous one.&lt;/p&gt;

&lt;p&gt;For a deeper breakdown of how the STT, LLM, TTS, tools, and latency budget interact, Smallest AI’s guide to &lt;a href="https://smallest.ai/blog/designing-voice-assistants-stt-llm-tts-tools-and-latency-budget?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;designing voice assistants around STT, LLMs, TTS, and latency&lt;/a&gt; is a useful companion to this implementation view.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create and store the API key
&lt;/h2&gt;

&lt;p&gt;If you are prototyping the pipeline with the &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;Smallest AI API&lt;/a&gt;, keep the API key in an environment variable rather than hard-coding it into the application.&lt;/p&gt;

&lt;p&gt;Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the &lt;code&gt;SMALLEST_API_KEY&lt;/code&gt; environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key-here"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every authenticated request sends the value through the &lt;code&gt;Authorization&lt;/code&gt; header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer &amp;lt;SMALLEST_API_KEY value&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the key on your server.&lt;/p&gt;

&lt;p&gt;Do not expose it in browser JavaScript, mobile application code, public repositories, screenshots, query parameters, or client-side logs.&lt;/p&gt;

&lt;p&gt;For production deployments, put the value in the secret-management system used by your infrastructure rather than a checked-in configuration file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building the core voice loop
&lt;/h2&gt;

&lt;p&gt;The easiest way to reason about the application is as several concurrent tasks rather than one long function.&lt;/p&gt;

&lt;p&gt;You typically have independent flows for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;microphone → STT
STT events → conversation state
conversation state → reasoning
reasoning output → TTS
TTS audio → playback
user interruption → cancellation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An asynchronous runtime such as Python’s &lt;code&gt;asyncio&lt;/code&gt;, Node.js, or another event-driven environment maps naturally to this model.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Capture and stream incoming audio
&lt;/h3&gt;

&lt;p&gt;Start with the audio source.&lt;/p&gt;

&lt;p&gt;Depending on the application, that may be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A browser microphone&lt;/li&gt;
&lt;li&gt;A native mobile microphone&lt;/li&gt;
&lt;li&gt;A WebRTC session&lt;/li&gt;
&lt;li&gt;A telephony stream&lt;/li&gt;
&lt;li&gt;A SIP connection&lt;/li&gt;
&lt;li&gt;Another realtime media transport&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not accumulate several seconds of speech before beginning transcription.&lt;/p&gt;

&lt;p&gt;Send audio frames to STT continuously.&lt;/p&gt;

&lt;p&gt;Frame sizes in the tens-of-milliseconds range are common because they balance responsiveness against packet and processing overhead. A 16 kHz stream is also a common baseline for speech-recognition workloads, although your actual settings should match the STT service and source audio you are using.&lt;/p&gt;

&lt;p&gt;The recognizer can then begin returning partial transcription results while the person is still speaking.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Handle VAD and endpointing separately
&lt;/h3&gt;

&lt;p&gt;Recognizing words is only part of the problem.&lt;/p&gt;

&lt;p&gt;The bot also has to determine when the user has finished a turn.&lt;/p&gt;

&lt;p&gt;This is where voice activity detection and endpointing become critical.&lt;/p&gt;

&lt;p&gt;Consider someone saying:&lt;/p&gt;

&lt;p&gt;“Can you… uh… move my appointment to Friday?”&lt;/p&gt;

&lt;p&gt;A simplistic silence timer may decide the utterance ended after “Can you.”&lt;/p&gt;

&lt;p&gt;Set the threshold too aggressively and the bot interrupts natural pauses.&lt;/p&gt;

&lt;p&gt;Set it too conservatively and every turn gains an uncomfortable delay.&lt;/p&gt;

&lt;p&gt;The correct value depends heavily on the environment and task.&lt;/p&gt;

&lt;p&gt;A tightly scripted support call may tolerate aggressive turn-taking. A conversational assistant whose users frequently pause to think may need more patience.&lt;/p&gt;

&lt;p&gt;Do not tune endpointing exclusively with clean test recordings. Use audio that represents actual microphones, background noise, accents, network conditions, and speaking patterns from your deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not wait for the entire model response
&lt;/h2&gt;

&lt;p&gt;After STT produces a finalized user turn, the transcript enters the reasoning layer.&lt;/p&gt;

&lt;p&gt;A common first implementation looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;complete transcript
        ↓
generate entire LLM response
        ↓
send complete response to TTS
        ↓
play audio
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It works, but it wastes time.&lt;/p&gt;

&lt;p&gt;If your reasoning system supports streaming output, begin preparing speech before the entire answer exists.&lt;/p&gt;

&lt;p&gt;You do not necessarily want to synthesize individual tokens. TTS needs enough context to produce natural phrasing and prosody.&lt;/p&gt;

&lt;p&gt;A better strategy is to buffer the model output into speakable chunks such as complete clauses or sentences:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LLM tokens
   ↓
small text buffer
   ↓
natural speech boundary
   ↓
TTS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;While TTS is synthesizing the first chunk, the model can continue generating the next.&lt;/p&gt;

&lt;p&gt;That overlap is one of the most important ways to reduce perceived response time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stream synthesized audio back immediately
&lt;/h2&gt;

&lt;p&gt;The same principle applies after text enters the TTS layer.&lt;/p&gt;

&lt;p&gt;Do not wait for the complete audio file if your provider supports streaming.&lt;/p&gt;

&lt;p&gt;Start playback as soon as enough audio has arrived.&lt;/p&gt;

&lt;p&gt;The end-to-end flow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User still speaking
       │
       ├── STT processing audio
       │
User finishes
       │
       ├── final transcript
       │
       ├── reasoning begins
       │
       ├── first usable text chunk
       │
       ├── TTS begins
       │
       └── first audio begins playing
              while later text
              and audio are still
              being generated
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For network delivery, codecs designed for interactive audio are usually better suited to real-time conversations than formats optimized primarily for downloadable media.&lt;/p&gt;

&lt;p&gt;Opus is commonly used in real-time communication because it performs well at relatively low bitrates and is designed for interactive audio. Raw PCM is convenient when debugging or when downstream systems need uncompressed samples.&lt;/p&gt;

&lt;p&gt;Whatever format you choose, avoid unnecessary encode/decode conversions between services. Every conversion creates another place for buffering, CPU work, or format mismatches.&lt;/p&gt;

&lt;h2&gt;
  
  
  Barge-in changes the architecture
&lt;/h2&gt;

&lt;p&gt;A voice bot that cannot be interrupted will quickly feel artificial.&lt;/p&gt;

&lt;p&gt;Suppose the bot starts saying:&lt;/p&gt;

&lt;p&gt;“Your current subscription includes—”&lt;/p&gt;

&lt;p&gt;and the user interrupts:&lt;/p&gt;

&lt;p&gt;“Actually, I want to cancel it.”&lt;/p&gt;

&lt;p&gt;The system should stop the outgoing response and process the new utterance.&lt;/p&gt;

&lt;p&gt;That means several things have to happen almost simultaneously:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;detect user speech
→ stop or cancel TTS generation
→ flush queued playback audio
→ preserve the new incoming speech
→ return control to the listening state
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important consequence is that your STT path cannot simply disappear while the bot is speaking.&lt;/p&gt;

&lt;p&gt;You need to continue monitoring incoming audio so the system can detect the interruption.&lt;/p&gt;

&lt;p&gt;This is also why echo cancellation matters. Without it, the microphone may feed the bot’s own synthesized speech back into STT, and the system can mistake its response for a new user utterance.&lt;/p&gt;

&lt;p&gt;Barge-in is not a small feature you bolt onto the end of development. It affects the audio pipeline, state machine, playback buffer, VAD behavior, and cancellation strategy.&lt;/p&gt;

&lt;p&gt;Design for it early.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat latency as a pipeline budget
&lt;/h2&gt;

&lt;p&gt;Developers often optimize whichever component has the most obvious latency number.&lt;/p&gt;

&lt;p&gt;That is not enough.&lt;/p&gt;

&lt;p&gt;A conversational delay is the sum of several pieces:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;endpoint detection
+ STT
+ network
+ reasoning
+ tool calls
+ TTS startup
+ playback buffering
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Improving one component may not noticeably change the experience if another stage dominates the total.&lt;/p&gt;

&lt;p&gt;Instrument the boundaries individually.&lt;/p&gt;

&lt;p&gt;For each turn, record timestamps such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;speech_started
speech_ended
final_transcript_received
llm_started
first_llm_chunk_received
tts_started
first_audio_received
playback_started
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then measure distributions rather than only averages.&lt;/p&gt;

&lt;p&gt;P50 tells you what a typical user experiences.&lt;/p&gt;

&lt;p&gt;P90 and P99 reveal the slow turns that users are more likely to remember as failures.&lt;/p&gt;

&lt;p&gt;A related metric is task completion. A low-latency system that misunderstands users or fails to complete the requested workflow is not successful simply because it responds quickly.&lt;/p&gt;

&lt;p&gt;Latency and task success have to be evaluated together.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four mistakes that break voice bots outside the demo
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Building a chatbot with a microphone attached
&lt;/h3&gt;

&lt;p&gt;Voice is not just a different input widget for a text application.&lt;/p&gt;

&lt;p&gt;You now have to handle:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Silence&lt;/li&gt;
&lt;li&gt;Natural pauses&lt;/li&gt;
&lt;li&gt;Background noise&lt;/li&gt;
&lt;li&gt;Echo&lt;/li&gt;
&lt;li&gt;Interruptions&lt;/li&gt;
&lt;li&gt;Audio encoding&lt;/li&gt;
&lt;li&gt;Turn detection&lt;/li&gt;
&lt;li&gt;Playback state&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ignoring these problems is why many prototypes work perfectly at a developer’s desk and fall apart during real calls.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Batching every stage
&lt;/h3&gt;

&lt;p&gt;If the system waits for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;full audio
→ full transcript
→ full model response
→ full TTS file
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;all of the delays are serialized.&lt;/p&gt;

&lt;p&gt;Stream wherever the component supports it.&lt;/p&gt;

&lt;p&gt;The goal is not merely faster APIs. It is overlapping stages.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Load testing HTTP requests instead of conversations
&lt;/h3&gt;

&lt;p&gt;Fifty text API requests are not equivalent to fifty simultaneous calls.&lt;/p&gt;

&lt;p&gt;Each live conversation can involve:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A persistent connection&lt;/li&gt;
&lt;li&gt;Continuous audio&lt;/li&gt;
&lt;li&gt;STT state&lt;/li&gt;
&lt;li&gt;VAD state&lt;/li&gt;
&lt;li&gt;Conversation history&lt;/li&gt;
&lt;li&gt;TTS generation&lt;/li&gt;
&lt;li&gt;Audio playback&lt;/li&gt;
&lt;li&gt;Cancellation events&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Load test the architecture using realistic concurrent audio sessions.&lt;/p&gt;

&lt;p&gt;Watch memory, open connections, queue depth, model latency, reconnection behavior, and tail latency.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Treating one voice as correct for every workflow
&lt;/h3&gt;

&lt;p&gt;The voice itself is part of the interface.&lt;/p&gt;

&lt;p&gt;A scheduling assistant, support agent, sales workflow, and collections system may have very different tone requirements.&lt;/p&gt;

&lt;p&gt;If your application needs a consistent custom voice, voice cloning is one option, but it should be evaluated alongside latency, language coverage, intelligibility, and the context in which the bot will speak.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production features that the MVP usually skips
&lt;/h2&gt;

&lt;p&gt;Getting audio through STT, a model, and TTS proves the basic pipeline.&lt;/p&gt;

&lt;p&gt;It does not make the bot production-ready.&lt;/p&gt;

&lt;h3&gt;
  
  
  Multi-turn conversation state
&lt;/h3&gt;

&lt;p&gt;A useful voice bot has to remember what happened earlier in the conversation.&lt;/p&gt;

&lt;p&gt;For shorter sessions, you can carry recent turns into the model context.&lt;/p&gt;

&lt;p&gt;Longer interactions require a strategy to keep context from growing indefinitely.&lt;/p&gt;

&lt;p&gt;Common approaches include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sliding window
summarization
retrieval of relevant past context
structured application state
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not rely entirely on a raw transcript when the application has business state such as an appointment date, order ID, account selection, or workflow stage.&lt;/p&gt;

&lt;p&gt;Store important state explicitly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Telephony integration
&lt;/h3&gt;

&lt;p&gt;If the bot will handle phone calls, speech generation is only part of the system.&lt;/p&gt;

&lt;p&gt;You eventually encounter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;SIP/PSTN connectivity&lt;/li&gt;
&lt;li&gt;Phone-number management&lt;/li&gt;
&lt;li&gt;DTMF&lt;/li&gt;
&lt;li&gt;Transfers&lt;/li&gt;
&lt;li&gt;Hold behavior&lt;/li&gt;
&lt;li&gt;Recording&lt;/li&gt;
&lt;li&gt;Regional routing&lt;/li&gt;
&lt;li&gt;Compliance requirements&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At this point, teams must decide whether they want to own the telephony and orchestration infrastructure themselves or use an agent platform that abstracts some of it.&lt;/p&gt;

&lt;p&gt;Smallest AI’s managed voice-agent platform is one option when you want the speech stack and agent infrastructure integrated rather than assembled independently.&lt;/p&gt;

&lt;h3&gt;
  
  
  Observability
&lt;/h3&gt;

&lt;p&gt;Log enough information to reconstruct why a bad conversation happened.&lt;/p&gt;

&lt;p&gt;Useful measurements include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;STT latency&lt;/li&gt;
&lt;li&gt;Endpointing delay&lt;/li&gt;
&lt;li&gt;LLM time to first output&lt;/li&gt;
&lt;li&gt;Tool-call latency&lt;/li&gt;
&lt;li&gt;TTS startup latency&lt;/li&gt;
&lt;li&gt;End-to-end time to first audio&lt;/li&gt;
&lt;li&gt;Interruption frequency&lt;/li&gt;
&lt;li&gt;Connection failures&lt;/li&gt;
&lt;li&gt;Task-completion rate&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Be thoughtful about what you store.&lt;/p&gt;

&lt;p&gt;Voice interactions can contain sensitive information, and observability should not become an excuse to dump raw transcripts, credentials, or unnecessary user data into logs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should you build the stack or use a managed platform?
&lt;/h2&gt;

&lt;p&gt;There is no universal answer.&lt;/p&gt;

&lt;p&gt;Building the pipeline yourself gives you direct control over:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;audio transport
STT
reasoning
tool execution
TTS
state
telephony
observability
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That control can be valuable if voice is a core part of your product infrastructure.&lt;/p&gt;

&lt;p&gt;But you also inherit every integration boundary.&lt;/p&gt;

&lt;p&gt;A managed platform moves more of that responsibility into the provider.&lt;/p&gt;

&lt;p&gt;Smallest AI’s stack combines Pulse for STT, Electron for conversational reasoning, Lightning for TTS, and its voice-agent tooling for orchestration and deployment.&lt;/p&gt;

&lt;p&gt;The important engineering question is not whether a unified or modular architecture is always better.&lt;/p&gt;

&lt;p&gt;It is how much of the real-time voice infrastructure your team wants to build, operate, monitor, and debug itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical build checklist
&lt;/h2&gt;

&lt;p&gt;Before calling the voice bot production-ready, verify that you have made explicit decisions about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Streaming versus batching at every stage&lt;/li&gt;
&lt;li&gt;STT endpointing and VAD thresholds&lt;/li&gt;
&lt;li&gt;Audio sample rate and encoding&lt;/li&gt;
&lt;li&gt;LLM versus deterministic routing&lt;/li&gt;
&lt;li&gt;TTS chunk boundaries&lt;/li&gt;
&lt;li&gt;Barge-in and cancellation&lt;/li&gt;
&lt;li&gt;Echo handling&lt;/li&gt;
&lt;li&gt;Conversation-state management&lt;/li&gt;
&lt;li&gt;Tool-call latency&lt;/li&gt;
&lt;li&gt;Reconnection behavior&lt;/li&gt;
&lt;li&gt;Concurrent session limits&lt;/li&gt;
&lt;li&gt;P50, P90, and P99 latency measurement&lt;/li&gt;
&lt;li&gt;Secret management&lt;/li&gt;
&lt;li&gt;Sensitive transcript logging&lt;/li&gt;
&lt;li&gt;Real-world audio testing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most production failures happen because one of these areas was treated as an implementation detail instead of an architectural decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real optimization is overlap
&lt;/h2&gt;

&lt;p&gt;The basic voice-bot loop has not changed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;listen → think → speak
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What separates a responsive voice application from a sluggish one is how much work happens concurrently.&lt;/p&gt;

&lt;p&gt;Stream audio into STT instead of waiting for a recording.&lt;/p&gt;

&lt;p&gt;Start reasoning as soon as the user’s turn is complete.&lt;/p&gt;

&lt;p&gt;Stream model output toward TTS instead of waiting for the full response.&lt;/p&gt;

&lt;p&gt;Start playback as soon as audio is available.&lt;/p&gt;

&lt;p&gt;Keep listening while the bot speaks so users can interrupt.&lt;/p&gt;

&lt;p&gt;Measure every boundary instead of treating latency as one opaque number.&lt;/p&gt;

&lt;p&gt;Once those pieces are in place, the technology choices become easier to evaluate because you can judge them inside the architecture that will actually run in production.&lt;/p&gt;

&lt;p&gt;If you want to prototype the pipeline using Pulse, Electron, and Lightning, &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;start building with the Smallest AI API&lt;/a&gt; and test it against your own audio, latency targets, and conversation patterns.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Which programming language should I use for a voice bot?
&lt;/h3&gt;

&lt;p&gt;Python is a practical choice when you are already working with AI APIs and asynchronous processing. Node.js is also a strong option for browser, WebSocket, and WebRTC-heavy backends.&lt;/p&gt;

&lt;p&gt;For high-concurrency deployments, other languages may make sense depending on your infrastructure.&lt;/p&gt;

&lt;p&gt;The best choice is usually the runtime your team can operate reliably while handling persistent connections and concurrent audio streams.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I reduce voice-bot latency?
&lt;/h3&gt;

&lt;p&gt;Start with architecture rather than micro-optimizations.&lt;/p&gt;

&lt;p&gt;Stream STT, reasoning output, and TTS. Tune endpointing carefully. Avoid unnecessary audio transformations. Reduce network hops where possible. Measure time to first output for each component, and inspect tail latency rather than relying only on averages.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do I need an LLM for every voice bot?
&lt;/h3&gt;

&lt;p&gt;No.&lt;/p&gt;

&lt;p&gt;A deterministic intent or rules layer can be the better choice for narrow, predictable workflows.&lt;/p&gt;

&lt;p&gt;Use an LLM when the application needs flexible language understanding, multi-turn reasoning, dynamic responses, or tool selection.&lt;/p&gt;

&lt;p&gt;Many production systems use a combination of deterministic logic and models.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between a voice bot and a voice assistant?
&lt;/h3&gt;

&lt;p&gt;The terms overlap.&lt;/p&gt;

&lt;p&gt;A voice bot usually describes a more task-specific system: customer support, scheduling, routing, qualification, or another bounded workflow.&lt;/p&gt;

&lt;p&gt;A voice assistant often implies broader, more open-ended interaction.&lt;/p&gt;

&lt;p&gt;Both can use the same underlying STT → reasoning → TTS architecture. The main difference is usually the scope of the jobs they are expected to handle.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>speechrecognition</category>
      <category>texttospeech</category>
      <category>websockets</category>
    </item>
    <item>
      <title>How to Build a Production-Ready Audio Transcription Pipeline in Python</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Wed, 26 Aug 2026 07:01:53 +0000</pubDate>
      <link>https://dev.to/smallestai/how-to-build-a-production-ready-audio-transcription-pipeline-in-python-2of9</link>
      <guid>https://dev.to/smallestai/how-to-build-a-production-ready-audio-transcription-pipeline-in-python-2of9</guid>
      <description>&lt;p&gt;Transcribing an audio file from Python looks simple in a demo:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open a file.&lt;/li&gt;
&lt;li&gt;Send it to an API.&lt;/li&gt;
&lt;li&gt;Print the returned text.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then real audio arrives.&lt;/p&gt;

&lt;p&gt;A user uploads an MP3 instead of a WAV. A phone recording is narrowband. Two speakers interrupt each other. A 90-minute recording times out halfway through. The network returns a transient error. Your application needs timestamps rather than one giant string.&lt;/p&gt;

&lt;p&gt;That is where transcription stops being an API-call problem and becomes a pipeline problem.&lt;/p&gt;

&lt;p&gt;This guide builds that pipeline step by step. We will preprocess audio, send pre-recorded files through a speech-to-text API, handle hosted audio, work with structured transcription data, add speaker diarization, split long recordings, and introduce safer retry patterns.&lt;/p&gt;

&lt;p&gt;For the implementation examples, we’ll use &lt;a href="https://smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest AI&lt;/a&gt; and its &lt;a href="https://smallest.ai/speech-to-text?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Pulse speech-to-text models&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the full transcription pipeline
&lt;/h2&gt;

&lt;p&gt;A production transcription workflow is more than “audio in, text out.”&lt;/p&gt;

&lt;p&gt;A useful mental model is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Receive or locate the audio.&lt;/li&gt;
&lt;li&gt;Inspect and normalize it when necessary.&lt;/li&gt;
&lt;li&gt;Choose the transcription mode and model.&lt;/li&gt;
&lt;li&gt;Authenticate the request securely.&lt;/li&gt;
&lt;li&gt;Send the audio bytes or hosted URL.&lt;/li&gt;
&lt;li&gt;Receive structured JSON.&lt;/li&gt;
&lt;li&gt;Extract transcript, timestamps, speakers, and metadata.&lt;/li&gt;
&lt;li&gt;Store or post-process the result.&lt;/li&gt;
&lt;li&gt;Handle failures, retries, duplicates, and long-running jobs.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That last part matters.&lt;/p&gt;

&lt;p&gt;The actual HTTP request may only occupy a few lines of Python. Most of the engineering work happens around it.&lt;/p&gt;

&lt;p&gt;A transcript might eventually power search, captions, meeting notes, call QA, downstream automation, analytics, or a voice application. Once that happens, timestamps, speaker boundaries, error handling, and repeatable input formats become as important as the transcript string itself.&lt;/p&gt;

&lt;p&gt;For a more API-focused look at recorded audio ingestion, the &lt;a href="https://smallest.ai/blog/audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;programmatic audio-to-text API workflow&lt;/a&gt; covers the broader recorded-audio pattern.&lt;/p&gt;

&lt;h2&gt;
  
  
  What “transcribe” means at the API boundary
&lt;/h2&gt;

&lt;p&gt;At the application layer, your responsibility is usually straightforward: provide valid audio and enough configuration for the transcription service to interpret it correctly.&lt;/p&gt;

&lt;p&gt;Behind that API boundary, an ASR system has to map an acoustic signal into linguistic units, decode those units into likely text, and produce a usable output representation.&lt;/p&gt;

&lt;p&gt;Depending on the system and configuration, that output can include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the complete transcript&lt;/li&gt;
&lt;li&gt;word-level timestamps&lt;/li&gt;
&lt;li&gt;utterance boundaries&lt;/li&gt;
&lt;li&gt;speaker labels&lt;/li&gt;
&lt;li&gt;language information&lt;/li&gt;
&lt;li&gt;confidence information&lt;/li&gt;
&lt;li&gt;processing metadata&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The structured fields are what make transcription useful to software.&lt;/p&gt;

&lt;p&gt;If you are building captions, for example, you need timing. If you are processing customer calls, you may need speaker labels. If transcripts trigger downstream actions, you may want to inspect confidence or other quality signals before trusting every token automatically.&lt;/p&gt;

&lt;p&gt;This also explains why testing only on clean demo audio is risky. Acoustic conditions, microphones, codecs, accents, background noise, overlapping speech, and domain-specific terminology can all change what the model receives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Set up the Python environment
&lt;/h2&gt;

&lt;p&gt;Start with an isolated virtual environment:&lt;/p&gt;

&lt;p&gt;(Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; venv stt-env

&lt;span class="c"&gt;# macOS / Linux&lt;/span&gt;
&lt;span class="nb"&gt;source &lt;/span&gt;stt-env/bin/activate

&lt;span class="c"&gt;# Windows&lt;/span&gt;
stt-env&lt;span class="se"&gt;\S&lt;/span&gt;cripts&lt;span class="se"&gt;\a&lt;/span&gt;ctivate

pip &lt;span class="nb"&gt;install &lt;/span&gt;requests python-dotenv pydub tenacity
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;requests handles the HTTP calls, python-dotenv is useful during local development, pydub handles common audio preprocessing tasks, and tenacity gives us controlled retry behavior later.&lt;/p&gt;

&lt;p&gt;If you use pydub with compressed formats such as MP3, make sure FFmpeg is installed on the machine running the application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Normalize unpredictable audio before transcription
&lt;/h2&gt;

&lt;p&gt;Tutorial audio is usually clean.&lt;/p&gt;

&lt;p&gt;Production audio rarely is.&lt;/p&gt;

&lt;p&gt;You may receive MP3, WAV, compressed call recordings, extracted video audio, stereo conversations, or recordings with inconsistent sample rates.&lt;/p&gt;

&lt;p&gt;For a predictable baseline, converting incoming files to mono, 16 kHz WAV is useful. Smallest AI’s current Pulse documentation recommends a 16 kHz sample rate, and converting to a known format also removes one variable when you debug failures.&lt;/p&gt;

&lt;p&gt;Do not interpret resampling as a way to recreate information that was never captured. Converting an 8 kHz telephone recording to 16 kHz does not restore frequencies lost during recording. The point is predictable input, not magic audio repair.&lt;/p&gt;

&lt;p&gt;Here is a small preprocessing function:&lt;/p&gt;

&lt;p&gt;(Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydub&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AudioSegment&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;preprocess_audio&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;input_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;output_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Convert an audio file to mono, 16 kHz, 16-bit PCM WAV.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;source&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;input_path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;is_file&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;FileNotFoundError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Audio file not found: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;audio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AudioSegment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;audio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_channels&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;audio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_frame_rate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;16_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;audio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_sample_width&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;export&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wav&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;output_path&lt;/span&gt;


&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;preprocess_audio&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input_audio.mp3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;preprocessed_audio.wav&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is local preprocessing only, so no API credentials are involved.&lt;/p&gt;

&lt;p&gt;You also should not preprocess blindly.&lt;/p&gt;

&lt;p&gt;If your input is already in a supported, appropriate format, another lossy conversion can do more harm than good. Inspect the audio first and normalize when your pipeline actually needs it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose the right transcription mode
&lt;/h2&gt;

&lt;p&gt;Smallest AI currently exposes Pulse and Pulse Pro through the same pre-recorded speech-to-text endpoint.&lt;/p&gt;

&lt;p&gt;The important distinction for this workflow is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;pulse-pro is intended for pre-recorded English transcription.&lt;/li&gt;
&lt;li&gt;pulse supports multilingual transcription and is also used when you need capabilities such as audio-by-URL, streaming, or speaker diarization.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both use the unified pre-recorded endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://api.smallest.ai/waves/v1/stt/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model is selected through the model query parameter.&lt;/p&gt;

&lt;p&gt;For developers implementing the pipeline, the &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest AI speech-to-text API&lt;/a&gt; provides the programmatic entry point used by these examples.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create and store the API key
&lt;/h2&gt;

&lt;p&gt;Keep the API key in an environment variable rather than hard-coding it into the application.&lt;/p&gt;

&lt;p&gt;Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key-here"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every authenticated request sends the value through the Authorization header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer &amp;lt;SMALLEST_API_KEY value&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the key on your server. Do not expose it in browser JavaScript, mobile application code, public repositories, screenshots, or client-side logs.&lt;/p&gt;

&lt;p&gt;For production deployments, store it in a server-side secrets manager provided by your cloud or infrastructure platform rather than committing a .env file to the repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the first transcription request from Python
&lt;/h2&gt;

&lt;p&gt;For a pre-recorded English file, we can send the raw file bytes using Pulse Pro.&lt;/p&gt;

&lt;p&gt;The request uses application/octet-stream and asks for word timestamps so the response can carry more structure than plain text.&lt;/p&gt;

&lt;p&gt;Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;


&lt;span class="n"&gt;ENDPOINT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.smallest.ai/waves/v1/stt/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;transcribe_audio&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Transcribe a pre-recorded English audio file with Pulse Pro.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;audio_path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;audio_path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;is_file&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;FileNotFoundError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Audio file not found: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;audio_path&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;api_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="n"&gt;params&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pulse-pro&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;language&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;en&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;word_timestamps&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;headers&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/octet-stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;audio_path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;audio_file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;ENDPOINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;audio_file&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;120&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;


&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;transcribe_audio&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;preprocessed_audio.wav&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;transcription&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;raise_for_status() deserves to stay in even the smallest example.&lt;/p&gt;

&lt;p&gt;Without it, your code can accidentally treat an authentication failure, rate limit, or server error as if it were a successful transcription with missing data.&lt;/p&gt;

&lt;p&gt;Word timestamps also become useful surprisingly quickly. They let you build searchable audio, synchronize captions, highlight matching sections of a recording, or connect transcript spans back to the original media.&lt;/p&gt;

&lt;h2&gt;
  
  
  Transcribe audio from a hosted URL
&lt;/h2&gt;

&lt;p&gt;Sometimes the file does not live on the transcription worker.&lt;/p&gt;

&lt;p&gt;It may already be stored in object storage behind a public or signed URL.&lt;/p&gt;

&lt;p&gt;For URL-based pre-recorded transcription, use Pulse and send a JSON payload containing the URL.&lt;/p&gt;

&lt;p&gt;The URL must be reachable by the transcription service. If you use private object storage, prefer a short-lived signed URL rather than making the object permanently public.&lt;/p&gt;

&lt;p&gt;Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;

&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;


&lt;span class="n"&gt;ENDPOINT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.smallest.ai/waves/v1/stt/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;transcribe_audio_url&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio_url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Transcribe hosted audio using the Pulse model.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="n"&gt;params&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pulse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;language&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;en&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;word_timestamps&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;headers&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;audio_url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;ENDPOINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;120&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice that the model choice changed.&lt;/p&gt;

&lt;p&gt;Pulse Pro accepts raw pre-recorded audio, while URL input belongs to the Pulse workflow. Treat model selection as part of the request contract rather than a cosmetic option.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat the response as structured data
&lt;/h2&gt;

&lt;p&gt;Do not throw away everything except the transcript string.&lt;/p&gt;

&lt;p&gt;For a response with word timestamps enabled, you may have timing and confidence information that can be valuable elsewhere in the product.&lt;/p&gt;

&lt;p&gt;A defensive parser should also assume optional fields can be absent:&lt;/p&gt;

&lt;p&gt;(Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;parse_transcript&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Print transcript text and available word metadata.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;full_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;transcription&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Transcript: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;full_text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;word&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;words&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]):&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;word&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;word&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;word&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;start&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;word&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;end&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;word&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;start_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;isinstance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;end_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;isinstance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;confidence_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;confidence&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;isinstance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;confidence&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;n/a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;start_text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; - &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;end_text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;] &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;(confidence: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;confidence_text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;language&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;language&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unknown&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Language: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Confidence values are better treated as signals than universal truth.&lt;/p&gt;

&lt;p&gt;Avoid assuming that a single threshold such as 0.7 works for every model, language, microphone, or domain. If confidence will route transcripts to human review, calibrate that threshold against your own labeled audio.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add speaker diarization when “who said it” matters
&lt;/h2&gt;

&lt;p&gt;For meetings, interviews, podcasts, and customer calls, plain transcription may not be enough.&lt;/p&gt;

&lt;p&gt;You also need to know which speaker produced each segment.&lt;/p&gt;

&lt;p&gt;That is speaker diarization.&lt;/p&gt;

&lt;p&gt;With the current Pulse pre-recorded API, diarization is enabled by using the Pulse model and passing:&lt;/p&gt;

&lt;p&gt;model=pulse, language=en, and diarize=true.&lt;/p&gt;

&lt;p&gt;You can combine diarization with word timestamps when you need both timing and speaker structure.&lt;/p&gt;

&lt;p&gt;The response can contain speaker information at the word and utterance levels. A local formatter can then rebuild a readable conversation:&lt;/p&gt;

&lt;p&gt;(Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;format_diarized_transcript&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Convert diarized utterances into readable speaker turns.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;lines&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;utterance&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utterances&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]):&lt;/span&gt;
        &lt;span class="n"&gt;speaker&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;utterance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;speaker&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unknown_speaker&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;utterance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;utterance&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;start&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;continue&lt;/span&gt;

        &lt;span class="n"&gt;start_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;isinstance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;lines&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;start_text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;] &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;speaker&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lines&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Diarization introduces its own failure modes.&lt;/p&gt;

&lt;p&gt;Overlapping speech is particularly difficult because two voices can occupy the same time interval. Recording separate channels upstream, when your telephony or conferencing stack makes that possible, can simplify later processing.&lt;/p&gt;

&lt;p&gt;Also remember that diarization is not the same as identity recognition. A label such as speaker_0 tells you that the segment belongs to one detected speaker; it does not automatically tell you that the person is “Alice.”&lt;/p&gt;

&lt;p&gt;If multi-speaker transcription is central to your application, the &lt;a href="https://smallest.ai/blog/the-complete-speaker-diarization-api-guide-how-it-works-and-best-practices?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;speaker diarization API guide&lt;/a&gt; goes deeper into those tradeoffs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Handle long recordings deliberately
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb08og9xvakpt7nz3p4so.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb08og9xvakpt7nz3p4so.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
Long audio creates a different reliability problem.&lt;/p&gt;

&lt;p&gt;A single large request gives you a large failure domain. If the request times out late in processing, you may have to repeat substantial work.&lt;/p&gt;

&lt;p&gt;There are two useful approaches.&lt;/p&gt;

&lt;p&gt;For long Pulse Pro transcription, an asynchronous webhook workflow can avoid holding one HTTP connection open for the entire job.&lt;/p&gt;

&lt;p&gt;Application-level chunking is another option when you want smaller independent work units.&lt;/p&gt;

&lt;p&gt;A practical chunking strategy is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;split the recording into manageable segments&lt;/li&gt;
&lt;li&gt;include a small overlap between adjacent segments&lt;/li&gt;
&lt;li&gt;transcribe segments independently&lt;/li&gt;
&lt;li&gt;preserve each segment’s original time offset&lt;/li&gt;
&lt;li&gt;reconcile duplicate words created by the overlap&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here is the local splitting step:&lt;/p&gt;

&lt;p&gt;(Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydub&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AudioSegment&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;split_audio_with_overlap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;input_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;output_dir&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;chunk_length_ms&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;60_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;overlap_ms&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Split audio into overlapping WAV chunks.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;audio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AudioSegment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;input_path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output_dir&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mkdir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exist_ok&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;chunk_paths&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="n"&gt;chunk_index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;chunk_length_ms&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

        &lt;span class="n"&gt;chunk_path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chunk_&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;chunk_index&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;04&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.wav&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;export&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wav&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;chunk_paths&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk_path&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;break&lt;/span&gt;

        &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;overlap_ms&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;chunk_index&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chunk_paths&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The overlap protects the boundary.&lt;/p&gt;

&lt;p&gt;Without it, a word beginning near the end of one chunk and finishing at the beginning of the next can be truncated. A small overlap gives both requests enough surrounding audio to decode the boundary more reliably.&lt;/p&gt;

&lt;p&gt;But overlapping chunks also create duplicate transcript content.&lt;/p&gt;

&lt;p&gt;Do not merge them by simply concatenating strings.&lt;/p&gt;

&lt;p&gt;Track each chunk’s original start time, offset the returned timestamps accordingly, and reconcile the overlapping region when building the final transcript.&lt;/p&gt;

&lt;p&gt;For especially long files, compare chunking against asynchronous transcription before automatically deciding that one large synchronous HTTP request is the right architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add retry logic without retrying everything
&lt;/h2&gt;

&lt;p&gt;Production networks fail.&lt;/p&gt;

&lt;p&gt;You should expect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;connection errors&lt;/li&gt;
&lt;li&gt;timeouts&lt;/li&gt;
&lt;li&gt;rate limits&lt;/li&gt;
&lt;li&gt;temporary upstream failures&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Retries help, but only when they are selective.&lt;/p&gt;

&lt;p&gt;Retrying an invalid request five times does not make it valid. Authentication errors and malformed payloads generally need intervention rather than exponential backoff.&lt;/p&gt;

&lt;p&gt;A better pattern is to retry connection failures, timeouts, rate limiting, and transient server responses.&lt;/p&gt;

&lt;p&gt;Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;tenacity&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;retry&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;retry_if_exception&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;stop_after_attempt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;wait_exponential&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="n"&gt;ENDPOINT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.smallest.ai/waves/v1/stt/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;TRANSIENT_STATUS_CODES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="mi"&gt;429&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;502&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;503&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;504&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;should_retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;BaseException&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;isinstance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Timeout&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;ConnectionError&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;isinstance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;HTTPError&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;
        &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
            &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;TRANSIENT_STATUS_CODES&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;


&lt;span class="nd"&gt;@retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;retry&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;retry_if_exception&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;should_retry&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;wait&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;wait_exponential&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;multiplier&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;min&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;max&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;stop&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;stop_after_attempt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;reraise&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;transcribe_with_retries&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Transcribe audio and retry only transient failures.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;audio_path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;audio_path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;is_file&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;FileNotFoundError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Audio file not found: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;audio_path&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;api_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;audio_path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;audio_file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;ENDPOINT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pulse-pro&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;language&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;en&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;word_timestamps&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/octet-stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;audio_file&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;120&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In a larger system, retries should also be observable.&lt;/p&gt;

&lt;p&gt;Record the request ID when the API provides one, track retry counts, distinguish permanent from transient failures, and make sure repeated jobs do not create duplicate downstream records.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prevent duplicate transcription work
&lt;/h2&gt;

&lt;p&gt;Reliability and cost control often point to the same design.&lt;/p&gt;

&lt;p&gt;If the same audio file can be submitted more than once, calculate a deterministic content hash and use it as an idempotency or cache key in your own application.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;(Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;sha256_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;digest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sha256&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;block&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;iter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;block&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before creating a new transcription job, check whether that hash already has a completed result.&lt;/p&gt;

&lt;p&gt;Other useful controls include checking recording duration before submission, setting application-level limits, storing completed responses durably, and separating transcription workers from the request path when jobs can take significant time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Accuracy is an application-level measurement
&lt;/h2&gt;

&lt;p&gt;Word Error Rate, or WER, is useful because it forces you to quantify transcription errors.&lt;/p&gt;

&lt;p&gt;At 5% WER, a 500-word transcript corresponds to roughly 25 word-level errors.&lt;/p&gt;

&lt;p&gt;Whether that is acceptable depends entirely on what happens next.&lt;/p&gt;

&lt;p&gt;A rough meeting summary might tolerate mistakes that would be unacceptable in a workflow where transcript content automatically updates records, triggers transactions, or feeds a compliance process.&lt;/p&gt;

&lt;p&gt;The important lesson is not to choose a universal “good” WER.&lt;/p&gt;

&lt;p&gt;Test the system on your own audio.&lt;/p&gt;

&lt;p&gt;Start with recordings from the actual environment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the microphones users really have&lt;/li&gt;
&lt;li&gt;the codecs used in production&lt;/li&gt;
&lt;li&gt;the accents and languages you expect&lt;/li&gt;
&lt;li&gt;real background noise&lt;/li&gt;
&lt;li&gt;overlapping speakers&lt;/li&gt;
&lt;li&gt;actual product names and domain terminology&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A short representative sample can expose obvious problems during prototyping, but production evaluation should grow into a larger labeled dataset.&lt;/p&gt;

&lt;p&gt;This is also why a clean demo clip is a weak benchmark. You are not deploying the benchmark. You are deploying your own acoustic environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  From prototype to production
&lt;/h2&gt;

&lt;p&gt;A dependable Python transcription pipeline usually comes back to a few engineering disciplines.&lt;/p&gt;

&lt;p&gt;Normalize inconsistent input when necessary.&lt;/p&gt;

&lt;p&gt;Treat the API response as structured data rather than one transcript string.&lt;/p&gt;

&lt;p&gt;Choose the transcription model according to the actual input and features you need.&lt;/p&gt;

&lt;p&gt;Keep credentials server-side.&lt;/p&gt;

&lt;p&gt;Handle long recordings deliberately instead of assuming one synchronous request will always succeed.&lt;/p&gt;

&lt;p&gt;Retry transient failures, not permanent ones.&lt;/p&gt;

&lt;p&gt;Cache completed work when the same recording can be submitted more than once.&lt;/p&gt;

&lt;p&gt;And above all, evaluate accuracy using the audio your application will actually receive.&lt;/p&gt;

&lt;p&gt;The first successful transcription request proves that the API works.&lt;/p&gt;

&lt;p&gt;Everything around that request determines whether your application works.&lt;/p&gt;

&lt;p&gt;If you want to test the pipeline against your own recording, &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=audio-to-text-api-how-to-convert-recorded-audio-into-accurate-transcripts-programmatically"&gt;start building with the Smallest AI API&lt;/a&gt;, create an API key, and run the pre-recorded Python example with representative audio from your application.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>speechrecognition</category>
      <category>machinelearning</category>
      <category>python</category>
    </item>
    <item>
      <title>How to Tune Voice Activity Detection for Low-Latency Voice Apps</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Wed, 26 Aug 2026 06:19:59 +0000</pubDate>
      <link>https://dev.to/smallestai/how-to-tune-voice-activity-detection-for-low-latency-voice-apps-1e7n</link>
      <guid>https://dev.to/smallestai/how-to-tune-voice-activity-detection-for-low-latency-voice-apps-1e7n</guid>
      <description>&lt;p&gt;When a real-time voice app feels slow or keeps misunderstanding users, it is tempting to start debugging the speech recognizer, the language model, or the network.&lt;/p&gt;

&lt;p&gt;Sometimes the problem is earlier.&lt;/p&gt;

&lt;p&gt;Before ASR processes a word, voice activity detection (VAD) has already decided whether the incoming audio is worth sending downstream.&lt;/p&gt;

&lt;p&gt;A good VAD configuration makes the pipeline feel responsive. A bad one can send keyboard noise into transcription, clip the first word of an utterance, delay turn-taking, or make an agent respond to its own audio.&lt;/p&gt;

&lt;p&gt;That makes VAD less of a preprocessing checkbox and more of a load-bearing part of real-time voice architecture.&lt;/p&gt;

&lt;p&gt;The broader &lt;a href="https://smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;Smallest AI&lt;/a&gt; stack covers real-time speech models and voice applications, but this article focuses specifically on the boundary before transcription: how to decide when speech begins, when it stops, and how aggressively the pipeline should react.&lt;/p&gt;

&lt;h2&gt;
  
  
  What VAD actually controls
&lt;/h2&gt;

&lt;p&gt;At its simplest, VAD repeatedly answers one question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is this frame speech or non-speech?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most implementations process audio in short windows, commonly around 10 to 30 milliseconds.&lt;/p&gt;

&lt;p&gt;Those frame-level decisions then control the rest of the system:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Start sending audio to ASR.&lt;/li&gt;
&lt;li&gt;Continue an active speech segment.&lt;/li&gt;
&lt;li&gt;Hold during a brief pause.&lt;/li&gt;
&lt;li&gt;Flush buffered audio.&lt;/li&gt;
&lt;li&gt;Stop the stream or begin endpointing logic.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In VoIP systems, VAD can avoid sending long stretches of silence.&lt;/p&gt;

&lt;p&gt;In transcription systems, it determines which audio reaches ASR.&lt;/p&gt;

&lt;p&gt;In conversational voice agents, VAD becomes one of the signals used to decide whether the user has started or stopped talking.&lt;/p&gt;

&lt;p&gt;The difficult part is that production audio rarely looks like a clean speech dataset.&lt;/p&gt;

&lt;p&gt;Your microphone may also capture HVAC noise, typing, traffic, music, another speaker, breathing, lip noise, or audio from the agent itself.&lt;/p&gt;

&lt;p&gt;VAD has to make its decision immediately, without knowing what the next few hundred milliseconds will contain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frame size is part of your latency budget
&lt;/h2&gt;

&lt;p&gt;Frame size is one of the first VAD parameters worth examining.&lt;/p&gt;

&lt;p&gt;Short frames, such as 10 ms windows, let the system detect speech onset quickly. That can reduce the delay before ASR begins receiving useful audio.&lt;/p&gt;

&lt;p&gt;The trade-off is context.&lt;/p&gt;

&lt;p&gt;With less audio inside each decision window, short frames may be easier to misclassify in difficult acoustic environments.&lt;/p&gt;

&lt;p&gt;Longer frames, such as 30 ms windows, contain more information and can make classification more stable. But they also delay the first speech decision.&lt;/p&gt;

&lt;p&gt;Thirty milliseconds does not sound significant in isolation.&lt;/p&gt;

&lt;p&gt;It becomes significant when you add it to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Audio capture and buffering&lt;/li&gt;
&lt;li&gt;Network transport&lt;/li&gt;
&lt;li&gt;Speech recognition&lt;/li&gt;
&lt;li&gt;Endpoint detection&lt;/li&gt;
&lt;li&gt;LLM inference&lt;/li&gt;
&lt;li&gt;Text-to-speech generation&lt;/li&gt;
&lt;li&gt;Playback buffering&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Voice engineering guidance has traditionally treated roughly 150 ms of one-way delay as an important quality boundary for highly interactive communication, while normal human turn-taking can occur on the order of a few hundred milliseconds.&lt;/p&gt;

&lt;p&gt;In that environment, several small delays can consume a meaningful percentage of the entire interaction budget.&lt;/p&gt;

&lt;p&gt;That does &lt;strong&gt;not&lt;/strong&gt; mean every system should use 10 ms frames.&lt;/p&gt;

&lt;p&gt;A controlled call-center deployment with standardized headsets has a very different noise profile from a mobile application being used in kitchens, cars, cafés, and sidewalks.&lt;/p&gt;

&lt;p&gt;Choose frame size against the audio your application actually receives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Classical VAD vs. neural VAD
&lt;/h2&gt;

&lt;p&gt;The right VAD architecture depends heavily on compute constraints and acoustic conditions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Classical VAD
&lt;/h3&gt;

&lt;p&gt;Classical detectors remain useful because they are inexpensive and predictable.&lt;/p&gt;

&lt;p&gt;The WebRTC VAD implementation, for example, uses Gaussian Mixture Models to compare speech and background-noise probabilities from audio features. Its implementation supports frame durations such as 10, 20, and 30 ms.&lt;/p&gt;

&lt;p&gt;You can inspect the &lt;a href="https://webrtc.googlesource.com/src/+/refs/heads/main/common_audio/vad/vad_core.c" rel="noopener noreferrer"&gt;WebRTC VAD implementation&lt;/a&gt; directly if you want to understand the statistical decision path.&lt;/p&gt;

&lt;p&gt;This kind of detector makes sense when you need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Low compute overhead&lt;/li&gt;
&lt;li&gt;On-device execution&lt;/li&gt;
&lt;li&gt;Predictable processing time&lt;/li&gt;
&lt;li&gt;Controlled acoustic environments&lt;/li&gt;
&lt;li&gt;No dependency on GPU inference&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Neural VAD
&lt;/h3&gt;

&lt;p&gt;Neural VAD usually uses a compact learned model rather than relying only on hand-designed spectral rules.&lt;/p&gt;

&lt;p&gt;That can make it more robust when background noise changes over time.&lt;/p&gt;

&lt;p&gt;Music, television, overlapping speakers, traffic, and other non-stationary noise can be particularly difficult for simpler statistical detectors.&lt;/p&gt;

&lt;p&gt;The trade-off is inference cost.&lt;/p&gt;

&lt;p&gt;For server-side pipelines that already have sufficient compute, that cost may be acceptable. On constrained devices or latency-sensitive edge deployments, it may not be.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hybrid VAD
&lt;/h3&gt;

&lt;p&gt;You do not necessarily have to choose one detector for every frame.&lt;/p&gt;

&lt;p&gt;A production system can use a lightweight detector as the first gate and invoke a more expensive model only for uncertain or difficult segments.&lt;/p&gt;

&lt;p&gt;The objective is not architectural purity.&lt;/p&gt;

&lt;p&gt;The objective is to avoid spending expensive inference on obvious silence while still handling noisy edge cases reliably.&lt;/p&gt;

&lt;h2&gt;
  
  
  VAD and endpointing solve different problems
&lt;/h2&gt;

&lt;p&gt;A common implementation mistake is treating VAD and endpointing as interchangeable.&lt;/p&gt;

&lt;p&gt;They are not.&lt;/p&gt;

&lt;p&gt;VAD answers:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is the user producing speech right now?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Endpointing answers:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Has the user finished their turn?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Those are different questions.&lt;/p&gt;

&lt;p&gt;Consider this sentence:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Can you book a meeting with... Sarah tomorrow?"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A speaker may naturally pause after "with" while remembering the name.&lt;/p&gt;

&lt;p&gt;During that pause, VAD may correctly classify several frames as non-speech.&lt;/p&gt;

&lt;p&gt;That does not mean the conversational turn has ended.&lt;/p&gt;

&lt;p&gt;A basic endpointing system might use a rule such as:&lt;/p&gt;

&lt;p&gt;800 ms of silence → end of turn&lt;/p&gt;

&lt;p&gt;That is easy to implement, but it can cause interruptions when users pause to think.&lt;/p&gt;

&lt;p&gt;More sophisticated endpointing systems can combine VAD state with additional information such as partial transcription, semantic completeness, or dedicated turn-taking models.&lt;/p&gt;

&lt;p&gt;VAD remains a low-level signal.&lt;/p&gt;

&lt;p&gt;Endpointing turns that signal into a conversational decision.&lt;/p&gt;

&lt;p&gt;If you are budgeting latency across the entire conversational pipeline, the distinction matters. Smallest AI's guide to &lt;a href="https://smallest.ai/blog/designing-voice-assistants-stt-llm-tts-tools-and-latency-budget?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;STT, LLM, TTS, tools, and latency budgeting&lt;/a&gt; looks at how those delays accumulate across the broader voice-agent stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  False triggers are usually where production VAD hurts
&lt;/h2&gt;

&lt;p&gt;VAD errors broadly fall into two categories.&lt;/p&gt;

&lt;h3&gt;
  
  
  False positives
&lt;/h3&gt;

&lt;p&gt;A false positive happens when non-speech audio is classified as speech.&lt;/p&gt;

&lt;p&gt;That can push unnecessary audio into ASR.&lt;/p&gt;

&lt;p&gt;Sometimes the result is harmless gibberish.&lt;/p&gt;

&lt;p&gt;A more dangerous failure happens when the recognizer produces plausible text from noise. Downstream components may then treat something that never happened as a real user utterance.&lt;/p&gt;

&lt;p&gt;For a conversational agent, a cough, television voice, keyboard sound, or door slam should not become an actionable request.&lt;/p&gt;

&lt;h3&gt;
  
  
  False negatives
&lt;/h3&gt;

&lt;p&gt;False negatives happen when actual speech is classified as non-speech.&lt;/p&gt;

&lt;p&gt;The most noticeable version is a clipped speech onset.&lt;/p&gt;

&lt;p&gt;If VAD opens the gate too late, the downstream recognizer may receive:&lt;/p&gt;

&lt;p&gt;...eed to change my booking&lt;/p&gt;

&lt;p&gt;instead of:&lt;/p&gt;

&lt;p&gt;I need to change my booking&lt;/p&gt;

&lt;p&gt;Users notice this immediately.&lt;/p&gt;

&lt;p&gt;They start repeating words, speaking unnaturally, or assuming the application is not listening.&lt;/p&gt;

&lt;p&gt;In production traffic, several inputs frequently cause trouble:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Non-stationary noise:&lt;/strong&gt; music, television, traffic, and overlapping speakers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Breath and lip sounds:&lt;/strong&gt; these can resemble speech-like acoustic events.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hesitation sounds:&lt;/strong&gt; "um," "uh," and other short, quiet speech can disappear when onset thresholds are too conservative.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Playback leakage:&lt;/strong&gt; an agent's own TTS may reach the microphone and trigger the detector.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Threshold-edge audio:&lt;/strong&gt; frames hovering near the activation threshold can cause rapid speech/non-speech switching.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The solution is not simply "increase the threshold."&lt;/p&gt;

&lt;p&gt;Every adjustment changes which class of errors you are accepting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical VAD tuning in a real pipeline
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F706w6saqx2x1yeny2wl5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F706w6saqx2x1yeny2wl5.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There is no universal production threshold.&lt;/p&gt;

&lt;p&gt;Tune against the failure distribution of your application.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Problem&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Likely cause&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;What to adjust&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;First word gets clipped&lt;/td&gt;
&lt;td&gt;Activation threshold too high or no leading buffer&lt;/td&gt;
&lt;td&gt;Lower the threshold and add pre-roll&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent responds too late&lt;/td&gt;
&lt;td&gt;Trailing padding or endpoint window is too long&lt;/td&gt;
&lt;td&gt;Reduce the trailing-silence window&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Background noise triggers ASR&lt;/td&gt;
&lt;td&gt;Threshold too low or weak upstream cleanup&lt;/td&gt;
&lt;td&gt;Increase the threshold or improve noise suppression&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;VAD rapidly flips states&lt;/td&gt;
&lt;td&gt;Signal is hovering near the boundary&lt;/td&gt;
&lt;td&gt;Add hysteresis smoothing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent detects its own TTS&lt;/td&gt;
&lt;td&gt;Playback is leaking into the microphone&lt;/td&gt;
&lt;td&gt;Add echo cancellation and playback-specific VAD behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Log the decisions, not just the transcripts
&lt;/h3&gt;

&lt;p&gt;If possible, capture VAD state alongside a representative sample of audio.&lt;/p&gt;

&lt;p&gt;You want to know:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which frames triggered speech?&lt;/li&gt;
&lt;li&gt;What did the audio actually contain?&lt;/li&gt;
&lt;li&gt;How often did speech onset get clipped?&lt;/li&gt;
&lt;li&gt;Which noise types produced false positives?&lt;/li&gt;
&lt;li&gt;How long did the system wait before closing an utterance?&lt;/li&gt;
&lt;li&gt;Were failures concentrated on specific devices or environments?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Synthetic noise tests are useful for regression.&lt;/p&gt;

&lt;p&gt;They are not a substitute for observing the acoustic environments your users actually create.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tune activation threshold and padding together
&lt;/h3&gt;

&lt;p&gt;Two settings tend to dominate day-to-day tuning:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Activation threshold&lt;/strong&gt; controls how confident the detector must be before a frame becomes speech.&lt;/p&gt;

&lt;p&gt;Lowering it improves sensitivity but usually increases false positives.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Padding&lt;/strong&gt; controls how much audio remains part of the active segment around speech boundaries.&lt;/p&gt;

&lt;p&gt;On the trailing edge, a practical conversational starting point is often around &lt;strong&gt;200–300 ms&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is long enough to absorb many natural micro-pauses without leaving the pipeline open indefinitely.&lt;/p&gt;

&lt;p&gt;It is a baseline, not a universal default.&lt;/p&gt;

&lt;p&gt;Your application still needs measurement.&lt;/p&gt;

&lt;p&gt;On the leading edge, sensitivity often deserves extra weight because a small pre-roll buffer is cheaper than losing the beginning of a user's sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why hysteresis matters
&lt;/h2&gt;

&lt;p&gt;Suppose your detector produces confidence values close to a threshold:&lt;/p&gt;

&lt;p&gt;0.48, 0.52, 0.49, 0.53, 0.47&lt;/p&gt;

&lt;p&gt;With a hard threshold at 0.50, the VAD state may repeatedly switch:&lt;/p&gt;

&lt;p&gt;off → on → off → on → off&lt;/p&gt;

&lt;p&gt;That state chatter is difficult for downstream components.&lt;/p&gt;

&lt;p&gt;Hysteresis introduces stability.&lt;/p&gt;

&lt;p&gt;Instead of using exactly the same transition rule in both directions, require sustained evidence before changing states.&lt;/p&gt;

&lt;p&gt;For example, entering speech may require several qualifying frames, while leaving speech may require a longer sequence below the deactivation boundary.&lt;/p&gt;

&lt;p&gt;The exact values depend on your detector, but the principle is broadly useful:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A frame-level classifier does not have to become a frame-level state transition.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Barge-in changes the operating conditions
&lt;/h2&gt;

&lt;p&gt;Voice agents introduce a problem that pure transcription systems often avoid.&lt;/p&gt;

&lt;p&gt;The system may be speaking while the user starts talking.&lt;/p&gt;

&lt;p&gt;Now the microphone contains both:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The user's interruption&lt;/li&gt;
&lt;li&gt;The agent's synthesized voice&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the barge-in problem.&lt;/p&gt;

&lt;p&gt;Acoustic echo cancellation is one of the first defenses because it reduces playback leakage before VAD evaluates the microphone signal.&lt;/p&gt;

&lt;p&gt;But echo cancellation alone does not eliminate the tuning problem.&lt;/p&gt;

&lt;p&gt;Your acoustic conditions during playback are different from your acoustic conditions during silence.&lt;/p&gt;

&lt;p&gt;That means the same VAD sensitivity may not be optimal in both states.&lt;/p&gt;

&lt;p&gt;If the playback-mode threshold is too high, real interruptions get missed.&lt;/p&gt;

&lt;p&gt;If it is too low, the system detects its own TTS and interrupts itself.&lt;/p&gt;

&lt;p&gt;Treat playback and non-playback as distinct operating modes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi-microphone systems can improve the input before VAD
&lt;/h2&gt;

&lt;p&gt;When multiple microphone channels are available, spatial processing can improve what the detector sees.&lt;/p&gt;

&lt;p&gt;Beamforming attempts to emphasize sound arriving from a target direction while suppressing other sources.&lt;/p&gt;

&lt;p&gt;That improves signal-to-noise ratio before classification.&lt;/p&gt;

&lt;p&gt;A cleaner input can reduce both false positives and false negatives without changing the VAD model itself.&lt;/p&gt;

&lt;p&gt;This illustrates a broader point:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not every VAD problem should be solved inside VAD.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Sometimes the correct fix is upstream.&lt;/p&gt;

&lt;p&gt;Noise suppression, acoustic echo cancellation, beamforming, microphone placement, gain control, and device-specific audio processing can all change the detector's error rate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where VAD belongs in the voice stack
&lt;/h2&gt;

&lt;p&gt;A simplified real-time speech pipeline might look like this:&lt;/p&gt;

&lt;p&gt;microphone → echo cancellation → noise suppression → VAD → ASR → endpointing/application logic&lt;/p&gt;

&lt;p&gt;The exact order depends on the system, but placement matters.&lt;/p&gt;

&lt;p&gt;Running noise suppression before VAD gives the detector a cleaner signal.&lt;/p&gt;

&lt;p&gt;Running VAD directly on raw audio means its threshold must tolerate every acoustic artifact that reaches the microphone.&lt;/p&gt;

&lt;p&gt;The same rule applies downstream.&lt;/p&gt;

&lt;p&gt;If VAD discards a frame containing speech, ASR cannot reconstruct audio it never received.&lt;/p&gt;

&lt;p&gt;If VAD sends noise into ASR, the recognizer has to decide what to do with audio that should have been filtered earlier.&lt;/p&gt;

&lt;p&gt;For a hosted downstream STT layer, &lt;a href="https://smallest.ai/speech-to-text?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;Pulse speech-to-text&lt;/a&gt; is Smallest AI's product for real-time and recorded transcription, including live-audio and voice-agent workloads.&lt;/p&gt;

&lt;p&gt;That does not make the STT model a replacement for your VAD design.&lt;/p&gt;

&lt;p&gt;It makes the VAD boundary easier to reason about: your recognizer can only process the audio your preprocessing layer decides to pass downstream.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing VAD against a real speech pipeline
&lt;/h2&gt;

&lt;p&gt;A useful VAD evaluation should not stop at frame-level accuracy.&lt;/p&gt;

&lt;p&gt;Measure what happens to the rest of the application.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Run representative audio through your VAD configuration.&lt;/li&gt;
&lt;li&gt;Preserve the accepted audio exactly as the recognizer would receive it.&lt;/li&gt;
&lt;li&gt;Send that audio through your STT layer.&lt;/li&gt;
&lt;li&gt;Compare transcripts across threshold and padding configurations.&lt;/li&gt;
&lt;li&gt;Measure speech-onset clipping.&lt;/li&gt;
&lt;li&gt;Measure false ASR activations caused by noise.&lt;/li&gt;
&lt;li&gt;Measure how VAD and trailing padding affect end-to-end response time.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This tells you something a standalone VAD score cannot:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;whether the detector's mistakes actually damage the product.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Developers who want to test the downstream transcription side can use the &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;Smallest AI API&lt;/a&gt; as the speech layer in this type of evaluation.&lt;/p&gt;

&lt;p&gt;The VAD itself can remain in your client, media server, or preprocessing service. The API then gives you a consistent downstream component against which you can evaluate how different gating decisions affect transcription.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create and store the API key
&lt;/h2&gt;

&lt;p&gt;Keep the API key in an environment variable rather than hard-coding it into the application.&lt;/p&gt;

&lt;p&gt;Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key-here"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every authenticated request sends the value through the Authorization header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer &amp;lt;SMALLEST_API_KEY value&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the key on your server.&lt;/p&gt;

&lt;p&gt;Do not expose it in browser JavaScript, client-side React code, mobile application code, public repositories, screenshots, query parameters, or client-side logs.&lt;/p&gt;

&lt;p&gt;For production deployments, store it in your hosting environment's server-side secrets manager rather than committing credentials to configuration files.&lt;/p&gt;

&lt;p&gt;This article intentionally does not invent an API request specifically for VAD because the VAD architecture described here is an upstream pipeline concern rather than a Smallest AI-specific VAD endpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to measure before shipping
&lt;/h2&gt;

&lt;p&gt;If VAD is going into production, evaluate it as a system component rather than a classifier in isolation.&lt;/p&gt;

&lt;p&gt;Useful measurements include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Speech-start detection delay&lt;/li&gt;
&lt;li&gt;Speech-end detection delay&lt;/li&gt;
&lt;li&gt;False-positive rate by noise category&lt;/li&gt;
&lt;li&gt;False-negative rate at utterance onset&lt;/li&gt;
&lt;li&gt;Percentage of utterances with clipped first words&lt;/li&gt;
&lt;li&gt;Number of unnecessary ASR activations&lt;/li&gt;
&lt;li&gt;Endpointing delay after actual user completion&lt;/li&gt;
&lt;li&gt;Barge-in success during TTS playback&lt;/li&gt;
&lt;li&gt;Behavior across microphones, devices, and environments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The correct configuration depends on the cost of each failure.&lt;/p&gt;

&lt;p&gt;A dictation application may tolerate a slightly slower onset if it reduces false activations.&lt;/p&gt;

&lt;p&gt;A voice agent may prefer a more sensitive leading edge because clipped words damage conversational flow immediately.&lt;/p&gt;

&lt;p&gt;Production tuning is about choosing those trade-offs deliberately.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;VAD is an upstream latency and quality control point, not just a silence detector.&lt;/li&gt;
&lt;li&gt;Frame size affects how quickly speech can be detected and how much acoustic context the classifier receives.&lt;/li&gt;
&lt;li&gt;False positives waste downstream work; false negatives can remove speech permanently.&lt;/li&gt;
&lt;li&gt;VAD and endpointing solve different problems.&lt;/li&gt;
&lt;li&gt;Trailing padding directly affects how quickly a conversational system can respond.&lt;/li&gt;
&lt;li&gt;Around 200–300 ms of trailing padding is a reasonable starting point for many conversational systems, but it should be validated against real traffic.&lt;/li&gt;
&lt;li&gt;Hysteresis can stabilize frame-level decisions without replacing the underlying detector.&lt;/li&gt;
&lt;li&gt;Barge-in requires echo handling and often different VAD behavior during playback.&lt;/li&gt;
&lt;li&gt;Beamforming and noise suppression can improve VAD performance before you touch the detector itself.&lt;/li&gt;
&lt;li&gt;The right threshold is the one that minimizes the errors that matter most to your application.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Voice activity detection is easy to describe because the output looks binary.&lt;/p&gt;

&lt;p&gt;Production behavior is not.&lt;/p&gt;

&lt;p&gt;Every threshold, frame size, pre-roll buffer, silence window, and state-transition rule changes what the rest of your voice pipeline receives and when it receives it.&lt;/p&gt;

&lt;p&gt;That is why VAD tuning should be evaluated alongside ASR and endpointing rather than as an isolated preprocessing benchmark.&lt;/p&gt;

&lt;p&gt;Measure with real audio. Log the boundaries. Inspect false triggers. Watch for clipped onsets. Test again while TTS is playing.&lt;/p&gt;

&lt;p&gt;Then optimize the failure mode that actually damages your application.&lt;/p&gt;

&lt;p&gt;If you want to evaluate how those VAD decisions affect downstream transcription, &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=designing-voice-assistants-stt-llm-tts-tools-and-latency-budget"&gt;create an API key and test the pipeline with your own audio using Smallest AI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>voiceai</category>
      <category>speechrecognition</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Building an AI Dubbing Pipeline That Survives Production: STT Translation TTS</title>
      <dc:creator>Smallest AI</dc:creator>
      <pubDate>Wed, 26 Aug 2026 05:42:43 +0000</pubDate>
      <link>https://dev.to/smallestai/building-an-ai-dubbing-pipeline-that-survives-production-stt-translation-tts-4jfh</link>
      <guid>https://dev.to/smallestai/building-an-ai-dubbing-pipeline-that-survives-production-stt-translation-tts-4jfh</guid>
      <description>&lt;p&gt;&lt;strong&gt;AI dubbing looks simple on a whiteboard:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;speech → text → translation → speech&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That description is technically correct. It is also where most of the engineering details disappear.&lt;/p&gt;

&lt;p&gt;Once you move beyond a single-speaker demo, every stage starts depending on metadata produced by the stage before it. A transcription error becomes a translation error. A missing speaker label assigns the wrong synthetic voice. A translated sentence that runs 30% longer than the source pushes the next line out of sync.&lt;/p&gt;

&lt;p&gt;The result is a pipeline where failures rarely stay local.&lt;/p&gt;

&lt;p&gt;For developers building their own stack, the useful mental model is not “connect three APIs.” It preserves &lt;strong&gt;enough information between those APIs that the final audio still matches the original performance&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Platforms such as YouTube and Prime Video have already expanded automated or AI-assisted dubbing workflows, but building a reliable version yourself still means solving several production problems around transcription, translation, timing, voice synthesis, and review.&lt;/p&gt;

&lt;p&gt;This article walks through that architecture from ingestion to final audio assembly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an AI dubbing pipeline actually does
&lt;/h2&gt;

&lt;p&gt;AI dubbing replaces speech in one language with synthesized speech in another.&lt;/p&gt;

&lt;p&gt;At the center are three stages:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Speech-to-Text (STT)&lt;/strong&gt; converts the source audio into structured text.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Translation&lt;/strong&gt; converts that text into the target language.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Text-to-Speech (TTS)&lt;/strong&gt; renders the translation back into audio.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;But production dubbing asks those stages to preserve more than words.&lt;/p&gt;

&lt;p&gt;You also need to carry:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;speaker identity&lt;/li&gt;
&lt;li&gt;word and segment timestamps&lt;/li&gt;
&lt;li&gt;emotional intent&lt;/li&gt;
&lt;li&gt;speaking style&lt;/li&gt;
&lt;li&gt;terminology&lt;/li&gt;
&lt;li&gt;target duration&lt;/li&gt;
&lt;li&gt;confidence and QA status&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without that metadata, you can generate translated speech but not necessarily a usable dub.&lt;/p&gt;

&lt;p&gt;This is why a production pipeline usually looks more like:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;audio preprocessing → STT → diarization → translation → timing validation → human QA → TTS → alignment → audio post-processing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you’re exploring speech infrastructure for this kind of workflow, &lt;a href="https://smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;Smallest AI&lt;/a&gt; exposes the speech components independently, which is useful when you want to control these stages yourself rather than treat dubbing as one opaque operation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 1: Speech-to-Text sets the quality ceiling
&lt;/h2&gt;

&lt;p&gt;A dubbing pipeline can rarely recover cleanly from a bad transcript.&lt;/p&gt;

&lt;p&gt;Suppose an STT model mishears a company name or technical term. The translation system receives the incorrect word, translates it fluently, and the TTS model says the mistake naturally.&lt;/p&gt;

&lt;p&gt;By the time somebody hears the final audio, the original transcription error has been polished by two more models.&lt;/p&gt;

&lt;p&gt;That makes STT quality an upstream constraint on everything that follows.&lt;/p&gt;

&lt;p&gt;For dubbing, plain transcript text is not enough. A useful STT result should ideally include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;word-level timestamps&lt;/li&gt;
&lt;li&gt;speaker labels&lt;/li&gt;
&lt;li&gt;sentence or utterance boundaries&lt;/li&gt;
&lt;li&gt;confidence information&lt;/li&gt;
&lt;li&gt;punctuation&lt;/li&gt;
&lt;li&gt;domain-specific vocabulary handling&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Why word timestamps matter
&lt;/h3&gt;

&lt;p&gt;Imagine the source contains this line:&lt;/p&gt;

&lt;p&gt;“Hello, welcome to the new headquarters.”&lt;/p&gt;

&lt;p&gt;The sentence begins at 0.52 seconds and ends at 3.10 seconds.&lt;/p&gt;

&lt;p&gt;That gives the downstream system roughly &lt;strong&gt;2.58 seconds&lt;/strong&gt; for the translated version.&lt;/p&gt;

&lt;p&gt;Without timestamps, the translation and TTS layers have no reliable timing window to target.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why diarization matters
&lt;/h3&gt;

&lt;p&gt;Multi-speaker content creates another problem.&lt;/p&gt;

&lt;p&gt;Interviews, podcasts, panel discussions, films, and training videos all require the system to know &lt;strong&gt;who said what&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If the STT layer outputs one continuous transcript, the synthesis layer cannot reliably determine which voice should speak each translated segment.&lt;/p&gt;

&lt;p&gt;For multi-speaker implementations, treat &lt;a href="https://smallest.ai/blog/the-complete-speaker-diarization-api-guide-how-it-works-and-best-practices?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;speaker diarization&lt;/a&gt; as part of the transcription architecture, not an optional post-processing feature.&lt;/p&gt;

&lt;p&gt;Smallest AI’s &lt;a href="https://smallest.ai/speech-to-text?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;Pulse speech-to-text&lt;/a&gt; supports streaming and pre-recorded transcription, including features such as speaker diarization and word timestamps. Whatever STT system you choose, benchmark it using the actual audio your application will process.&lt;/p&gt;

&lt;p&gt;Podcast audio, accented speech, overlapping speakers, film dialogue, and noisy field recordings can behave very differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 2: Translation has to fit speech, not a document
&lt;/h2&gt;

&lt;p&gt;Translation APIs make converting a sentence from one language to another relatively straightforward.&lt;/p&gt;

&lt;p&gt;Dubbing introduces a different requirement:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The translated sentence has to be speakable inside approximately the same time window as the original.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That changes how the translation layer should be designed.&lt;/p&gt;

&lt;p&gt;A translated sentence can preserve meaning perfectly and still fail the dubbing workflow because it takes too long to say.&lt;/p&gt;

&lt;p&gt;Research into professional dubbing also shows that naturalness, translation quality, timing, and preservation of speech characteristics interact in more complicated ways than simply forcing equal character counts. The large-scale study &lt;a href="https://aclanthology.org/2023.tacl-1.25/" rel="noopener noreferrer"&gt;Dubbing in Practice&lt;/a&gt; is a useful reference for understanding those tradeoffs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Translate segments, not entire transcripts
&lt;/h3&gt;

&lt;p&gt;If the STT result already contains timed segments, preserve them.&lt;/p&gt;

&lt;p&gt;Instead of sending a complete 20-minute transcript through translation and trying to reconstruct alignment afterward, translate individual utterances or tightly grouped segments.&lt;/p&gt;

&lt;p&gt;For example, your internal pipeline might normalize STT output into a structure like this: (Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"segments"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"speaker"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"A"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"start"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.52&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"end"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;3.10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Hello, welcome to the new headquarters."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.98&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is an internal pipeline representation, not a provider-specific API request.&lt;/p&gt;

&lt;p&gt;Keeping the segment boundaries intact gives every later stage a stable ID, speaker, and timing window.&lt;/p&gt;

&lt;h3&gt;
  
  
  Add a length-validation layer
&lt;/h3&gt;

&lt;p&gt;After translation, compare the target-language output with the available source duration.&lt;/p&gt;

&lt;p&gt;A simple first-pass implementation can estimate spoken duration using historical speaking-rate data or a language-specific character/token heuristic.&lt;/p&gt;

&lt;p&gt;Then flag segments that exceed your tolerance.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;source_duration = 2.58s
estimated_translation_duration = 3.21s
difference = +24.4%
status = REVIEW
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact threshold should come from testing your content rather than being treated as universal. A system might initially flag anything more than 10–15% outside its target window, then tune that rule based on actual listening results.&lt;/p&gt;

&lt;p&gt;The important point is that timing problems should be detected &lt;strong&gt;before TTS generation&lt;/strong&gt;, not after you have rendered hundreds of unusable clips.&lt;/p&gt;

&lt;h3&gt;
  
  
  Preserve register and context
&lt;/h3&gt;

&lt;p&gt;Dubbing also needs more than literal translation.&lt;/p&gt;

&lt;p&gt;A production translation layer should preserve:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;formality&lt;/li&gt;
&lt;li&gt;slang&lt;/li&gt;
&lt;li&gt;technical terminology&lt;/li&gt;
&lt;li&gt;character relationships&lt;/li&gt;
&lt;li&gt;idioms&lt;/li&gt;
&lt;li&gt;proper nouns&lt;/li&gt;
&lt;li&gt;brand terminology&lt;/li&gt;
&lt;li&gt;emotional intent&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A casual speaker should not suddenly sound formal simply because the translation model selected a grammatically valid but stylistically inappropriate phrase.&lt;/p&gt;

&lt;p&gt;For technical, educational, or branded content, terminology enforcement is especially important. A glossary or controlled vocabulary can prevent the translation system from rewriting names and domain-specific terms differently from one segment to the next.&lt;/p&gt;

&lt;h3&gt;
  
  
  Human QA still matters
&lt;/h3&gt;

&lt;p&gt;Treat machine translation as draft dialogue.&lt;/p&gt;

&lt;p&gt;Certain outputs should be routed for review automatically, especially when they contain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;low-confidence source transcription&lt;/li&gt;
&lt;li&gt;idioms&lt;/li&gt;
&lt;li&gt;ambiguous names&lt;/li&gt;
&lt;li&gt;unusually large length differences&lt;/li&gt;
&lt;li&gt;culturally specific references&lt;/li&gt;
&lt;li&gt;terminology that must remain consistent&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal is not to manually review every generated word forever. It is to make the pipeline capable of recognizing when automation has lower confidence.&lt;/p&gt;

&lt;p&gt;For a broader look at these localization tradeoffs, Smallest AI’s guide to &lt;a href="https://smallest.ai/blog/ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;AI dubbing pipelines for translation, timing, and TTS&lt;/a&gt; covers the same problem from a localization perspective.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 3: TTS has to solve voice and timing together
&lt;/h2&gt;

&lt;p&gt;Once the translated dialogue has passed validation, TTS turns it back into speech.&lt;/p&gt;

&lt;p&gt;Generating audio is not the difficult part.&lt;/p&gt;

&lt;p&gt;Generating audio that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;sounds like the intended speaker&lt;/li&gt;
&lt;li&gt;preserves emotional intent&lt;/li&gt;
&lt;li&gt;fits the source timing window&lt;/li&gt;
&lt;li&gt;stays consistent across hundreds of segments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;is much harder.&lt;/p&gt;

&lt;h3&gt;
  
  
  Voice cloning changes speaker consistency
&lt;/h3&gt;

&lt;p&gt;A basic dubbing system can assign a preset voice to each speaker.&lt;/p&gt;

&lt;p&gt;A more advanced pipeline can create a voice representation from the source speaker and reuse it across translated segments.&lt;/p&gt;

&lt;p&gt;This matters because speaker identity is part of the original content.&lt;/p&gt;

&lt;p&gt;If someone appears throughout a 30-minute video, the translated version should not sound like three different people because different chunks were synthesized independently.&lt;/p&gt;

&lt;p&gt;Current &lt;a href="https://smallest.ai/text-to-speech?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;Lightning text-to-speech&lt;/a&gt; models from Smallest AI support voice cloning as part of the TTS workflow.&lt;/p&gt;

&lt;p&gt;For dubbing, voice quality should still be evaluated language by language. A voice that works well with one target language may not preserve the same accent, rhythm, or pronunciation behavior in another.&lt;/p&gt;

&lt;h3&gt;
  
  
  Timing is not just “increase the speed”
&lt;/h3&gt;

&lt;p&gt;Suppose the original line lasts 2.58 seconds but the translated speech naturally takes 3.1 seconds.&lt;/p&gt;

&lt;p&gt;You have several possible interventions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;shorten the translation&lt;/li&gt;
&lt;li&gt;increase speaking rate slightly&lt;/li&gt;
&lt;li&gt;modify pauses&lt;/li&gt;
&lt;li&gt;regenerate the translation with stricter length constraints&lt;/li&gt;
&lt;li&gt;allow small timeline drift&lt;/li&gt;
&lt;li&gt;correct the remaining difference in post-production&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these approaches works for every line.&lt;/p&gt;

&lt;p&gt;If you aggressively speed up every long translation, the dub starts sounding rushed. If you rewrite every sentence until its character count matches, meaning and naturalness can suffer.&lt;/p&gt;

&lt;p&gt;Production systems usually combine translation constraints, synthesis controls, and post-processing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Emotional fidelity is another constraint
&lt;/h3&gt;

&lt;p&gt;A speaker who is excited, sarcastic, uncertain, or angry carries information that is not contained in the transcript alone.&lt;/p&gt;

&lt;p&gt;TTS can generate the correct sentence while still changing the perceived performance.&lt;/p&gt;

&lt;p&gt;That is why dubbing evaluation needs listening tests rather than text-only checks.&lt;/p&gt;

&lt;p&gt;The pipeline needs to ask two different questions:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Did the system say the right thing?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;and&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Did it sound appropriate for the original scene?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Those are not the same test.&lt;/p&gt;

&lt;h2&gt;
  
  
  A production architecture for AI dubbing
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foj9s56rmt5gxx5bjrfbi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foj9s56rmt5gxx5bjrfbi.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The three core models become much easier to reason about when the surrounding pipeline is explicit.&lt;/p&gt;

&lt;p&gt;A practical batch architecture can look like this:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Ingest the source
&lt;/h3&gt;

&lt;p&gt;Accept the source video or audio and create an immutable reference asset.&lt;/p&gt;

&lt;p&gt;Keep the original timeline available throughout processing.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Preprocess audio
&lt;/h3&gt;

&lt;p&gt;Depending on the source, preprocessing might include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;extracting the dialogue track&lt;/li&gt;
&lt;li&gt;normalizing levels&lt;/li&gt;
&lt;li&gt;reducing noise&lt;/li&gt;
&lt;li&gt;detecting silence&lt;/li&gt;
&lt;li&gt;splitting extremely long files into manageable units&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Be careful with operations that change timing.&lt;/p&gt;

&lt;p&gt;If you remove silence before transcription, for example, STT timestamps may no longer map directly to the original video.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Run STT with timestamps and diarization
&lt;/h3&gt;

&lt;p&gt;Generate structured transcript data containing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;segment ID&lt;/li&gt;
&lt;li&gt;speaker ID&lt;/li&gt;
&lt;li&gt;start timestamp&lt;/li&gt;
&lt;li&gt;end timestamp&lt;/li&gt;
&lt;li&gt;text&lt;/li&gt;
&lt;li&gt;confidence&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Store this as structured data rather than flattening it into a text document.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Translate each segment
&lt;/h3&gt;

&lt;p&gt;Translate with enough surrounding context to preserve meaning while retaining the original segment IDs.&lt;/p&gt;

&lt;p&gt;Do not lose the mapping between source and translated dialogue.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Validate duration
&lt;/h3&gt;

&lt;p&gt;Estimate whether the translated line can fit the original timing window.&lt;/p&gt;

&lt;p&gt;Send problematic segments into a retry or review path.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Review risky segments
&lt;/h3&gt;

&lt;p&gt;A review interface should allow somebody to inspect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;source audio&lt;/li&gt;
&lt;li&gt;source transcript&lt;/li&gt;
&lt;li&gt;translation&lt;/li&gt;
&lt;li&gt;timing window&lt;/li&gt;
&lt;li&gt;confidence&lt;/li&gt;
&lt;li&gt;speaker&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This checkpoint is far cheaper than discovering translation mistakes after synthesis and final mixing.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. Generate target speech
&lt;/h3&gt;

&lt;p&gt;Route each approved translated segment to the correct speaker voice.&lt;/p&gt;

&lt;p&gt;Your own internal synthesis queue might contain data such as: (Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"segment_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"speaker_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"A"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"translated_text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Hola, bienvenido a la nueva sede."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"target_duration_seconds"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;2.58&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"voice_profile"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"speaker-A"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"delivery"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"friendly"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Again, this is an example of &lt;strong&gt;your application’s internal data model&lt;/strong&gt;, not a Smallest AI API request schema.&lt;/p&gt;

&lt;p&gt;The actual request body should follow whichever TTS provider’s current API documentation you are using.&lt;/p&gt;

&lt;h3&gt;
  
  
  8. Reassemble the timeline
&lt;/h3&gt;

&lt;p&gt;Place generated segments back at their corresponding source timestamps.&lt;/p&gt;

&lt;p&gt;Then restore or mix:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;music&lt;/li&gt;
&lt;li&gt;room tone&lt;/li&gt;
&lt;li&gt;environmental sound&lt;/li&gt;
&lt;li&gt;sound effects&lt;/li&gt;
&lt;li&gt;non-dialogue audio&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  9. Normalize and export
&lt;/h3&gt;

&lt;p&gt;Run final loudness and quality checks before producing the deliverable audio or remuxing it into the video.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create and store the API key
&lt;/h2&gt;

&lt;p&gt;If you prototype the speech stages using the &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;Smallest AI API&lt;/a&gt;, keep the API key in an environment variable rather than hard-coding it into the application.&lt;/p&gt;

&lt;p&gt;Before running the snippet, create a &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;Smallest.ai API key&lt;/a&gt; in the dashboard and store it in the SMALLEST_API_KEY environment variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;SMALLEST_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key-here"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Authenticated server-side requests should pass the value through the Authorization header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer &amp;lt;SMALLEST_API_KEY value&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the key on your server. Do not expose it in browser JavaScript, mobile application code, public repositories, screenshots, query parameters, or client-side logs.&lt;/p&gt;

&lt;p&gt;For production deployments, store credentials in your cloud or infrastructure provider’s server-side secrets-management system rather than committing them to source control.&lt;/p&gt;

&lt;p&gt;The exact STT and TTS endpoints, model names, and request fields can change over time, so use the current API documentation rather than copying an old request schema into a new production integration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where dubbing pipelines actually break
&lt;/h2&gt;

&lt;p&gt;Calling three APIs sequentially is not usually the difficult part.&lt;/p&gt;

&lt;p&gt;Production failures tend to happen in the state you carry between calls.&lt;/p&gt;

&lt;h2&gt;
  
  
  Timestamp drift
&lt;/h2&gt;

&lt;p&gt;Audio preprocessing can modify the timeline used by transcription.&lt;/p&gt;

&lt;p&gt;Suppose you remove a two-second silence before sending a clip to STT. Every timestamp after that edit is now offset relative to the original video.&lt;/p&gt;

&lt;p&gt;You need either:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a mapping between processed and original timestamps, or&lt;/li&gt;
&lt;li&gt;a preprocessing strategy that preserves the original timeline&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Otherwise the translated speech can be correct and still appear at the wrong moment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Speaker IDs changing between chunks
&lt;/h2&gt;

&lt;p&gt;Diarization systems often label speakers relative to the audio chunk being processed.&lt;/p&gt;

&lt;p&gt;That means:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;chunk 1 → Speaker A = Alice
chunk 2 → Speaker A = Bob
chunk 3 → Speaker B = Alice
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your TTS routing blindly trusts those labels, Alice’s voice clone can suddenly start reading Bob’s lines.&lt;/p&gt;

&lt;p&gt;For long or chunked media, add a speaker-reconciliation step before voice assignment. This can involve comparing speaker representations across chunks and mapping local diarization labels to stable global speaker IDs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Translation quality varying by language pair
&lt;/h2&gt;

&lt;p&gt;A pipeline validated on English → Spanish should not automatically be considered validated for English → Arabic, Thai, Hindi, Japanese, or another target language.&lt;/p&gt;

&lt;p&gt;Sentence structure, spoken duration, pronunciation behavior, translation-resource availability, and cultural adaptation requirements vary.&lt;/p&gt;

&lt;p&gt;Evaluate the &lt;strong&gt;entire pipeline&lt;/strong&gt; for each language pair you intend to support.&lt;/p&gt;

&lt;p&gt;That means measuring more than translation accuracy.&lt;/p&gt;

&lt;p&gt;Also evaluate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;timing fit&lt;/li&gt;
&lt;li&gt;pronunciation&lt;/li&gt;
&lt;li&gt;voice consistency&lt;/li&gt;
&lt;li&gt;prosody&lt;/li&gt;
&lt;li&gt;speaker identity&lt;/li&gt;
&lt;li&gt;cultural appropriateness&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For multilingual production specifically, this guide to &lt;a href="https://smallest.ai/blog/multilingual-voice-dubbing-for-product-videos-how-to-localize-audio-without-re-recording?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;localizing product videos without re-recording&lt;/a&gt; goes deeper into the localization workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Voice inconsistency across segments
&lt;/h2&gt;

&lt;p&gt;When a long recording is synthesized as hundreds of independent requests, subtle changes in pacing or delivery can accumulate.&lt;/p&gt;

&lt;p&gt;This can make a single speaker sound different between scenes even when the same voice profile is being used.&lt;/p&gt;

&lt;p&gt;Evaluate consistency across the complete program, not just isolated samples.&lt;/p&gt;

&lt;p&gt;If one voice needs to remain recognizable across a large content library, &lt;a href="https://smallest.ai/blog/voice-cloning-for-brand-consistency-how-teams-can-scale-a-single-voice-across-products-and-channels?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;voice cloning and brand consistency&lt;/a&gt; becomes an architectural concern rather than simply a model feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lip-sync adds another system
&lt;/h2&gt;

&lt;p&gt;Audio alignment and visual lip-sync are related, but they are not the same problem.&lt;/p&gt;

&lt;p&gt;Basic dubbing can align translated speech to approximately the same time window while leaving the original video untouched.&lt;/p&gt;

&lt;p&gt;True visual dubbing goes further by modifying facial motion to match the new phonemes.&lt;/p&gt;

&lt;p&gt;That requires an additional video-generation or facial-animation layer.&lt;/p&gt;

&lt;p&gt;For example, NVIDIA’s &lt;a href="https://docs.nvidia.com/nim/maxine/audio2face-2d/latest/overview.html" rel="noopener noreferrer"&gt;Audio2Face-2D documentation&lt;/a&gt; describes a system that uses audio to generate facial motion and synchronize mouth movement.&lt;/p&gt;

&lt;p&gt;Whether you need this depends heavily on the source material.&lt;/p&gt;

&lt;p&gt;Lip-sync may be less important for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;narrated screen recordings&lt;/li&gt;
&lt;li&gt;animated explainers&lt;/li&gt;
&lt;li&gt;podcasts converted to video&lt;/li&gt;
&lt;li&gt;slides with voice-over&lt;/li&gt;
&lt;li&gt;videos where the speaker’s face is rarely visible&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It becomes much more noticeable when a face occupies most of the frame.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-time dubbing is a different architecture
&lt;/h2&gt;

&lt;p&gt;Batch dubbing gives the system a major advantage: complete context.&lt;/p&gt;

&lt;p&gt;The transcription engine can process finished sentences. The translation model can see the entire utterance. The TTS system knows how much audio it needs to generate.&lt;/p&gt;

&lt;p&gt;Real-time dubbing removes that luxury.&lt;/p&gt;

&lt;p&gt;A live pipeline may need to perform:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;streaming STT → incremental translation → streaming TTS&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;before the speaker has even finished the complete thought.&lt;/p&gt;

&lt;p&gt;That creates new problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;partial transcripts can be revised&lt;/li&gt;
&lt;li&gt;sentence meaning can change at the end&lt;/li&gt;
&lt;li&gt;translation may require words that have not arrived yet&lt;/li&gt;
&lt;li&gt;synthesis needs to start before the final segment is available&lt;/li&gt;
&lt;li&gt;latency accumulates across each stage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For that architecture, &lt;a href="https://smallest.ai/blog/streaming-tts-explained-for-developers-ux-latency-and-cost-guide?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;streaming TTS&lt;/a&gt; changes how audio should be buffered and scheduled.&lt;/p&gt;

&lt;p&gt;Smallest AI also exposes &lt;a href="https://smallest.ai/speech-to-speech?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;Hydra speech-to-speech&lt;/a&gt; for low-latency, full-duplex speech applications.&lt;/p&gt;

&lt;p&gt;However, speech-to-speech and an inspectable dubbing pipeline solve different problems.&lt;/p&gt;

&lt;p&gt;If your workflow requires explicit translated text for QA, analytics, terminology control, moderation, or editing, keeping STT, translation, and TTS as explicit stages gives you much more control over the intermediate state.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build QA into the pipeline instead of adding it later
&lt;/h2&gt;

&lt;p&gt;A useful production design has checkpoints at three places.&lt;/p&gt;

&lt;h3&gt;
  
  
  After transcription
&lt;/h3&gt;

&lt;p&gt;Review or automatically flag:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;low-confidence words&lt;/li&gt;
&lt;li&gt;unclear names&lt;/li&gt;
&lt;li&gt;speaker changes&lt;/li&gt;
&lt;li&gt;overlapping speech&lt;/li&gt;
&lt;li&gt;domain vocabulary&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  After translation
&lt;/h3&gt;

&lt;p&gt;Review:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;meaning&lt;/li&gt;
&lt;li&gt;terminology&lt;/li&gt;
&lt;li&gt;register&lt;/li&gt;
&lt;li&gt;cultural adaptation&lt;/li&gt;
&lt;li&gt;timing fit&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  After synthesis
&lt;/h3&gt;

&lt;p&gt;Listen for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;clipped speech&lt;/li&gt;
&lt;li&gt;unnatural pacing&lt;/li&gt;
&lt;li&gt;incorrect pronunciation&lt;/li&gt;
&lt;li&gt;voice inconsistency&lt;/li&gt;
&lt;li&gt;missing emotion&lt;/li&gt;
&lt;li&gt;timeline drift&lt;/li&gt;
&lt;li&gt;audio-level differences&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At scale, you do not necessarily need humans to listen to every second.&lt;/p&gt;

&lt;p&gt;Confidence thresholds and automated checks can reduce the review set.&lt;/p&gt;

&lt;p&gt;But eliminating QA completely usually just moves the cost downstream, where mistakes are more expensive to repair.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline is really about preserving state
&lt;/h2&gt;

&lt;p&gt;An AI dubbing stack has three obvious models:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;STT → translation → TTS&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;But those models are not what make the system reliable.&lt;/p&gt;

&lt;p&gt;The useful architecture is the information that survives between them:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;speaker identity → timestamps → translation context → timing constraints → voice identity → QA status&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Lose any of those and the pipeline becomes harder to control.&lt;/p&gt;

&lt;p&gt;Keep them structured, and each stage becomes independently testable.&lt;/p&gt;

&lt;p&gt;For a first prototype, keep the scope deliberately small:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one source language&lt;/li&gt;
&lt;li&gt;one target language&lt;/li&gt;
&lt;li&gt;short clips&lt;/li&gt;
&lt;li&gt;one or two speakers&lt;/li&gt;
&lt;li&gt;explicit transcript review&lt;/li&gt;
&lt;li&gt;explicit translation review&lt;/li&gt;
&lt;li&gt;deterministic segment IDs&lt;/li&gt;
&lt;li&gt;final listening QA&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Once that works, expand into longer files, additional speakers, more language pairs, streaming, or visual lip-sync.&lt;/p&gt;

&lt;p&gt;The fastest path to a useful dubbing system is not adding more models. It is making the seams between the existing models observable.&lt;/p&gt;

&lt;p&gt;If you want to prototype the speech side with your own audio, create an API key and &lt;a href="https://app.smallest.ai/?utm_source=dev.to&amp;amp;utm_medium=vizup&amp;amp;utm_campaign=ai-dubbing-pipelines-localizing-video-audio-with-translation-timing-and-tts"&gt;start building with the Smallest AI developer platform&lt;/a&gt;, then validate the complete STT → translation → TTS path against the language pairs and content your application will actually process.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>texttospeech</category>
      <category>speechrecognition</category>
    </item>
  </channel>
</rss>
