<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Neeraj Sharma</title>
    <description>The latest articles on DEV Community by Neeraj Sharma (@neeraj_sharma_757fbb49aa7).</description>
    <link>https://dev.to/neeraj_sharma_757fbb49aa7</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4156727%2Fa4e3050c-181e-455b-8a50-2398b0f2b950.png</url>
      <title>DEV Community: Neeraj Sharma</title>
      <link>https://dev.to/neeraj_sharma_757fbb49aa7</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/neeraj_sharma_757fbb49aa7"/>
    <language>en</language>
    <item>
      <title>Showing live voice agent transcripts in a Flutter app with LiveKit</title>
      <dc:creator>Neeraj Sharma</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:53:20 +0000</pubDate>
      <link>https://dev.to/neeraj_sharma_757fbb49aa7/showing-live-voice-agent-transcripts-in-a-flutter-app-with-livekit-d4b</link>
      <guid>https://dev.to/neeraj_sharma_757fbb49aa7/showing-live-voice-agent-transcripts-in-a-flutter-app-with-livekit-d4b</guid>
      <description>&lt;p&gt;If you build a voice agent with LiveKit Agents and a Flutter frontend, at some point you want the conversation on screen: what the user said and what the agent is saying, updating live. The docs cover it, but the pieces are spread out and the older API in many examples is deprecated. Here's what works with the current SDKs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1b5i9tipzho1dbvuh53z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1b5i9tipzho1dbvuh53z.png" alt="How transcripts reach a Flutter app: LiveKit room, agent, lk.transcription text stream, handler, transcript map" width="800" height="410"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the text comes from
&lt;/h2&gt;

&lt;p&gt;When &lt;code&gt;AgentSession&lt;/code&gt; runs STT, it publishes transcriptions to the room on the &lt;code&gt;lk.transcription&lt;/code&gt; text stream topic. The agent's own speech goes out on the same topic, synced with audio playback, so it can show up word by word as the agent talks. If the user interrupts, the agent text is cut to what was actually spoken.&lt;/p&gt;

&lt;p&gt;The sender identity on each stream is the participant who was transcribed, so you can tell user and agent apart without extra metadata.&lt;/p&gt;

&lt;p&gt;The older &lt;code&gt;TranscriptionReceived&lt;/code&gt; event and &lt;code&gt;publish_transcription()&lt;/code&gt; are deprecated. They use a separate delivery path, so anything written to &lt;code&gt;lk.transcription&lt;/code&gt; never reaches a &lt;code&gt;TranscriptionEvent&lt;/code&gt; listener. Build on text streams.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Flutter handler
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight dart"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="s"&gt;'dart:convert'&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="n"&gt;transcript&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;String&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="kt"&gt;String&lt;/span&gt; &lt;span class="n"&gt;speaker&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;String&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;bool&lt;/span&gt; &lt;span class="n"&gt;isFinal&lt;/span&gt;&lt;span class="p"&gt;})&amp;gt;{};&lt;/span&gt;

&lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;listenForTranscripts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Room&lt;/span&gt; &lt;span class="n"&gt;room&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="n"&gt;me&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;room&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;localParticipant&lt;/span&gt;&lt;span class="o"&gt;?.&lt;/span&gt;&lt;span class="na"&gt;identity&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="n"&gt;room&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;registerTextStreamHandler&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;'lk.transcription'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;TextStreamReader&lt;/span&gt; &lt;span class="n"&gt;reader&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;String&lt;/span&gt; &lt;span class="n"&gt;participantIdentity&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="kd"&gt;async&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="n"&gt;info&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;reader&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;info&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;info&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;attributes&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;'lk.transcribed_track_id'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// chat message, not a transcript&lt;/span&gt;

    &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="n"&gt;segmentId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;info&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;attributes&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;'lk.segment_id'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="n"&gt;info&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="n"&gt;speaker&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;participantIdentity&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;me&lt;/span&gt; &lt;span class="o"&gt;?&lt;/span&gt; &lt;span class="s"&gt;'You'&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s"&gt;'Agent'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;var&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;''&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="n"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;bool&lt;/span&gt; &lt;span class="n"&gt;isFinal&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="n"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;segmentId&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nl"&gt;speaker:&lt;/span&gt; &lt;span class="n"&gt;speaker&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nl"&gt;text:&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nl"&gt;isFinal:&lt;/span&gt; &lt;span class="n"&gt;isFinal&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="n"&gt;onTranscriptChanged&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// setState, notifyListeners, etc.&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="n"&gt;reader&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;utf8&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// stream aborted, e.g. during a reconnect&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// read this after the stream closes: agent streams set it in the trailer&lt;/span&gt;
    &lt;span class="n"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;info&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;attributes&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;'lk.transcription_final'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s"&gt;'true'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The docs example uses &lt;code&gt;reader.readAll()&lt;/code&gt;. That waits for the stream to close, and an agent segment stays open until the agent stops talking, so the whole reply lands at once. Reading chunks gives you the word-by-word version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Interim vs final
&lt;/h2&gt;

&lt;p&gt;User speech and agent speech arrive differently. For user speech you get a new stream for each interim result, then a final one, all with the same &lt;code&gt;lk.segment_id&lt;/code&gt;. Key your map by that id and replace the entry every time, or the same sentence shows up three times.&lt;/p&gt;

&lt;p&gt;Agent speech is one stream per segment that grows as the agent talks. It opens with &lt;code&gt;lk.transcription_final&lt;/code&gt; set to &lt;code&gt;false&lt;/code&gt; and flips it to &lt;code&gt;true&lt;/code&gt; in the trailer when it closes, which is why the handler checks it after the loop.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1raq1rz3cj6ehoarq015.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1raq1rz3cj6ehoarq015.png" alt="Interim vs final transcription segments for user speech and agent speech" width="800" height="380"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you only want finished sentences, skip entries where &lt;code&gt;isFinal&lt;/code&gt; is false.&lt;/p&gt;

&lt;h2&gt;
  
  
  Register once
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;registerTextStreamHandler&lt;/code&gt; throws if a handler for that topic is already set. In a widget that can rebuild or reconnect, register once after connecting and call &lt;code&gt;room.unregisterTextStreamHandler('lk.transcription')&lt;/code&gt; in &lt;code&gt;dispose&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;If the transcript still feels laggy after this, look at end-of-turn detection on the agent side. I wrote about tuning that in &lt;a href="https://axionry.com/blog/voice-agent-turn-detection-barge-in" rel="noopener noreferrer"&gt;turn detection and barge-in for voice agents&lt;/a&gt;, and if you are estimating what a production voice agent costs per minute, there's a &lt;a href="https://axionry.com/tools/voice-ai-cost-calculator" rel="noopener noreferrer"&gt;voice AI cost calculator&lt;/a&gt; on our site.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The code was checked against the client-sdk-flutter source and the LiveKit docs. Originally published on &lt;a href="https://axionry.hashnode.dev/showing-live-voice-agent-transcripts-in-a-flutter-app-with-livekit" rel="noopener noreferrer"&gt;Axionry Engineering&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>flutter</category>
      <category>dart</category>
      <category>webrtc</category>
      <category>ai</category>
    </item>
    <item>
      <title>Estimate your AI feature's model bill before you build it</title>
      <dc:creator>Neeraj Sharma</dc:creator>
      <pubDate>Fri, 02 Oct 2026 11:42:19 +0000</pubDate>
      <link>https://dev.to/neeraj_sharma_757fbb49aa7/estimate-your-ai-features-model-bill-before-you-build-it-465f</link>
      <guid>https://dev.to/neeraj_sharma_757fbb49aa7/estimate-your-ai-features-model-bill-before-you-build-it-465f</guid>
      <description>&lt;p&gt;Most AI MVP budgets I see cover the build and forget the monthly line. That line is the model bill, and at a few thousand users it can pass what the build cost.&lt;/p&gt;

&lt;p&gt;One formula gets you a usable estimate.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9ddnz4n4t00p59i3h2hv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9ddnz4n4t00p59i3h2hv.png" alt="The monthly model bill formula with a worked example" width="800" height="380"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The formula
&lt;/h2&gt;

&lt;p&gt;monthly cost = requests per month x (input tokens x input price + output tokens x output price)&lt;/p&gt;

&lt;p&gt;Requests per month is users times how often they use the feature. Input tokens include everything you send: system prompt, retrieved context, chat history. Resending all of that on every request is why input usually dominates.&lt;/p&gt;

&lt;h2&gt;
  
  
  A worked example
&lt;/h2&gt;

&lt;p&gt;10,000 monthly users, 40 requests each, 4,000 input and 400 output tokens per request. That's 400,000 requests, 1.6B input tokens and 160M output tokens a month.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model (list price per 1M tokens)&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Monthly&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5 ($2 / $10)&lt;/td&gt;
&lt;td&gt;$3,200&lt;/td&gt;
&lt;td&gt;$1,600&lt;/td&gt;
&lt;td&gt;$4,800&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.1 Flash-Lite ($0.25 / $1.50)&lt;/td&gt;
&lt;td&gt;$400&lt;/td&gt;
&lt;td&gt;$240&lt;/td&gt;
&lt;td&gt;$640&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same feature, 7.5x apart.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fosdkmccx2thf1ikvq8x7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fosdkmccx2thf1ikvq8x7.png" alt="Monthly model bill for the same feature: Sonnet 5 $4,800, routed $1,472, Flash-Lite $640, Flash-Lite with caching $370" width="800" height="300"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Routing and caching
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Routing.&lt;/strong&gt; Most requests don't need the big model. Send the easy 80% to Flash-Lite and the hard 20% to Sonnet, and the $4,800 bill drops to about $1,470.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Caching.&lt;/strong&gt; If 3,000 of those 4,000 input tokens are the same system prompt and reference docs every time, cache them. Gemini 3.1 Flash-Lite bills cached input at $0.025 per 1M instead of $0.25, which takes the Flash-Lite bill from $640 to about $370 a month. Storage is $0.50 per 1M tokens per hour, so keeping a 3,000-token cache alive all month adds about $1.&lt;/p&gt;

&lt;p&gt;Both are build decisions. Adding them after launch means rewriting the request path.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do before writing code
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Write down users, requests per user and token counts per request, even rough ones.&lt;/li&gt;
&lt;li&gt;Run the formula for a big and a small model.&lt;/li&gt;
&lt;li&gt;If the big-model number scares you, design routing and caching in from day one.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you want the build cost and the run cost side by side, our &lt;a href="https://axionry.com/tools/ai-product-cost-estimator" rel="noopener noreferrer"&gt;AI product cost estimator&lt;/a&gt; does both, including the monthly bill at your user count. I also wrote up how we scope and price &lt;a href="https://axionry.com/services/ai-product-development" rel="noopener noreferrer"&gt;AI product development&lt;/a&gt; in checkpoints.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Prices are list prices from the Anthropic and Google pricing pages, October 2026.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>startup</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Our voice agent costs 2.5¢ a minute. Here's the whole bill.</title>
      <dc:creator>Neeraj Sharma</dc:creator>
      <pubDate>Fri, 02 Oct 2026 11:38:48 +0000</pubDate>
      <link>https://dev.to/neeraj_sharma_757fbb49aa7/our-voice-agent-costs-25c-a-minute-heres-the-whole-bill-3a17</link>
      <guid>https://dev.to/neeraj_sharma_757fbb49aa7/our-voice-agent-costs-25c-a-minute-heres-the-whole-bill-3a17</guid>
      <description>&lt;p&gt;We run an AI interviewer that talks to candidates over phone calls. It started on a managed voice platform at roughly 10¢ a minute. The same interviews now run on a custom LiveKit pipeline at about 2.5¢ a minute, all in. For reference, Retell's pricing calculator defaults to 11¢ a minute and LiveKit Cloud's own calculator to about 4.8¢.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsnqq8dx0ran7sg3j463d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsnqq8dx0ran7sg3j463d.png" alt="Phone voice agent architecture: Telnyx SIP, LiveKit room, agent worker with VAD, turn detector, xAI STT, Gemini 2.5 Flash-Lite, xAI TTS" width="800" height="470"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The stack
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Telnyx SIP trunk into LiveKit SIP&lt;/li&gt;
&lt;li&gt;Silero VAD plus LiveKit's turn detector (v1-mini) running on CPU&lt;/li&gt;
&lt;li&gt;xAI streaming STT, Gemini 2.5 Flash-Lite, xAI TTS&lt;/li&gt;
&lt;li&gt;Self-hosted agent workers and Redis for session state&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where the money goes
&lt;/h2&gt;

&lt;p&gt;Per minute of an inbound call, at published list prices, assuming 600 characters of agent speech and about 3,000 input / 175 output tokens a minute:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Line item&lt;/th&gt;
&lt;th&gt;Rate&lt;/th&gt;
&lt;th&gt;Per minute&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;xAI TTS&lt;/td&gt;
&lt;td&gt;$15 per 1M characters&lt;/td&gt;
&lt;td&gt;$0.0090&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LiveKit SIP fee (Ship plan)&lt;/td&gt;
&lt;td&gt;$0.004 per minute&lt;/td&gt;
&lt;td&gt;$0.0040&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;xAI STT (streaming)&lt;/td&gt;
&lt;td&gt;$0.20 per hour&lt;/td&gt;
&lt;td&gt;$0.0033&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Telnyx inbound local&lt;/td&gt;
&lt;td&gt;$0.0032 per minute&lt;/td&gt;
&lt;td&gt;$0.0032&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash-Lite&lt;/td&gt;
&lt;td&gt;$0.10 in / $0.40 out per 1M tokens&lt;/td&gt;
&lt;td&gt;$0.0004&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Turn detection (local CPU)&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;$0.0000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workers, Redis, logs&lt;/td&gt;
&lt;td&gt;self-hosted&lt;/td&gt;
&lt;td&gt;~$0.0051&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$0.025&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9clvdgzta3ua5pao9q2x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9clvdgzta3ua5pao9q2x.png" alt="Per-minute cost breakdown of a 2.5 cent voice agent" width="800" height="380"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A few things I didn't expect:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TTS is the biggest line.&lt;/strong&gt; Almost a cent a minute, billed by character. Shorter agent turns save more than any model swap would.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The LLM barely matters for cost.&lt;/strong&gt; Four hundredths of a cent. Pick it for latency and quality, not price.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rate cards only get you to 2¢.&lt;/strong&gt; The last half cent is workers, Redis and logging. No vendor lists that.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhd0p2khpizbw1rbohztx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhd0p2khpizbw1rbohztx.png" alt="Cents per minute: Retell default 11, old managed platform 10, LiveKit Cloud default 4.79, custom pipeline 2.5" width="800" height="280"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  When it isn't worth building
&lt;/h2&gt;

&lt;p&gt;On per-minute cost alone, building wins easily. Count engineer time and it gets closer. With about $600 a month of fixed infra and two engineer-days a month of upkeep (at $150k a year, 260 working days), a custom pipeline only beats an 11¢ vendor above roughly 20,600 minutes a month. Below that, stay on a managed platform and spend the time on the product.&lt;/p&gt;

&lt;p&gt;The full migration, with the numbers behind the 75% cut, is in our &lt;a href="https://axionry.com/case-studies/voice-ai-cost-2-5-cents" rel="noopener noreferrer"&gt;2.5¢ voice AI case study&lt;/a&gt;. If you want to run your own volumes, there's a &lt;a href="https://axionry.com/tools/voice-ai-cost-calculator" rel="noopener noreferrer"&gt;voice AI cost calculator&lt;/a&gt; on our site.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Rates are from the xAI, Google, Telnyx and LiveKit pricing pages. Originally published on &lt;a href="https://axionry.hashnode.dev/our-voice-agent-costs-2-5-a-minute-here-s-the-whole-bill" rel="noopener noreferrer"&gt;Axionry Engineering&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>voiceai</category>
      <category>webrtc</category>
      <category>llm</category>
    </item>
    <item>
      <title>Why your NAT gateway bill keeps growing (and the 20 minute fix)</title>
      <dc:creator>Neeraj Sharma</dc:creator>
      <pubDate>Fri, 02 Oct 2026 08:22:14 +0000</pubDate>
      <link>https://dev.to/neeraj_sharma_757fbb49aa7/why-your-nat-gateway-bill-keeps-growing-and-the-20-minute-fix-36na</link>
      <guid>https://dev.to/neeraj_sharma_757fbb49aa7/why-your-nat-gateway-bill-keeps-growing-and-the-20-minute-fix-36na</guid>
      <description>&lt;p&gt;If your workloads run in private subnets, a big part of your NAT gateway bill is probably traffic to S3. The fix takes a few minutes and costs nothing.&lt;/p&gt;

&lt;p&gt;NAT gateway pricing in us-east-1 is $0.045 per hour plus $0.045 per GB processed. The hourly part is small. The per-GB part grows with every image pull, backup and log upload that leaves a private subnet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq82yy2r1aa7svl6u15kx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq82yy2r1aa7svl6u15kx.png" alt="S3 traffic path from a private subnet, through a NAT gateway versus an S3 gateway endpoint" width="800" height="410"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Check if this is you
&lt;/h2&gt;

&lt;p&gt;In Cost Explorer, filter Service to "EC2-Other" and group by Usage Type. Look for &lt;code&gt;NatGateway-Bytes&lt;/code&gt;. In us-east-1 it has no prefix; other regions get one, like &lt;code&gt;USW2-NatGateway-Bytes&lt;/code&gt;. If it's near the top, keep reading.&lt;/p&gt;

&lt;p&gt;To see where the bytes go, turn on VPC Flow Logs for the NAT gateway's network interface and sum bytes by destination. If most destinations fall in the S3 ranges from &lt;code&gt;ip-ranges.amazonaws.com/ip-ranges.json&lt;/code&gt;, you found it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 1: S3 and DynamoDB gateway endpoints (free)
&lt;/h2&gt;

&lt;p&gt;Gateway endpoints for S3 and DynamoDB have no hourly or data charge. Create one and attach your private route tables:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws ec2 create-vpc-endpoint &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--vpc-id&lt;/span&gt; vpc-0abc123 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--vpc-endpoint-type&lt;/span&gt; Gateway &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--service-name&lt;/span&gt; com.amazonaws.us-east-1.s3 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--route-table-ids&lt;/span&gt; rtb-0priv1 rtb-0priv2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It only covers buckets in the same region. Creating it drops open TCP connections to S3, so don't do it mid-backup, and check that your clients reconnect. S3 will also see your private IPs instead of the NAT's public IP, so bucket policies that allow by source IP need updating. ECR pulls get cheaper too, because image layers are served from S3.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 2: interface endpoints, but only above break-even
&lt;/h2&gt;

&lt;p&gt;ECR API, STS, CloudWatch Logs and most other services need interface endpoints instead. These are not free: $0.01 per hour per AZ plus $0.01 per GB. Across two AZs that is $14.60 a month (2 x 730 hours x $0.01) before any traffic.&lt;/p&gt;

&lt;p&gt;Every GB you move off NAT saves $0.035 ($0.045 minus $0.01). So one endpoint in two AZs pays for itself at about 420 GB a month. Below that, leave the traffic on NAT.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3aya4u4nscocr7ef70c5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3aya4u4nscocr7ef70c5.png" alt="Monthly cost of NAT processing versus an interface endpoint in two AZs, break-even at about 417 GB" width="800" height="410"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 3: watch cross-AZ traffic to the NAT
&lt;/h2&gt;

&lt;p&gt;With one NAT gateway for the whole VPC, traffic from the other AZs pays inter-AZ transfer just to reach it: $0.01/GB in each direction, so $0.02 for every GB that crosses. A second NAT gateway costs about $33 a month in hours, so it pays off once roughly 1.6 TB a month crosses from that AZ.&lt;/p&gt;

&lt;p&gt;Do fix 1 first. It's free and it's usually the biggest chunk.&lt;/p&gt;

&lt;p&gt;The full list of leaks I check (NAT, egress, idle EBS, public IPv4 and more) is in this &lt;a href="https://axionry.com/blog/aws-cost-optimization-checklist" rel="noopener noreferrer"&gt;AWS cost optimization checklist&lt;/a&gt;, and there's a free &lt;a href="https://axionry.com/tools/cloud-cost-calculator" rel="noopener noreferrer"&gt;cloud cost calculator&lt;/a&gt; if you want to plug in your own numbers.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Prices are us-east-1 list prices.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>devops</category>
      <category>cloud</category>
      <category>finops</category>
    </item>
  </channel>
</rss>
