<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Yvone</title>
    <description>The latest articles on DEV Community by Yvone (@yvone_f8de85837dee3e4cd0f).</description>
    <link>https://dev.to/yvone_f8de85837dee3e4cd0f</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149976%2F0365f5e4-a1e0-4622-9eb0-596f0a752d98.png</url>
      <title>DEV Community: Yvone</title>
      <link>https://dev.to/yvone_f8de85837dee3e4cd0f</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yvone_f8de85837dee3e4cd0f"/>
    <language>en</language>
    <item>
      <title>Why we don’t proxy live audio WebSockets through Vercel (and what we do instead)</title>
      <dc:creator>Yvone</dc:creator>
      <pubDate>Tue, 29 Sep 2026 14:33:15 +0000</pubDate>
      <link>https://dev.to/yvone_f8de85837dee3e4cd0f/why-we-dont-proxy-live-audio-websockets-through-vercel-and-what-we-do-instead-7go</link>
      <guid>https://dev.to/yvone_f8de85837dee3e4cd0f/why-we-dont-proxy-live-audio-websockets-through-vercel-and-what-we-do-instead-7go</guid>
      <description>&lt;p&gt;When we shipped live captions for &lt;a href="https://www.transcribechirp.online" rel="noopener noreferrer"&gt;TranscribeChirp&lt;/a&gt;, the first instinct was tempting:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Put a WebSocket on the Next.js API route, pipe mic audio through Vercel, fan out to Deepgram.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That path fails on serverless. Here is the split we ended up with — and why it is a better default for any “live while the user is still talking” feature on Vercel.&lt;/p&gt;

&lt;h2&gt;
  
  
  The constraint
&lt;/h2&gt;

&lt;p&gt;Vercel (and similar platforms) want &lt;strong&gt;request → response&lt;/strong&gt;. A live mic stream is &lt;strong&gt;minutes of bidirectional bytes&lt;/strong&gt;. If you terminate that socket on a serverless function you get:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;cold starts mid-session&lt;/li&gt;
&lt;li&gt;hard execution time limits&lt;/li&gt;
&lt;li&gt;awkward billing for idle keep-alives&lt;/li&gt;
&lt;li&gt;one more place your API keys can leak into long-lived connections&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the rule we wrote into our template: &lt;strong&gt;do not proxy long-lived audio (or high-frequency events) through the app server&lt;/strong&gt;. Live bytes belong on a provider edge or a dedicated long-running Node process.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern that works
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Browser (logged in)
  │
  ├─ POST /api/.../live/session  → short-lived provider token + job id
  │
  └─ WebSocket ─────────────────→ Deepgram (or other live ASR)
         │
         ▼ partial captions
  React UI
         │
  POST /api/.../live/meter      (optional, every few seconds)
  POST /api/.../live/finalize   → persist, settle credits, close job
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four jobs, clear owners:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Your API&lt;/strong&gt; — auth, credit pre-check, create a job row, mint a &lt;strong&gt;scoped, short-TTL&lt;/strong&gt; token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Provider&lt;/strong&gt; — browser connects directly; partial transcripts stream back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your API (optional)&lt;/strong&gt; — idempotent mid-session metering so you can warn or cut off before the wallet hits zero.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your API&lt;/strong&gt; — finalize: save the transcript, charge the last partial minute, mark the job done.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Batch upload / URL import / “record then upload” stay untouched. Live is an &lt;strong&gt;additive&lt;/strong&gt; button, not a rewrite of the async path.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the session endpoint returns
&lt;/h2&gt;

&lt;p&gt;Keep the mint response boring and small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// shape only — names vary by provider&lt;/span&gt;
&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;LiveSession&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;jobId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;token&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;      &lt;span class="c1"&gt;// short TTL, scoped to this session&lt;/span&gt;
  &lt;span class="nl"&gt;expiresAt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="c1"&gt;// provider endpoint / model hints as needed&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Never ship the long-lived master API key to the browser. Never let the token outlive the session by much. If the tab dies, finalize (or a sweeper) closes the job so credits do not leak.&lt;/p&gt;

&lt;h2&gt;
  
  
  Metering without owning the socket
&lt;/h2&gt;

&lt;p&gt;We still need fair billing. For live we use a simple floor/ceil pattern:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pre-check: user needs at least one minute’s worth of credits before start.&lt;/li&gt;
&lt;li&gt;Mid-session ticks (~5s): charge by &lt;strong&gt;floor&lt;/strong&gt; minutes so far (zero until 60s).&lt;/li&gt;
&lt;li&gt;Finalize: &lt;strong&gt;ceil&lt;/strong&gt; the last partial minute (same idea as batch).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because metering is HTTPS posts from the client (or a trusted finalize), the WebSocket can stay on the provider. Your database remains the source of truth for “how many billable minutes did this job consume?”&lt;/p&gt;

&lt;h2&gt;
  
  
  When you &lt;em&gt;do&lt;/em&gt; need your own Node process
&lt;/h2&gt;

&lt;p&gt;Provider-direct is enough for speech-to-text and many chat streams. Reach for Fly / Railway / a VM when you need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;custom audio mixing or VAD before the vendor sees bytes&lt;/li&gt;
&lt;li&gt;multi-party fan-in that no single provider session covers&lt;/li&gt;
&lt;li&gt;a protocol the browser cannot speak safely with a short token&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Until then, mint + direct connect is less ops and fewer failure modes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;If the product promise is “text appears while they are still speaking,” treat Vercel as the &lt;strong&gt;control plane&lt;/strong&gt; (auth, jobs, credits) and the ASR vendor as the &lt;strong&gt;data plane&lt;/strong&gt; (audio + partials). That split kept our live path shippable next to batch without pretending serverless is a always-on media relay.&lt;/p&gt;

&lt;p&gt;We run this for live transcription and live translation (captions plus sentence-level translate) on &lt;a href="https://www.transcribechirp.online" rel="noopener noreferrer"&gt;TranscribeChirp&lt;/a&gt;. If you are wiring the same shape on Next.js, the mental model above is the part worth stealing — not the brand name.&lt;/p&gt;

&lt;p&gt;Questions / sharper designs welcome in the comments.&lt;/p&gt;

</description>
      <category>nextjs</category>
      <category>webdev</category>
      <category>javascript</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
