When we shipped live captions for TranscribeChirp, the first instinct was tempting:
Put a WebSocket on the Next.js API route, pipe mic audio through Vercel, fan out to Deepgram.
That path fails on serverless. Here is the split we ended up with — and why it is a better default for any “live while the user is still talking” feature on Vercel.
The constraint
Vercel (and similar platforms) want request → response. A live mic stream is minutes of bidirectional bytes. If you terminate that socket on a serverless function you get:
- cold starts mid-session
- hard execution time limits
- awkward billing for idle keep-alives
- one more place your API keys can leak into long-lived connections
So the rule we wrote into our template: do not proxy long-lived audio (or high-frequency events) through the app server. Live bytes belong on a provider edge or a dedicated long-running Node process.
The pattern that works
Browser (logged in)
│
├─ POST /api/.../live/session → short-lived provider token + job id
│
└─ WebSocket ─────────────────→ Deepgram (or other live ASR)
│
▼ partial captions
React UI
│
POST /api/.../live/meter (optional, every few seconds)
POST /api/.../live/finalize → persist, settle credits, close job
Four jobs, clear owners:
- Your API — auth, credit pre-check, create a job row, mint a scoped, short-TTL token.
- Provider — browser connects directly; partial transcripts stream back.
- Your API (optional) — idempotent mid-session metering so you can warn or cut off before the wallet hits zero.
- Your API — finalize: save the transcript, charge the last partial minute, mark the job done.
Batch upload / URL import / “record then upload” stay untouched. Live is an additive button, not a rewrite of the async path.
What the session endpoint returns
Keep the mint response boring and small:
// shape only — names vary by provider
type LiveSession = {
jobId: string;
token: string; // short TTL, scoped to this session
expiresAt: string;
// provider endpoint / model hints as needed
};
Never ship the long-lived master API key to the browser. Never let the token outlive the session by much. If the tab dies, finalize (or a sweeper) closes the job so credits do not leak.
Metering without owning the socket
We still need fair billing. For live we use a simple floor/ceil pattern:
- Pre-check: user needs at least one minute’s worth of credits before start.
- Mid-session ticks (~5s): charge by floor minutes so far (zero until 60s).
- Finalize: ceil the last partial minute (same idea as batch).
Because metering is HTTPS posts from the client (or a trusted finalize), the WebSocket can stay on the provider. Your database remains the source of truth for “how many billable minutes did this job consume?”
When you do need your own Node process
Provider-direct is enough for speech-to-text and many chat streams. Reach for Fly / Railway / a VM when you need:
- custom audio mixing or VAD before the vendor sees bytes
- multi-party fan-in that no single provider session covers
- a protocol the browser cannot speak safely with a short token
Until then, mint + direct connect is less ops and fewer failure modes.
Takeaway
If the product promise is “text appears while they are still speaking,” treat Vercel as the control plane (auth, jobs, credits) and the ASR vendor as the data plane (audio + partials). That split kept our live path shippable next to batch without pretending serverless is a always-on media relay.
We run this for live transcription and live translation (captions plus sentence-level translate) on TranscribeChirp. If you are wiring the same shape on Next.js, the mental model above is the part worth stealing — not the brand name.
Questions / sharper designs welcome in the comments.
Top comments (0)