DEV Community

Autor Technologies Inc.
Autor Technologies Inc.

Posted on

4 Months of vox-bench in Production — What 14,000 Benchmark Runs Taught Us About Voice AI Latency

In June, we open-sourced vox-bench, the latency benchmarking tool we built after a TTS spike caused 12% of Loquent callers to talk over our AI mid-response. Four months and 14,000 automated benchmark runs later, we have data on voice AI provider latency that nobody else is publishing — and some of our original assumptions were wrong.

Quick context if you missed the original release

Loquent is our production voice AI platform. It handles thousands of automated phone calls per month for healthcare and dental clients across Canada — appointment booking, patient intake, after-hours triage. The pipeline is Twilio (media stream) → Deepgram (STT) → Anthropic Claude (LLM) → ElevenLabs (TTS) → Twilio (audio back).

We built vox-bench because we needed per-stage, per-percentile latency tracking with regression detection. It runs as a TypeScript CLI, benchmarks each pipeline stage independently and in combination, and fires alerts when p95 latency at any stage crosses your configured threshold. We've been running it via GitHub Actions every 6 hours since May.

The tool is at github.com/Autor-Technologies/vox-bench. This article is about what the data told us.

Assumption #1 we got wrong: TTS is NOT the most variable stage anymore

When we released vox-bench in June, our data clearly showed TTS as the most unpredictable pipeline stage. ElevenLabs had a p99/p50 ratio of 2.69x — meaning the worst 1% of requests took nearly three times longer than the median. We built our alerting thresholds around that assumption.

Four months later, the picture has shifted. Here's the p99/p50 variability ratio across our current stack, aggregated from 14,000 runs:

  • Deepgram Nova-3 STT: 1.72x (was 1.72x in June — rock solid, no change)
  • Anthropic Claude Haiku LLM: 2.85x (was 2.29x in June — got MORE variable)
  • ElevenLabs Turbo v2.5 TTS: 1.92x (was 2.69x — significantly improved)

The LLM layer is now our most variable stage. Not because Claude got worse — the median actually dropped from 210ms to 185ms TTFB over the same period. But the tail got fatter. Our p99 went from 480ms to 525ms. We think this correlates with model updates that improved average performance but introduced more variance under load.

ElevenLabs, meanwhile, quietly got much more consistent. Their Turbo v2.5 release in July cut our p99 from 350ms to 255ms and the time-of-day pattern we reported in June (2-4pm ET spikes) mostly disappeared. Credit where it's due — their infrastructure team clearly addressed capacity issues.

The takeaway: don't set your alerting thresholds once and forget them. Provider performance characteristics change faster than you think. We now re-baseline our vox-bench thresholds monthly.

The 680ms wall

After 14,000 runs and correlating benchmark data with actual Loquent call logs, we've identified what we internally call "the 680ms wall." Here's what it is.

When total pipeline round-trip (caller stops speaking → AI audio starts playing) stays below 680ms, our call completion rate is 94%. The caller doesn't notice the delay. Conversations flow. When we cross 680ms, completion drops to 87%. Cross 1,100ms and it falls to 71%.

These aren't theoretical benchmarks. This is from 8 months of production call data across 40,000+ calls, correlated with the per-call latency measurements vox-bench introduced.

What surprised us: the drop isn't linear. There's a genuine cliff between 680ms and 700ms. Below the wall, callers behave normally. Above it, they start exhibiting what we call "confirmation seeking" — saying "hello?" or repeating their last sentence. Once a caller enters that mode, the rest of the call degrades even if subsequent responses are fast. The caller has lost trust in the AI's responsiveness.

We published our original article saying the threshold was 800ms. We were wrong. It's lower, and the penalty for crossing it is harsher than we expected.

Provider changes we made based on the data

vox-bench includes a provider comparison mode — run the same benchmark workload against multiple providers simultaneously. We've been using it continuously to evaluate alternatives. Three changes we made:

STT: Stayed with Deepgram Nova-3. We evaluated OpenAI Whisper (API) and Google Cloud Speech-to-Text v2 as alternatives. Whisper's accuracy on medical terminology was 3-4% better than Deepgram, but its p50 latency was 340ms versus Deepgram's 175ms. That 165ms difference eats almost our entire latency budget at the STT stage. Google was competitive on latency (195ms p50) but struggled with Canadian English accents — our word error rate was 18% higher on calls from patients in Quebec and Atlantic Canada.

LLM: Added a fast-path for simple responses. vox-bench data showed that 62% of our LLM calls were for straightforward responses — confirming appointment times, reading back addresses, answering "what are your hours?" These don't need Claude's full reasoning. We added a response cache and a classifier that routes simple queries to cached templates, hitting the LLM only for complex or contextual responses. Result: those 62% of responses now have 0ms LLM latency. Our blended p50 total pipeline time dropped from 620ms to 410ms.

TTS: Tested but didn't switch from ElevenLabs. We benchmarked Google Cloud TTS, Amazon Polly, and OpenAI TTS. Google had the lowest latency (90ms p50 vs ElevenLabs' 125ms), but every provider we tested had an uncanny quality gap compared to ElevenLabs on healthcare vocabulary. Pronouncing medication names, clinic addresses, and doctor names correctly matters when your caller is a patient. ElevenLabs' custom pronunciation dictionary handles this better than any alternative we tested.

The data pattern nobody warned us about

Here's something we didn't expect to find in four months of continuous benchmarking: provider latency is seasonal. Not in the weather sense — in the tech industry calendar sense.

We have 14,000 data points across May through September. The trend is clear:

  • Late June through July: All three providers (Deepgram, Claude, ElevenLabs) showed 10-15% lower latency than their April-May baselines. Our theory: lighter API traffic during summer.
  • Early September: Latency climbed 20% above baseline across the board. Every provider, simultaneously. This coincided with the post-Labor Day tech industry return-to-work surge.
  • Tuesday through Thursday consistently runs 8-12% slower than weekends across all providers.

If you're benchmarking voice AI providers and running your evaluation on a Saturday afternoon, you're getting a misleadingly optimistic picture. Benchmark during peak hours (Tuesday-Thursday, 10am-3pm ET) or your production experience will be worse than your evaluation predicted.

We've added time-of-week bucketing to vox-bench's reporting to surface this automatically.

What 47 GitHub issues taught us about other people's voice AI stacks

Since open-sourcing vox-bench, we've had contributions from teams building voice AI across different domains — customer service, real estate, legal intake, restaurant ordering. Their benchmark data (shared voluntarily via issues and PRs) revealed patterns we wouldn't have seen from healthcare alone:

Restaurant and retail voice AI runs 25-30% faster end-to-end than healthcare. The vocabulary is simpler, utterances are shorter ("I'd like to order a large pepperoni pizza" vs. "I need to reschedule my appointment with Dr. Krishnamurthy for my follow-up on the blood work from last Tuesday"), and there's less ambiguity for the LLM to resolve. If you're reading voice AI latency benchmarks from a non-healthcare context, discount them for medical applications.

The STT bottleneck shifts by language. Two contributors benchmarking Spanish-language voice AI found that Deepgram's Spanish model runs 40% slower than English Nova-3 on equivalent utterance lengths. One switched to a self-hosted Whisper model and got better latency AND accuracy for Spanish. We added multi-language benchmark profiles to vox-bench based on their work.

WebSocket connection reuse matters more than we realized. A contributor profiling their pipeline found that 180ms of their "STT latency" was actually WebSocket connection establishment — they were opening a new connection per utterance instead of keeping a persistent stream. vox-bench now measures connection overhead separately from processing time.

Updated numbers: September 2026 production stack

For anyone comparing their own pipeline, here are our current production numbers from vox-bench, benchmarked Tuesday-Thursday 10am-3pm ET (worst-case real-world conditions):

Stage p50 p95 p99
Deepgram Nova-3 STT 175ms 240ms 300ms
Claude Haiku LLM (complex) 185ms 355ms 525ms
Response cache (simple) <1ms <1ms <1ms
ElevenLabs Turbo v2.5 TTS 125ms 195ms 255ms
Total (complex path) 580ms 870ms 1,090ms
Total (cached path) 365ms 510ms 630ms
Blended (62% cached) 410ms 645ms 800ms

Our blended p50 is now comfortably below the 680ms wall. The complex path p95 still crosses it, which means roughly 5% of complex-response calls have a noticeable delay. We're okay with that — pushing it lower would require sacrificing response quality, and our completion rate data shows 5% above the wall is an acceptable tradeoff.

Key findings after 4 months

  1. Re-baseline your latency thresholds monthly. Provider performance characteristics change with model updates, infrastructure changes, and traffic patterns. The thresholds you set in month one will be wrong by month three.

  2. The conversational latency wall is 680ms, not 800ms. Our original estimate was too generous. Below 680ms, callers don't notice. Above it, call quality degrades non-linearly.

  3. Cache simple responses aggressively. 62% of our voice AI calls don't need LLM inference at all. A response cache with a lightweight classifier cut our blended latency by 34%.

  4. Benchmark during peak hours or your data is lying. Tuesday-Thursday, 10am-3pm gives you realistic production numbers. Weekend benchmarks are 15-20% more optimistic than what your users will experience.

  5. Provider variability shifts over time. TTS was our most variable stage in June. Four months later, it's the LLM. Continuous benchmarking catches these shifts before your callers do.

What's next for vox-bench

We're working on two additions for the next release. First, a Grafana dashboard template that visualizes vox-bench data over time — right now you have to build your own or use the JSON/Markdown output. Second, a "latency budget calculator" that takes your target total round-trip time and suggests per-stage allocations based on current provider performance data.

vox-bench is at github.com/Autor-Technologies/vox-bench. If you're running it against your own stack, we'd genuinely like to see your numbers — open an issue or PR with anonymized benchmark data and we'll incorporate it into the cross-domain analysis.

If you're building something similar, we'd love to hear about it. Reach out at hello@autor.ca or visit autor.ca.

Top comments (0)