DEV Community

RAXXO Studios
RAXXO Studios

Posted on Originally published at raxxo.shop

Eleven v4 vs v4 Turbo vs v3: Which to Use

  • Eleven v4 shipped September 28, 2026 and entered the Artificial Analysis leaderboard at #1 of 92 with 1319 Elo

  • v4 sits 150 Elo above Eleven v3 and doubles the per-generation limit to 10,000 characters

  • Use eleven_v4 for produced voiceover and eleven_v4_turbo for live agents; Flash v2.5 only wins on price

  • Before moving a live pipeline, render one paragraph with your clone on v3 and v4 and compare on headphones

Eleven v4 is the new text to speech model from ElevenLabs, released September 28, 2026, in two variants: eleven_v4 for produced audio and eleven_v4_turbo for real-time voice agents. By the afternoon of launch day it was sitting at #1 on the Artificial Analysis text to speech leaderboard, ahead of Cartesia, Google and 89 other models.

I haven't run v4 through my own listening test yet. So every number below is either ElevenLabs's own or Artificial Analysis's, and I say which. The decision at the end of each section is mine.

What Eleven v4 Changes Compared to v3

ElevenLabs describes v4 as a new architecture with "higher audio quality and a wider emotional range." Every launch page says something like that. The line I actually care about comes right after it: "Speaker identity is now stable across regenerations."

If you've produced anything long on Eleven v3, you know why that sentence is there. Regenerate one bad line and the replacement could sound like a close cousin of your narrator. On headphones it's obvious. Viewers pick up on it too, even when they can't say what's off.

That fix alone would get me to switch.

The rest of the list is solid, if less exciting. The model docs put the per-generation limit at 10,000 characters for eleven_v4 against 5,000 for eleven_v3, which is roughly 10 minutes of audio in one pass. Past that, ElevenLabs says "context stitching" keeps the pacing consistent between generations. Language support goes to 90+, up from the 70+ v3 launched with, and one voice can switch languages while keeping its identity.

Professional Voice Clones work on v4. ElevenLabs goes further and claims its Instant Voice Clones, made from 10 seconds of audio, "now outperform the Professional Voice Clones of Multilingual v2." That's a vendor claim about the thing they sell, so I'd test it before cancelling any PVC work.

Audio tags carry over from v3 and are supposed to land more reliably now. The usual [laughs], [whispers], [sighs] and [long pause] are there, plus scene cues like [door slams] or [light rain], and you can direct delivery in plain words, for example [said angrily in French accent]. IPA handling in pronunciation dictionaries got better too, which matters if you've got brand names in your scripts.

All 17,500+ voices in the voice library work with v4. Nobody has to rebuild a voice roster just to try it.

Eleven v4 vs v4 Turbo vs v3: The Numbers

Here's the side-by-side on launch day. Ranks and Elo come from the Artificial Analysis Speech Arena, which scores models from blind pairwise listener votes. Everything else comes from ElevenLabs's docs and launch page.

Eleven v4 Eleven v4 Turbo Eleven v3
Model ID eleven_v4 eleven_v4_turbo eleven_v3
Built for Produced content Voice agents, real time Expressive produced content
Artificial Analysis rank #1 of 92 (1319 Elo) Not listed yet #18 (1169 Elo)
Characters per generation 10,000 Not published 5,000
Latency (vendor figure) Not the goal ~100 ms inference, ~150 ms to first speech Not a real-time model
Languages 90+ 90+ 70+

The leaderboard gap is bigger than it looks. Cartesia Sonic 3.6 is second at 1276, Google's Gemini 3.8 Flash TTS third at 1267. From #2 down to #10 the whole spread is 70 points, and v4 opened a 43-point lead over #2 on day one.

It buries ElevenLabs's own back catalogue as well. Eleven v3 Conversational is #13 with 1197. Multilingual v2, which plenty of production pipelines still run on, is #38 with 1094, and Flash v2.5 is #45 with 1076. Going from Multilingual v2 straight to v4 is a 225-point jump.

For latency, ElevenLabs published its own time-to-first-speech chart: 150 ms for v4 Turbo, 262 ms for Cartesia Sonic 3.6, 814 ms for OpenAI's GPT-4o mini TTS. Those are lab numbers. Your network sits on top.

Price is the part I can't pin down yet. The pricing page doesn't list a v4 credit rate, and the launch page only says v4 uses the same credit system as the other TTS models. Artificial Analysis's price column puts v4 about 20% below v3 per million characters and about 60% above Flash v2.5. If that holds, v3 to v4 is an upgrade that also costs less. I don't see that combination often.

The free plan gives you 10,000 credits a month, around 10 minutes of audio, no card required. Enough for every test in the last section.

Which Model ID to Use for Which Job

Switching is one parameter. Picking the value is the actual work.

Voiceover and audiobooks: eleven_v4

This is the job v4 was built for, and the 10,000-character ceiling changes how you cut a script. A 12-minute episode at roughly 1,000 characters per minute is two generations on v4 and three on v3. One seam fewer to listen for. I wrote up my episode structure in ElevenLabs Studio Workflow: 4 Patterns for 12-Minute Solo Episodes, and all four patterns still apply. You just split less often.

Voice agents and phone bots: eleven_v4_turbo

Turbo streams in both directions, so text goes in while audio is already coming out. The v4 family also outputs ยต-law for telephony next to MP3 and WAV/PCM. ElevenLabs says Turbo keeps the full expressive range of v4, which means an agent can laugh or pause on cue instead of reading everything in one flat register.

There's a naming trap here. eleven_v4_turbo is not eleven_turbo_v2_5. The ElevenLabs docs still say "We recommend using the Flash models over Turbo models in all use cases," and that line is about the old v2.5 Turbo. It doesn't apply to v4 Turbo. Anyone skimming the docs, human or model, can mix the two up.

Short-form hooks and scored shorts: eleven_v4 with audio tags

On a 30-second short, the tags are the whole point. A [whispers] on the hook line and [light rain] under the setup used to mean a voice pass and then a separate sound pass. Now both live in the script. You still want a real music bed, and The Viral AI Sound: How Creators Score Videos in 2026 covers how I layer generated sound against trending audio.

Thousands of cheap lines: Flash v2.5, for now

Game barks and notification lines, say, where every clip gets heard once and "clear enough" is the bar. By the Artificial Analysis numbers Flash v2.5 is still cheaper per character. I'd revisit that the day ElevenLabs publishes a credit rate for v4 Turbo. If it lands near Flash, the case for Flash gets thin fast.

For how ElevenLabs compares to the rest of the field beyond this launch, see ElevenLabs vs Other AI Voice Tools: An Honest Comparison.

What I Would Re-Test Before Switching

Arena votes come from strangers listening to short clips. Your pipeline has your own voice and your product names in it. Swapping model_id takes ten seconds. Finding out three weeks later that a product name sounds wrong in 40 published videos takes a lot longer.

So, before anything live moves over:

  1. Re-run your tagged scripts. ElevenLabs says v4 follows audio tags more reliably, and the flip side is that a tag v3 quietly ignored might now fire. Search your scripts for leftovers before you batch-render.

  2. Render one reference paragraph with your clone on both models. Same text, same voice, back to back on headphones. If you depend on a PVC, this is also where you find out whether the 10-second Instant Clone claim holds for your voice. Some voices clone better than others.

  3. Play every entry in your pronunciation dictionary. RAXXO is the first word I'd check, because invented names are exactly where a new model guesses differently.

  4. Listen across the seam. Put the last sentence of part one next to the first sentence of part two and see if the energy drops.

  5. Measure latency from where your servers actually run. The 150 ms figure is ElevenLabs's. Time to first audio byte from your region is what your users feel.

  6. Set model_id explicitly in every call. Otherwise a future default change swaps your narrator without anyone touching the code.

Number 2 is the one I wouldn't skip.

Bottom Line

For produced audio, I'd make Eleven v4 the default today. The stable speaker identity is the reason, more than the #1 rank. For live agents, eleven_v4_turbo is the one to test, and the ~150 ms claim is worth checking on your own network before you build around it.

I wouldn't start a new project on v3. The only reason I can see to stay is matching old renders you can't re-record.

The free tier's 10,000 monthly credits cover every test above. Once the voiceover is locked, the video still needs a title, a caption, hashtags and a music direction. That's what RAXXO Studio does: upload the finished clip and it writes those for whichever platform the clip goes to.

Which would you move first, the long-form narration or the live agent?

This article contains affiliate links. If you sign up through them, I may earn a small commission at no extra cost to you. (Ad)

Top comments (0)