ElevenLabs raises the bar for expressive AI speech

ElevenLabs is rolling out Eleven v4, a new speech model that the company says tracks directorial cues more accurately than its predecessor and keeps a narrator's or character's voice stable across long-form productions. Alongside it, the company is shipping a faster variant aimed squarely at real-time applications, where responsive voice agents have historically had to compromise on expression.

The release lands roughly a year after Eleven v3, which already introduced support for audio tags but followed those instructions with less precision. With v4, ElevenLabs is betting that the difference between a usable synthetic voice and a genuinely performative one comes down to how faithfully a model interprets the directions a writer or producer gives it — and how little that performance drifts over the course of an audiobook, a dubbing project, or a long interactive session.

Direction cues become far more dependable

The core promise of Eleven v4 is control. Users can annotate a script with tags that specify emotion, pauses, and sound effects on a line-by-line basis, or they can simply describe what they want in plain sentences. According to ElevenLabs, the new version handles these instructions considerably better than v3 did.

In practice, that means non-verbal sounds — laughter, whispers, the slam of a door — are generated more reliably and in the right place. These are exactly the kinds of details that separate a flat synthetic read from a performance with texture. A scripted chuckle that lands on the wrong syllable, or a whisper that arrives with the wrong intensity, breaks the illusion immediately; ElevenLabs is positioning v4 as the version that closes much of that gap.

Pronunciation controls have also been tightened. Writers can use phonetic spelling to pin down how names, brand terms, and technical vocabulary are spoken, and the company says those controls now work more consistently than before. For anyone producing content full of acronyms, chemical names, or non-English proper nouns, that reliability is arguably as valuable as the emotional range.

Elevenlabs interface showing a dialogue script for two speakers, with orange-highlighted tags like warm, long pause, Gong sounds, and nervous laugh in the text, along with settings for voice, model Eleven v4, stability, and similarity.
Tags in the script let users set emotions, pauses, and sound effects for each line. | Image: Elevenlabs

An architecture designed to hold a voice steady

ElevenLabs attributes the improvements to a new model architecture that analyzes a script's tone, pacing, and context rather than treating each line as an isolated unit of text. That contextual awareness matters most when a production runs long.

Narrators and characters are expected to sound consistent throughout a project, even in workflows where individual lines get regenerated multiple times during editing. Anyone who has patched together a voice-over from dozens of separate takes knows the failure mode: a phrase here or there that subtly shifts in timbre, energy, or accent, leaving an audible seam. ElevenLabs says v4 is built to prevent that drift.

Each request can handle up to 10,000 characters, which works out to roughly ten minutes of finished audio. Longer works such as audiobooks are assembled from multiple segments, and the company expects pacing and delivery to remain stable across those transitions. In dialogue, the model is said to respond to the context of an entire scene instead of performing each speaker's line in isolation — a change that should make back-and-forth exchanges feel more like a conversation and less like alternating monologues.

Broader language coverage and revived clone tools

Eleven v4 supports more than 90 languages, an increase from the roughly 70 offered by v3. The expansion is paired with a specific claim about cloned voices: they should be able to speak other languages with native-sounding accents without gradually reverting to the accent of the original recording. That kind of drift has long been a complaint in multilingual dubbing, where a voice built from English source material can betray its origins when asked to speak, say, Spanish or Japanese.

Bar chart showing median time to first audible speech, with Eleven v4 Turbo at 150 ms ahead of Cartesia Sonic 3.6 at 262 ms, xAI TTS at 362 ms, Gemini 3.8 Flash-Lite TTS at 685 ms, and OpenAI GPT-4o mini TTS at 814 ms.
In Elevenlabs' benchmarks, Eleven v4 Turbo starts producing speech much faster than the competing models tested. | Image: Elevenlabs

ElevenLabs also says Professional Voice Clones work again in v4 after being unsupported in v3, restoring a capability that higher-end production workflows had lost. At the other end of the spectrum, an Instant Voice Clone requires only ten seconds of audio, keeping rapid prototyping accessible to users who do not have studio-quality recordings on hand.

Turbo pushes toward real-time voice agents

The second half of the announcement is about speed. Eleven v4 Turbo is built for real-time uses such as customer service calls and game characters, where a pause of more than a fraction of a second makes an interaction feel artificial.

According to ElevenLabs' testing, Turbo begins producing audible speech in about 150 milliseconds. The company frames this as resolving a longstanding trade-off: voice agent developers previously had to pick between a fast response and an expressive one. Turbo is intended to deliver both, which could make conversational agents sound less like scripted phone trees and more like participants in a live exchange.

Eleven v4 at a glance

  • More accurate adherence to direction tags and plain-language instructions for emotion, pauses, and sound effects
  • Laughter, whispers, and effects such as slamming doors generated more reliably than in v3
  • Consistent narrator and character voices across long productions, including when individual lines are regenerated
  • Up to 10,000 characters per request, approximately ten minutes of audio
  • Support for more than 90 languages, up from about 70 in v3
  • Cloned voices expected to hold native accents in other languages without drifting back
  • Professional Voice Clones restored after being unsupported in v3; Instant Voice Clone needs ten seconds of audio
  • Eleven v4 Turbo starts speaking in roughly 150 milliseconds, which ElevenLabs says beats competing models

Why the details matter more than the headline

The competitive frontier in speech synthesis has shifted. Fluent, human-sounding audio is now table stakes; what separates platforms is controllability — whether a director can ask for a specific delivery and get it back unchanged, take after take. Eleven v4's pitch is essentially about predictability at scale, from a ten-minute chapter download to a live agent answering a customer within a fifth of a second.

For producers of audiobooks, dubbed video, and interactive characters, the practical questions will be how well Turbo's speed survives real conversational conditions and whether segment-by-segment assembly truly holds its tonal footing across a full-length title. Those are the claims ElevenLabs is making now; the coming months of production work will determine how well they hold up.

This article is based on reporting by The Decoder. Read the original article.

Originally published on the-decoder.com