ElevenLabs has launched Eleven v4, a text-to-speech model designed to interpret tone, pacing, emotion, character, and conversational context, alongside Eleven v4 Turbo for lower-latency applications. The company says Eleven v4 ranked first on Artificial Analysis’s Provider Voice Arena leaderboard and was preferred by about 75% of listeners in blind comparisons with competing systems, though these are company-reported evaluations. Eleven v4 Turbo has a median inference latency of about 100 milliseconds and a median time to first speech of about 150 milliseconds, according to measurements described in the article. Both models support natural-language delivery instructions, inline audio tags, and improved International Phonetic Alphabet pronunciation control. They are designed to produce more consistent multi-speaker dialogue and preserve speaker identity across conversations, narration, regenerated lines, and longer projects. Both models support more than 90 languages, including cross-language speech that retains the original speaker’s identity while adopting a native accent. Voice cloning can use 10 seconds of audio for Instant Voice Clones, while Professional Voice Clones are also supported. Eleven v4 and Eleven v4 Turbo are available through ElevenAgents, ElevenCreative, and the ElevenAPI.
