hn.today

Eleven v4 by ElevenLabs

elevenlabs.io9 points1 comments
Screenshot of Eleven v4 by ElevenLabs

Eleven v4 is a new text-to-speech model and a low-latency variant, Eleven v4 Turbo, designed to generate highly emotive, context-aware speech that preserves speaker identity across languages and multi-speaker dialogue. Built on a new architecture, it interprets tone, pacing, character and scene context to produce delivery that can be dramatic, tender, urgent or comedic while maintaining natural conversational dynamics rather than stitched-together lines. Users can fine-tune delivery with natural-language direction and inline audio tags (laughs, said angrily in French accent, light rain, etc.), plus improved IPA support for reliable custom pronunciations. Use cases highlighted include audiobooks, character performances, dubbing and responsive voice agents across industries such as healthcare and gaming.

Benchmarks and specific improvements are emphasized: ranked #1 by Artificial Analysis and preferred in blind tests (winning 65-81% of head-to-head comparisons), with Eleven v4 Turbo offering median time-to-first-speech around 150 ms and median inference latency ~100 ms for fast responses. Voice cloning needs as little as 10 seconds for high-fidelity Instant Voice Clones, with Professional Voice Clone support and stronger cross-language accent adherence across 90+ languages. Request-stitching, speaker consistency across generations, and integration with ElevenAgents, ElevenCreative and ElevenAPI are presented as practical advances for deploying expressive, low-latency TTS at scale.

Read on elevenlabs.io1 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.