Content
ElevenLabs Unveils v4 Speech Models with Faster Voice Cloning and Broader Language Support
New Architecture and Faster Voice Cloning
Improved Expression and Context Awareness
Support for More Than 90 Languages
Lower Latency for Enterprise Voice Agents
Growing Competition in Speech AI
Funding, Revenue Growth, and Expansion
Potential IPO Plans
ElevenLabs v4 Speech Model Adds Greater Control Over Expression and Support for 90 Languages
Time: Sep, 28, 2026

ElevenLabs Unveils v4 Speech Models with Faster Voice Cloning and Broader Language Support

ElevenLabs launched two new speech models on Monday: ElevenLabs v4 and v4 Turbo. The models offer greater control over vocal expression, reduced latency for voice agents, and support for more than 90 languages.

New Architecture and Faster Voice Cloning

After releasing its v3 model last year and previewing its successor at an event in Warsaw earlier this year, ElevenLabs has introduced a new architecture for the v4 generation. The updated design provides more precise control and faster voice cloning, allowing users to replicate a voice from only 10 seconds of audio.

Improved Expression and Context Awareness

For creative applications, v4 preserves voice identity more consistently across longer passages. It also considers textual context when reading aloud, enabling more appropriate shifts in expression.

ElevenLabs introduced inline expression tags with v3. In v4, this feature has been expanded so users can combine multiple tags, with the model following them in sequence.

Support for More Than 90 Languages

The previous model supported 70 languages, while v4 expands coverage to more than 90. According to ElevenLabs, the largest improvements in quality were observed in Japanese, Brazilian Portuguese, Mandarin, and Cantonese.

Lower Latency for Enterprise Voice Agents

ElevenLabs has rapidly expanded its enterprise calling business over the past year, with large companies now accounting for more than 55% of its business. The company says v4 is particularly well suited to voice agents because its lower latency enables smoother, more natural conversations.

The model can begin producing audio as soon as the underlying large language model starts generating a response. It can also respond differently to confrontations, escalations, and holds, supporting more effective issue resolution.

Growing Competition in Speech AI

Competition in the speech-model market has intensified as startups including Cartesia, Deepgram, Fish Audio, Boson, and WellSaid Labs develop increasingly expressive systems. Major technology companies such as Google and OpenAI have also continued to improve their voice models.

Funding, Revenue Growth, and Expansion

Earlier this year, ElevenLabs raised $500 million in a Sequoia-led funding round that valued the company at $11 billion. Rumors have since emerged of a follow-up round that could double its valuation to $22 billion.

The company’s annualized revenue run rate has risen from approximately $330 million at the beginning of the year to more than $600 million. ElevenLabs has also hired aggressively across markets including India, Europe, and Brazil, bringing its workforce to more than 800 employees.

Potential IPO Plans

Co-founder and CEO Mati Staniszewski recently said ElevenLabs is targeting an initial public offering “in the next years,” although he did not provide a specific timeline.

2
Live Chat