Back to feed

ElevenLabs v4: Speech Synthesis in 90 Languages with Emotion Control

ElevenLabs has introduced new speech synthesis models, Eleven v4 and v4 Turbo, positioning them as the most emotionally expressive solutions on the market.

The new models support over 90 languages, expanding content localization capabilities. The system can clone a voice from a 10-second audio clip while preserving the speaker's original accent.

The model maintains voice integrity in long texts and automatically adjusts intonation to fit the context.

Emotion control is achieved through special tags in the text, such as [excited, happy] or [laughs]. Users can combine tags for sequential emotion playback and add background sounds like rain or phone rings.

The Turbo version is optimized for voice agents. First sound latency is around 150 ms, and generation begins before the large language model finishes formulating its response.

7.6K views

More from this channel Black Triangle Channel Telegram

Similar in this category Technology