New speech model clones voices from ten seconds of audio and supports more than 90 languages

Text-to-speech company ElevenLabs launched its newest speech model, Eleven v4, on 28 September 2026, together with a low-latency variant called v4 Turbo built for real-time agents.

Illustration: a single reel of audio tape splits into dozens of thin colored threads fanning out across warm paper — one voice multiplied into many languages.
Illustration
Gift article

New speech model clones voices from ten seconds of audio and supports more than 90 languages

Text-to-speech company ElevenLabs launched its newest speech model, Eleven v4, on 28 September 2026, together with a low-latency variant called v4 Turbo built for real-time agents. According to the independent benchmarking firm Artificial Analysis, v4 took first place on their audio leaderboard hours after launch — but most of the quality and latency figures in the launch are the company's own measurements.

What's new in v4

The most visible change is how users control the speech. Where previous versions used SSML tags (Speech Synthesis Markup Language), SSML is now disabled in v4, according to the product page's frequently asked questions as reproduced by Unite.AI. Instead, the model reads natural-language instructions and inline audio tags directly in the text, such as [laughs], [said angrily in French accent], [light rain] and [phone buzzing].

TechCrunch reports that v4 is built on a new architecture that expands the expressiveness tags introduced with v3, and which the company says delivers better voice consistency across long texts — a well-known problem where synthetic voices can drift away from the original's timbre partway through a recording.

Languages and voice cloning

Language support increases from 70 to more than 90 languages. ElevenLabs states that the largest quality improvements were observed in Japanese, Brazilian Portuguese, Mandarin and Cantonese — according to TechCrunch's account of the company's own assessments, not an independent measurement.

The cloning threshold has also been cut sharply: with v4, the company says (as reported by TechCrunch), ten seconds of audio is enough to clone a voice. That makes the tool far more accessible — and far easier to misuse, which we return to below.

Turbo: latency for real-time agents

Eleven v4 Turbo targets real-time speech agents, where speed to the first word determines whether a conversation feels natural. Here it is important to distinguish two figures that ElevenLabs provides, and which Windows Report separates:

  • ~100 ms median inference latency — the time the model spends on generation itself.
  • ~150 ms median time-to-first-speech — the time from request to first speech, i.e. what the user notices.

Both figures are company-provided. According to the methodology in the launch post, as described by Unite.AI, the measurements were carried out in September 2026 via WebSocket streaming with identical scripts and standard configurations, with network latency measured and removed for all systems. In the same comparison, ElevenLabs states that Cartesia Sonic 3.6 comes in at 262 ms and OpenAI GPT-4o mini TTS at 814 ms in time to first speech; competitors xAI TTS and Google Gemini Flash-Lite TTS were also tested. Since the figures come from ElevenLabs' own benchmarks, they have not been independently verified.

What is independent, and what is the company's own numbers?

There is one independent data point: Artificial Analysis' audio leaderboard, where listeners rate anonymous samples side by side, placed v4 in first place hours after launch, according to Martin Cid Magazine.

By contrast, the figure that roughly three in four listeners preferred v4 in blind comparisons against competitors is based on ElevenLabs' own tests and must be attributed to the company. The same applies to all the latency and language quality figures above. Readers who want to assess the model for themselves should therefore place the most weight on the leaderboard result — which is based on listeners' rankings of anonymous samples — and wait for more independent measurements of latency and quality.

Practical details: access and pricing

Both models are available immediately in ElevenAgents, ElevenCreative and via ElevenLabs' API, with access included in the free tier, according to Unite.AI. The free tier is, however, small: around ten minutes of generated audio per month, according to Unite.AI's review as reproduced by Martin Cid Magazine. Paid plans start at $6 per month.

The launch post was written by the company's founders Mati Staniszewski and Piotr Dabkowski, according to Unite.AI.

Misuse and safety

Ten seconds of audio is enough to create a convincing clone, and that makes fraud and voice impersonation a real risk — a point Martin Cid Magazine highlights in its coverage. Safeguards include a requirement of verified consent for instant clones, and generated audio being flagged by ElevenLabs' AI Speech Classifier, according to the product page's FAQ. But as the magazine points out, part of the responsibility rests on whoever uploads audio actually having the right to it. A system that accepts ten seconds has limited ability to stop someone who holds ten seconds of another person's voice.

The context: rapid growth toward a stock listing

The launch comes in a busy period for the company. ElevenLabs earlier this year raised $500 million from Sequoia in a round that valued the company at $11 billion, TechCrunch writes. The company's annualized revenue has risen from around $330 million at the start of the year to over $600 million. According to TechCrunch, rumors have moreover swirled about a new round at a $22 billion valuation — rumors that have been neither confirmed nor denied.

Founder and CEO Mati Staniszewski has told TechCrunch that the company aims for a stock listing in the years ahead, without committing to a date.

What remains to be seen

V4 launches into a market with fierce competition from Google, OpenAI, Cartesia, Deepgram and Fish Audio, and Artificial Analysis' first ranking is only an early indicator. What remains unresolved is how stable the quality is over longer texts and across all 90+ languages, whether competitors respond with latency improvements of their own, and how well the consent and watermarking schemes work in practice now that the cloning threshold is ten seconds. Until more independent benchmarks are available, it makes sense to read the company's figures as measurements — not as fact.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.