![A generated hero card for ElevenLabs Eleven v4 with two panels: on the left, a script block showing inline audio tags such as [laughs], [whispers] and [said angrily in French accent]; on the right, a panel reading rank 1, Elo 1,319, $80.00 per 1M characters and 1,674 appearances on Artificial Analysis' Provider Voice Arena, with a footer reading 'Provider Voice rank, Elo and price per Artificial Analysis, September 2026; tag syntax and latency vendor-documented.' The OrcaRouter logo sits in the bottom-right corner.](https://cms.orcarouter.ai/api/media/file/1-1462.png)
Eleven v4's Audio Tags and Ten-Second Clones: What Changes in Your Pipeline
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 982 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 197 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
The most consequential thing about ElevenLabs Eleven v4 is not its Elo. It is that the model reads direction written into the script — [laughs], [whispers], [sighs], [said angrily in French accent], even [light rain] — and that ElevenLabs Eleven v4 Turbo, the low-latency sibling shipped the same day, follows the same tags at a median time to first speech the vendor puts at around 150 ms. Both went to general availability on September 28, 2026. Artificial Analysis' Provider Voice Arena already has Eleven v4 at rank 1 with an Elo of 1,319 across 1,674 appearances, at a board price of $80.00 per million characters. The ranking is the headline. The direction syntax and the ten-second clone are the part that changes what you have to build.
![A generated four-card summary of what ElevenLabs Eleven v4 changed in the pipeline: a Script direction card listing [laughs], [whispers], [sighs], [said angrily in French accent] and natural-language delivery prompts; a Sound and scene card listing [light rain], [phone buzzing], multi-speaker dialogue with turn-level context and IPA phonemes for custom pronunciations; a Voice creation card listing Instant Voice Cloning from 10 s of audio, a Professional Voice Cloning tier above it, and 90-plus languages with a 10,000-character cap; and a Two endpoints card describing Eleven v4 as the most emotive, highest-quality model against Eleven v4 Turbo at about 100 ms median inference and about 150 ms to first speech. Footer: 'Tag syntax, cloning bar, language count, character cap and Turbo latency are vendor-documented; no independent score is quoted on this card.' The OrcaRouter logo sits in the bottom-right corner.](https://cms.orcarouter.ai/api/media/file/2-1397.png)
Two models, one launch, different jobs
ElevenLabs shipped two endpoints on September 28, 2026, and they are not the same product with a speed switch. ElevenLabs Eleven v4 is the expressive model: 90-plus languages, a 10,000-character cap per request, improved International Phonetic Alphabet support for custom pronunciations, and multi-speaker dialogue that the company says preserves speaker identity across a scene. ElevenLabs Eleven v4 Turbo keeps the 90-plus languages and the 10,000-character cap but is tuned for agents — the announcement quotes a median inference latency of about 100 ms and a median time to first speech of about 150 ms, measured by ElevenLabs in September 2026.

The reactive question — is the flagship fast enough for a phone agent — has no published answer. ElevenLabs gives latency figures for Turbo, not for Eleven v4 itself, and the two are distinct endpoints rather than one model under two names. If your agent's budget is 200 ms to first audio, Turbo is the endpoint in the documentation and the flagship is not.
The tags are a real interface, and a real dependency
Inline audio tags are not new to speech synthesis, but the useful change here is accuracy of adherence. The announcement states that Eleven v4 follows audio tags and natural-language direction prompts more precisely than prior models, and that this is how you direct a laugh, a whisper, or a delivery in a specific accent without editing audio afterwards.
Three practical consequences follow, and they matter more than the demo reels:
• Your script becomes a templated artifact, not a string. A tag is a token in your content pipeline: it needs escaping if untrusted text is ever passed through, and it needs to survive your CMS, your translation layer and your diff review. A translator who drops [whispers] has changed a performance, silently.
• Tag adherence is a vendor claim with a version attached. The claim that this generation follows tags better than the last one is ElevenLabs' own, tested by ElevenLabs, and it will not hold identically across languages. Validate on your own scripts and your own locales before you build a style guide on it.
• Sound-effect tags like [light rain] or [phone buzzing] move a little of the work of audio post-production into the text prompt. That is a cost shift worth measuring: if a tag replaces a foley pass, it is cheap; if it replaces nothing and only adds variance, it is a new failure mode in your QA.
Ten seconds is the number that changes onboarding
Instant Voice Cloning in this release captures a voice with high fidelity from ten seconds of audio, with the professional cloning tier still available above it. Ten seconds is short enough that it collapses a workflow step: consent capture, sample upload and first synthesis can happen in one sitting, on a phone, without a studio.
The consent problem does not shrink with the sample. It grows, because the bar to produce a convincing clone has fallen to a clip anyone can record. If you are cloning voices you do not own, the compliance question is now entirely yours to answer upstream of the API call — the model will do what you ask. Treat the ten-second figure as an operational threshold, not a permission slip.
What it costs while the window is open
ElevenLabs repriced at launch, and the promotional rate is the one on the page today. On the API pricing table, Eleven v4 lists at $0.022 per 1,000 characters against a stated list of $0.08, and Eleven v4 Turbo at $0.011 against $0.04. Both carry the same label: 72% off until October 12. The same page prices the previous flagship, Eleven v3, at $0.08 per 1,000 characters.
• Board price on the arena, per million characters — Eleven v4 $80.00, the highest in the top ten of that board.
• Vendor rate card, per thousand characters — $0.022 until October 12, then $0.08.
• What that means in practice — a script-heavy month of five million characters costs $110 during the window and $400 after it, on the same endpoint with the same code.

Anyone budgeting for Q1 against the number on the page today is budgeting against a number that expires in two weeks. Use $0.08.
Where a router earns its place in a launch like this
OrcaRouter does not host ElevenLabs, so there is nothing in this launch to call through us, and we will not imply otherwise — the place to call Eleven v4 is ElevenLabs' own API. What a single key is genuinely useful for in the fortnight after a launch is the comparison. If you are deciding whether the new flagship beats what you already run, you want to A/B the two on your own scripts rather than on a leaderboard, and doing that across two vendors means two accounts, two billing relationships and two integration paths before you have an answer.
That is the case we handle: one API key across 200-plus models with provider list prices passed through at 0% markup, so a vendor price change — including a promotional window with an expiry date — is live on our side the same day rather than at the next invoice cycle. Where a voice path is customer-facing, automatic failover moves traffic to a healthy upstream when one endpoint degrades, which is how you trial a brand-new model without betting a release on it.
What to test first, and what to ignore
Three checks will tell you more than any ranking. First, run your longest realistic script through the 10,000-character cap and see where the seams land — the cap is per request, and multi-speaker dialogue is the feature most likely to be split by it. Second, put five of your own sentences with tags into both v4 and v4 Turbo and listen for adherence drift between the expressive and the low-latency endpoint; they are separate models, and a tag that lands on one may not land identically on the other. Third, price your actual monthly character volume at $0.08 rather than $0.022, because that is the number your next quarter runs on.
The ~75% blind-preference figure in the announcement is a vendor measurement and we are not going to launder it into anything else. The independent number in this launch is the one on the board: first place at 1,319 Elo on 1,674 appearances, and an interval of roughly ±19 points around it. That is a strong result on a real board. It is not a specification, and it will not tell you whether your tag syntax survives a translation pass.
