
Eleven v4 vs Qwen Audio 3.1: The Cheapest Voice on the Board Won It
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 982 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 197 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
On the board where every entrant is given the same eight cloned voices, Alibaba Qwen-Audio-3.1-TTS-Plus finished first at Elo 1,178 and ElevenLabs Eleven v4 finished second at 1,157. The model that won costs $19.30 per million characters on that board. The model that came second costs $80.00. That is a 21-point quality gap and a four-fold price gap pointing in opposite directions, and it is the most interesting result in the whole launch week — more interesting than the number-one spot ElevenLabs took on the other arena, because on that one the ordering is conventional and the winner is the most expensive voice in the top ten.
The two boards are not interchangeable and the reason this matchup reads the way it does is entirely in which one you look at. Provider Voice has listeners judge each vendor's own voices. Controlled Voice clones the same eight voices — four US, four UK — for every model, so nobody gets to win on their default voice cast. ElevenLabs wins the first by 61 points. Alibaba wins the second by 21. Neither board is wrong, and the second one is the one that predicts a product shipping a cloned brand voice.

The controlled-voice scoreboard, in full
• Blind preference, matched clones — Qwen-Audio-3.1-TTS-Plus rank 1, Elo 1,178, $19.30 per million characters vs Eleven v4 rank 2, Elo 1,157, $80.00.
• Blind preference, each vendor's own voices — Eleven v4 rank 1, Elo 1,319 vs Qwen-Audio-3.0-TTS-Plus rank 4, Elo 1,258. A 61-point gap the other way.
• Price, arena normalisation — Qwen $19.30 vs Eleven v4 $80.00. Qwen is 76% cheaper.
• Price, vendor rate card — Eleven v4 $0.022 per thousand characters promotional and $0.08 after October 12. Qwen Audio 3.1's own rate card is not something we could read, so the arena normalisation is the only price we can compare on.
• Version on each board — 3.1 on Controlled Voice, 3.0 on Provider Voice. They are different models and the boards are not interchangeable.
• The 3.0 entry on Controlled Voice — rank 5, Elo 1,124, same $19.30 price. The 3.1 release is worth 54 Elo over its own predecessor at an identical price.
• Qwen's earlier generations on Provider Voice — Qwen3 TTS Flash rank 79, Elo 944; Qwen3 TTS rank 82, Elo 930.
• ElevenLabs' own previous flagship on Controlled Voice — Eleven v3 rank 7, Elo 1,073.
• Grouped-bar sanity check — the eight-way spread from rank 1 to rank 10 on Controlled Voice is 121 Elo. Qwen's 21-point lead over Eleven v4 is under a fifth of the whole field's range, which is the right scale to hold it at.

Two rows there do most of the work. The first is the 3.0-versus-3.1 comparison: Alibaba moved 54 Elo in one version bump without changing the price at all. That is a faster cadence than ElevenLabs' 84-point v3-to-v4 move, and it means the specific rank ordering in this article is a snapshot that both vendors are actively churning.
The second is the 121-point total spread. A 21-point gap in a board whose useful range is about 120 points is a real but modest difference — roughly the same magnitude as the gap between Qwen's own 3.1 and its 3.0. If your use case sits in a language neither model was tuned against, that margin is well within the range a few hard test scripts can overturn.
The version problem nobody is flagging
The result circulating as "Eleven v4 is second on Controlled Voice" is accurate and incomplete, and the omission matters for anyone planning around it. The model sitting above Eleven v4 on that board is Alibaba's 3.1 release. The Qwen entry on the Provider Voice board, meanwhile, is the 3.0 release — the 3.1 model does not appear in the top rows of that board as we read it. So the two headlines — "Eleven v4 is number one" and "Eleven v4 is number two" — are statements about two different Qwen models on two different tasks, and the Qwen version that beat it on one board is not the Qwen version it beat on the other.
This is not a gotcha about either vendor; it is a reason to distrust any single-sentence summary of a week like this. If you are choosing between these two systems, the comparison you want is 3.1 against v4 on your own cloned voice, in your own language, and neither vendor has published that.
What we can and cannot say about Qwen Audio 3.1
Honesty about evidence is worth more here than a full spec table, so here is the boundary of what this article is resting on. From the arena boards we have, for Qwen-Audio-3.1-TTS-Plus: a controlled-voice Elo of 1,178, a price normalisation of $19.30 per million characters, and the fact that it is the current release in a line whose previous version scored 54 points lower at the same price.
We do not have, and will not guess at: its language coverage, its latency, its input cap per request, whether it supports inline delivery directives, whether it offers voice cloning beyond the arena's eight voices, or what it bills on its own rate card. Those are all documented for Eleven v4 — 90-plus languages, a 10,000-character cap, audio tags, ten-second instant clones, IPA phoneme support — and their absence on the Qwen side is a gap in what is publicly readable, not a strike against the model. A purchase decision made on this article alone would be a decision made on two numbers.

One contrast does survive that limitation, because both sides are vendor-stated or both are board-stated. On price, the board normalisation is the common denominator and it favours Qwen by 76%. On quality under matched speakers, the same board favours Qwen by 21 Elo. Those two rows are comparable. Everything else here is not, and the article will not pretend otherwise.
Why a controlled board should carry more weight than usual
Most voice comparisons default to whatever leaderboard puts the model you already like on top. In this matchup there is a principled reason to prefer Controlled Voice, and it is specific rather than general.
Alibaba and ElevenLabs are competing with wildly different assets. ElevenLabs' advantage is a voice library built over years of curated casting; Alibaba's is a research organisation shipping model versions on a fast cadence. A board that lets each entrant bring its own voices measures the library as much as the synthesiser, and on that board ElevenLabs wins by 61 points. Strip the libraries out — same eight cloned speakers for everyone — and the advantage changes hands. If the voice your users hear is a clone of a specific person, you are not buying ElevenLabs' library, so the controlled board is the one that describes your purchase. If you are picking from a menu of stock voices, the opposite is true and Eleven v4's 61-point margin is the number that pays.
There is a second, less comfortable reason to weight the controlled result. A vendor-voice board rewards a distinctive default cast, and a distinctive cast can be polarising in ways an average Elo hides — one listener's characterful is another's affected. A cloned-voice board with matched speakers removes that variable entirely, which makes it the more reproducible measurement of the two.
Running either one, and what we actually route
OrcaRouter hosts neither ElevenLabs nor Alibaba's speech models. Neither vendor is in our catalogue, so there is no way to call either one through us and this article will not suggest otherwise — you go to the vendor's own API and several third-party platforms for both.
What we do cover is the layer above. If your speech stack is one node in a larger pipeline, we serve 200-plus models behind a single API key with provider list prices passed through at 0% markup, so a repricing at any upstream — including ElevenLabs' promotional window and the much larger list price that follows it on October 12 — is reflected on our side the day it happens rather than at the next billing cycle. For a decision that turns on a 4x price ratio, that timing is the difference between a comparison that ages well and one that reads as wrong in three weeks.
And where the voice path cannot go down, automatic failover moves traffic to a healthy upstream instead of failing the request. In a matchup this close on quality and this lopsided on price, the sensible architecture is both models behind one endpoint with routing deciding the split — not a permanent selection made from a leaderboard snapshot that both vendors will invalidate within a quarter.
The verdict, and its expiry date
Pick Qwen-Audio-3.1-TTS-Plus if you are cloning a voice and watching cost. It is first on the board that holds speakers constant, it is roughly a quarter of the price, and its line is moving faster — 54 Elo in one version step at an unchanged rate suggests the next version is a better bet than the current one is on paper.
Pick Eleven v4 if your users hear stock voices, if you need coverage in languages you can verify before shipping, or if the documented feature set — audio tags, phoneme control, 10,000-character requests, ten-second clones — is doing work in your product that the arena boards do not measure. Its 61-point lead on the vendor-voice board is not nothing, and the two features you can write into a script are capabilities, not scores.
Set a review date. Alibaba moved 54 Elo in a version bump this cycle and ElevenLabs moved 84 across a full generation, and Eleven v4's promotional rate expires on October 12. Whatever this comparison says today, the two facts most likely to flip it — a 3.2 release and a reverted price — both land before the end of the quarter.
