A hero card comparing Qwen-Audio-3.1-TTS with ElevenLabs Eleven v3, listing Qwen's 16 languages plus 20 dialects, 70% price cut and absence of an arena score against ElevenLabs' 74 languages, roughly $100 per million characters and Elo of 1,168, above a ribbon reading 'Announced at Apsara, September 23 2026'.
Guides & Insights

Qwen-Audio-3.1 vs ElevenLabs: A 70% Price Cut Meets the Incumbent

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The vendor cut the price of its Qwen-Audio-3.1-TTS API by roughly 70% on September 23, 2026, announced on stage at the Apsara Conference in Hangzhou alongside four other speech models. That single number is the reason this comparison is worth writing: ElevenLabs Eleven v3 has been the default production voice for most Western teams at an effective rate near $100 per million characters, and a credible challenger just repriced the same job at a fraction of it while widening language coverage at the same time. The two systems are not equivalent — Qwen-Audio-3.1-TTS is a hosted-only API from the vendor's Tongyi Lab with a smaller voice ecosystem, while ElevenLabs is a full content platform with the deepest voice library in the industry. But if your decision has ever been decided by the invoice, the arithmetic changed this week.

What actually shipped on September 23

Qwen-Audio-3.1 is not one model. It is a five-model refresh of Alibaba's entire speech stack, and the tiers matter because they have different pricing and different jobs:

Qwen-Audio-3.1-TTS — the direct successor to the synthesis model most teams use. Adds seven languages for a total of 16, keeps the 20 Chinese dialect regions, adds natural cross-language timbre transfer (one voice carrying across Mandarin, Cantonese, English, Japanese), instruction-based control of emotion, pace and delivery style, and handles noisy, echoed or otherwise poor reference audio without a separate denoising mode.

Qwen-Audio-3.1-TTS-Next — a new tier, and a different product. Alibaba calls it an "Audiogen" model: it generates voice, sound effects and background ambience in a single pass from text, timestamps and reference audio. Multi-speaker dialogue, podcasts, film and TV soundscapes, 48 kHz output. Per Alibaba Cloud's Model Studio documentation, it takes up to 3,000 characters of input, caps generated audio at 240 seconds for podcasts and 120 seconds otherwise, runs at 3 requests per second, and returns WAV, MP3 or PCM.

Qwen-Audio-3.1-ASR, Qwen-Audio-3.1-ASR-Next and Qwen-Audio-3.1-Realtime — recognition, audio understanding (emotion, music, environmental sound, event localisation) and full-duplex conversation respectively. Realtime is the one Alibaba is pushing into hardware: Qwen Office, Qoder, QwenNote A2, the QwenNote Eva desktop robot and the Qwen AI glasses.

The price cuts ran across the whole line, not just TTS: synthesis down about 70%, realtime down about 85%, recognition down as much as 95%. Alibaba also open-sourced Qwen-Audio-Agent, a framework for building realtime voice agents, and debuted Qwen3.8-LiveTranslate with latency pushed under 2.5 seconds.

A screenshot of Alibaba Cloud Model Studio's English documentation site showing the reference page for qwen3.8-livetranslate-flash-realtime, a real-time audio and video translation model that accepts audio and images and outputs text and audio across 60 input languages and 29 speech output languages, under an Apsara 2026 Conference banner, with the Speech-to-Speech branch of the navigation listing the qwen-audio-3.0 realtime models.

The scoreboard

A two-column scoreboard for Qwen-Audio-3.1-TTS and ElevenLabs Eleven v3 across price, languages, weights, voice library, delivery control and arena score, with a footer stating that the Qwen figures are vendor-reported and unscored and the ElevenLabs figures are per Artificial Analysis.

Price — Qwen-Audio-3.1-TTS: roughly 70% below the previous Qwen-Audio generation, which put the 3.0 Plus tier near $27.60 per 1M characters and the Flash tier near $15. ElevenLabs Eleven v3: about $100 per 1M characters as tracked by Artificial Analysis, with Flash at roughly $50.

Languages — Qwen-Audio-3.1-TTS: 16 languages plus 20 Chinese dialect regions. ElevenLabs Eleven v3: 74 languages.

Weights — Qwen-Audio-3.1-TTS: closed, hosted only on Alibaba Cloud Model Studio. ElevenLabs Eleven v3: closed, hosted only.

Voice library — Qwen-Audio-3.1-TTS: native cloning plus cross-language timbre transfer; no comparable public marketplace. ElevenLabs: the largest shared voice library in the industry, plus instant cloning from roughly 30 seconds of audio.

Delivery control — Qwen-Audio-3.1-TTS: instruction-based emotion, pace and style, inheriting the 86 inline delivery tags introduced with 3.0. ElevenLabs Eleven v3: audio tags plus the surrounding platform (dubbing, agents, studio).

Independent benchmark — Qwen-Audio-3.1-TTS: none yet. ElevenLabs Eleven v3: 1,168 Elo on Artificial Analysis's Provider Voice Arena leaderboard as of August 2026.

Where the challenger genuinely wins

Three things, and only one of them is price.

Price, obviously. A 70% cut against a $100-per-million-character incumbent is not a rounding difference. It is the difference between a product you prototype with and a product you ship to every user. If you are generating narration at volume, the annual saving is not a line item you have to argue for internally — it argues for itself.

Chinese, unambiguously. Twenty Chinese dialect regions is a moat nothing ElevenLabs offers comes close to. If your product serves mainland Chinese users, or Cantonese speakers, or a diaspora audience that code-switches mid-sentence, the comparison is over before it starts.

Cross-language timbre transfer. Qwen-Audio-3.1-TTS takes one voice and carries it naturally across Mandarin, Cantonese, English and Japanese. That is a specific capability with a specific use case — a single branded voice that survives a language switch — and it is a real differentiator for anyone localising a single presenter rather than casting a new one per market.

Where ElevenLabs still wins, and it is not close

Language breadth is the first thing, and it is the biggest. Seventy-four languages against sixteen is not a gap you close with a pricing announcement. If you ship in Polish, Turkish, Vietnamese and Portuguese, ElevenLabs covers you today and Qwen-Audio-3.1-TTS does not.

The second is the ecosystem. ElevenLabs is not a synthesis endpoint; it is a content platform. A shared voice library you can license rather than clone, dubbing workflows, agent infrastructure, a studio product, a documented API surface with years of client libraries behind it. Migrating away from that is a project, not a config change, and the teams that have built on it know exactly what they would be giving up.

The third is something harder to name and worth naming anyway: the perception of naturalness on a single English voice. Qwen's synthesis has historically been excellent on Chinese and merely very good on English, and while the 3.1 release claims to improve multilingual delivery, there is no independent listening test on the new model yet. Anyone who tells you they know how the new Qwen voice sounds in English is guessing.

The number nobody should quote yet

This is the part of the story that will get mangled in coverage, so here it is plainly. Qwen-Audio-3.1-TTS has no independent benchmark score. Its predecessor, Qwen-Audio-3.0-TTS-Plus, sat at 1,234 Elo on Artificial Analysis's Provider Voice Arena leaderboard as of mid-August 2026, ranking between second and fifth on that board. Qwen-Audio-3.1-TTS has not been added to any arena yet. Every quality claim in Alibaba's launch materials is vendor-reported and unreproduced.

That does not make the claims false. It makes them unverified, and it means the honest framing of this matchup right now is: one model with a price and no score, against one model with a score and a much higher price. Anything stronger than that is speculation wearing a benchmark's clothes.

A screenshot of the Artificial Analysis Provider Voice Arena leaderboard in its August 2026 state, ranking text-to-speech models by Elo: Cartesia Sonic 3.6 first at 1,277, Alibaba's Qwen-Audio-3.0-TTS-Plus at 1,234 and ElevenLabs Eleven v3 at 1,168.

The same discipline applies to latency. Alibaba's headline first-token figure for the ASR side of this family — 92 milliseconds — comes from the company's own high-concurrency benchmark on a 0.6B specification, not from third-party testing, and analysis published since the launch notes that ASR and TTS are separate products in a pipeline rather than one end-to-end model, and that community reports describe latency accumulating in long dialogues under a pseudo-streaming pattern. Treat vendor latency numbers across this whole family as ceilings, not as measured experience.

What the price cut is worth in practice

Take a realistic workload: a product that synthesises 20 million characters of narration a month across English and Mandarin. At ElevenLabs Eleven v3's roughly $100 per million characters, that is about $2,000 a month, or $24,000 a year. At the Qwen-Audio-3.0-TTS-Plus rate of roughly $27.60 per million, the same volume was already about $552 a month. Apply the ~70% cut Alibaba announced on September 23 and the same workload lands in the low hundreds of dollars a month.

None of those figures are list prices you can sign against — ElevenLabs bills in credits through subscription tiers, and Alibaba's cuts are percentage reductions against rates that differ by region and by tier. But the shape of the comparison is unambiguous, and the shape is what you budget against.

One caveat worth flagging for anyone modelling this carefully: Alibaba Cloud moved its realtime transcription billing on the Singapore node from duration-based metering to token-based metering, at $0.93 per million input tokens and $0.70 per million output tokens. Token billing is less predictable than duration billing when you do not control how many tokens a given audio slice becomes, and noise, silence and multi-turn context all inflate token consumption. A 70% headline cut does not automatically translate to a 70% saving on your specific traffic.

How to actually run the switch

If you are evaluating this, the useful thing to know is what a migration costs you in engineering time. Both models are closed hosted APIs, so you are integrating against an HTTP endpoint either way, and the work is a client, a retry policy and a voice inventory — not a re-architecture.

That is where OrcaRouter's shape helps, with one honest caveat first: we do not route Qwen-Audio-3.1-TTS or any Qwen-Audio model. Alibaba's speech endpoints are out of scope for us, and the same is true of ElevenLabs. What we do cover is the inference layer around them — one API key across 200-plus text models at 0% markup, meaning provider list price passed straight through, so a vendor price cut is live on our side the same day rather than after a repricing cycle. For teams building the application that drives a speech pipeline — the summariser, the script generator, the content planner — that means one contract instead of several, automatic failover when an upstream degrades, and a routing DSL for composing models into a single call. It does not mean you call Qwen's voice through us. Anyone who tells you otherwise has not read the catalogue.

Who should move, and who should wait

Move now if your workload is Chinese-first, if dialect coverage is on your requirements list, if volume cost is the constraint that has kept you on a cheaper-but-worse voice, or if you need one branded voice to survive a language switch. The September 23 pricing makes the economics hard to argue with and the Chinese-language quality was already competitive before it.

Wait if you ship in more than about sixteen languages, if your product depends on a licensable voice library rather than cloning, if you are deep in a platform's dubbing or agent tooling, or if your procurement process requires an independent quality score before sign-off. None of those are permanent blockers, but none of them were addressed by a price cut.

And watch one thing specifically: whether Qwen-Audio-3.1-TTS appears on a blind-listening arena. The moment it does, this comparison stops being about price and language counts and starts being about whether the cheaper voice is actually as good. That is the question the launch did not answer.