A generated hero title card for StepAudio 3 TTS vs Qwen-Audio-3.0-TTS with the kicker 'OrcaRouter · model radar — head to head', the headline 'StepAudio 3 TTS vs Qwen-Audio-3.0-TTS' and the subtitle 'One has a blind-listening arena score. The other shipped seven weeks after the last vote', above two rounded model cards either side of a versus mark.
Guides & Insights

StepAudio 3 TTS vs Qwen-Audio-3.0-TTS: One of These Has an Arena Score, and It Is Not the New One

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Here is the whole comparison in one line. Qwen-Audio-3.0-TTS sits on the Text-to-Speech Arena at an Elo of 1,234 for its Plus tier — a score earned across 2,502 blind listening votes. StepAudio 3 TTS, released by StepFun on September 15, 2026, is not on that board at all. Its predecessor is: StepAudio 2.5 TTS holds an Elo of 1,209 from 1,306 votes.

So the honest question is not "which sounds better". It is whether a 25-point Elo gap, measured on the previous generation, survives a release that StepFun says improved prosody and cut the price by more than half. Nobody has run that vote yet, and everything below is arranged around that gap in the evidence rather than pretending it away.

The board, as it actually reads today

Worth seeing the real standings, because several circulating write-ups cite a stale version of them. The blind-listening arena ranks synthesis models by which sample voters preferred. Read across the top of the provider voice board:

• Cartesia Sonic 3.6 — Elo 1,277, August 2026, $49.0 per million characters

• Inworld Realtime TTS-2 — Elo 1,243, August 2026, $20.8 per million characters

• SpeechifyAI Simba 3.2 — Elo 1,238, July 2026, $6.6 per million characters

Alibaba Qwen-Audio-3.0-TTS-Plus — Elo 1,234, July 2026, $27.6 per million characters, 8 arena voices

• VUI Labs Luna TTS — Elo 1,229, June 2026, $80.0 per million characters

• Inworld Realtime TTS-2 Flash — Elo 1,213, August 2026, $10.4 per million characters

• StepFun StepAudio 2.5 TTS — Elo 1,209, August 2026, $85.0 per million characters, 8 arena voices

Google Gemini 3.1 Flash TTS — Elo 1,204, April 2026, $18.3 per million characters

Screenshot of the Artificial Analysis Provider Voice Arena text-to-speech leaderboard, showing Cartesia Sonic 3.6 first at Elo 1,277, Alibaba Qwen-Audio-3.0-TTS-Plus fourth at Elo 1,234 over 2,502 samples at $27.6 per million characters, and StepFun StepAudio 2.5 TTS seventh at Elo 1,209 over 1,306 samples at $85.0 per million characters, with confidence intervals of -14/+14 and -16/+16 respectively.

Two things fall out of that list immediately. The first is that Qwen's Plus tier is not the leader — it is fourth, behind Cartesia, Inworld and Speechify, and the 1,234 figure that gets quoted as a crown is a mid-table score on a board that has moved since. The second is that StepAudio 2.5 TTS was re-run in August 2026 and placed seventh, 25 Elo behind Qwen, at more than three times the price.

And a third thing that matters more than any of it: look at the confidence intervals. Qwen's Plus tier is listed at −14/+14 across 2,502 samples. StepAudio 2.5 TTS is listed at −16/+16 across 1,306. Those intervals overlap. A 25-point gap between two models whose error bars span 14 and 16 points in either direction is not a demonstrated quality difference — it is two models the arena cannot currently separate, ranked in an order that a few hundred more votes could reverse.

A generated scoreboard titled 'StepAudio 3 TTS vs Qwen-Audio-3.0-TTS — the scoreboard'. The StepAudio column reads Arena Elo 'not scored yet', Samples 'none', Price '2.5 yuan / 10,000 chars', Voices 'not published', Languages 'Chinese-first', Request size 'not published'. The Qwen column reads Arena Elo '1,234, Plus tier', Samples '2,502', Price '$27.6 / 1M chars', Voices 'over 1,000', Languages '16 + 20 dialects', Request size '20,000 chars'. Footer: 'Qwen figures per the Artificial Analysis text-to-speech arena; StepAudio 3 TTS has no arena score yet.'

What the missing data point costs you

StepAudio 3 TTS is the model StepFun replaced StepAudio 2.5 TTS with, and the launch claims control over timbre, intonation, rhythm and breathing pauses, plus paralinguistic behaviour — laughter, hesitation, stuttering, repetition, self-correction — delivered through a streaming architecture where generation and playback overlap.

That is exactly the set of qualities the arena measures. It is also exactly why the absence of a score is frustrating rather than fatal: the benchmark that would answer the question exists, is publicly run, and simply has not been re-run since the new model shipped. Anyone telling you StepAudio 3 TTS sounds better than Qwen-Audio-3.0-TTS is extrapolating from a predecessor that lost by a margin inside its own error bars.

Price is the part that is fully documented, and it moved a lot

StepAudio 3 TTS bills at 2.5 yuan per 10,000 characters on StepFun's own platform. StepAudio 2.5 TTS billed at 5.8. On the arena's own character-based pricing column, the older model was $85.00 per million characters; 2.5 yuan per 10,000 characters works out to roughly $35 per million at recent exchange rates — a conversion that is ours rather than a published figure, and worth redoing against your own region and FX. That is a cut of close to 60% on the same line item.

It is still above Qwen. Qwen-Audio-3.0-TTS-Plus is listed at $27.6 per million characters, and the Flash tier is cheaper still, with Alibaba Cloud Model Studio billing synthesis by input character rather than by audio produced — output is not charged separately. Third-party platforms listing the Flash tier quote rates in the high teens per million, though regional variation makes any single headline number unreliable.

So the shape of this is: StepFun roughly halved synthesis cost, closed most of the distance to Qwen, and did not cross it. If price per character is your binding constraint, the argument narrows but the answer does not change. If it is not, price stopped being the reason to pick either model.

Voice inventory and language coverage are not close

This is where the comparison stops being a judgement call and becomes a specification question.

• Voice inventory — Qwen-Audio-3.0-TTS over 1,000 voices across its tiers vs StepAudio 3 TTS no comparable published count

• Languages — Qwen-Audio-3.0-TTS 16 languages plus 20 Chinese dialects, with the Flash tier tuned for low-resource languages and dialect authenticity vs StepAudio 3 TTS Chinese-first, with no equivalent coverage claim

• Request size — Qwen-Audio-3.0-TTS 20,000 characters per request vs StepAudio 3 TTS not published

• Latency — Qwen-Audio-3.0-TTS Flash targets first-packet latency within 200 ms per Alibaba's own docs vs StepAudio 3 TTS streaming, no published figure

• Tiering — Qwen-Audio-3.0-TTS Plus for expressiveness and Flash for real-time, six flagship voices locked to one tier or the other vs StepAudio 3 TTS a single tier

• Control surface — Qwen-Audio-3.0-TTS natural-language delivery direction plus inline expression tags such as [excited], [laughing], [whispers] embedded in the input text vs StepAudio 3 TTS model-inferred prosody from context

• Voice cloning — StepAudio 3 TTS a flat 9.9 yuan per cloned voice across its TTS tiers vs Qwen-Audio-3.0-TTS cloning through Alibaba Cloud Model Studio

• Output — Qwen-Audio-3.0-TTS 24 kHz default sample rate vs StepAudio 3 TTS not published

That control-surface row is the real design difference, and it is not a quality question. Qwen's expression tags are explicit and synchronous: you write the emotion into the text and the model renders it, which makes them testable and regression-friendly. StepFun's description is of a model that infers the appropriate hesitation or laugh from context, which sounds better when it is right and is much harder to pin down when it is wrong. Pick according to whether you would rather author the performance or delegate it.

Screenshot of Alibaba Cloud Model Studio documentation for qwen-audio-3.0-tts-flash, showing the model capabilities table with text input and audio output, function calling and fine-tuning unsupported, and a note that the Flash version controls first-packet latency within 200 ms.

How to decide without the measurement

You cannot reason your way to an answer from published numbers, because the decisive one does not exist. That leaves testing, and testing these two has a specific shape worth planning for.

Both are hosted-only commercial APIs — Qwen-Audio-3.0-TTS through Alibaba Cloud Model Studio, StepAudio 3 TTS through StepFun's open platform. Neither is open weights, so there is no self-hosted route to evaluate. Being precise about what OrcaRouter covers here: the text-to-speech models on our catalogue today are the OpenAI tts-1 family, gpt-4o-mini-tts and Google's Gemini TTS previews. Neither of the models in this article is on it, and Qwen-Audio-3.0-TTS is not either — nothing here should be read as a claim that we serve them.

The part that does sit on OrcaRouter is the text half of the pipeline. The scripts your synthesis layer reads are generated by a language model, and that is the component that changes most often, costs the most to re-contract, and is easiest to leave un-pinned — nearly 200 text models behind one API, automatic failover, a routing DSL that composes several into one call, and model fusion where one model's judgment is not enough, with provider list price passed through at no markup. If you run a blind evaluation between these two synthesis models, generate the corpus through one endpoint so both candidates read identical input, and keep the synthesis layer behind an interface you can swap.

Where this leaves the decision

Take Qwen-Audio-3.0-TTS today if you need breadth — over a thousand voices, 16 languages and 20 dialects, a documented 20,000-character request ceiling, a latency tier and an expressiveness tier, a 200 ms first-packet target, and a rate below StepFun's. It is the model with a measurement behind it and the wider surface, and its fourth-place Elo is a real result even if it is not the crown it gets quoted as.

Take StepAudio 3 TTS if you are building Chinese-first, want one tier instead of a tier-selection decision, want cloning at a flat published fee, and are willing to be early on a model that has not been blind-tested. You are paying roughly 27% more per character than Qwen's Plus tier in exchange for a generation of claimed prosody improvement over a model that sat 25 Elo back with overlapping error bars.

What would settle it is one re-run of the arena with StepAudio 3 TTS in the pool. Until that vote happens, this is a measured model against an unmeasured one, and the 25 points that separate their predecessors are inside the noise — which is not a ranking, and should not be sold as one.