
StepAudio 3 ASR vs Whisper Large v3 Turbo: 1.7% Against 4.6%, and Nobody Has Run the Test
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
The comparison you want does not exist. StepAudio 3 ASR and Whisper Large v3 Turbo sit on the same Artificial Analysis AA-WER index for non-streaming speech to text — 1.7% word error rate against figures in the 4 to 5.5% range — and as of late September 2026 no one has published a head-to-head on identical audio. Not StepFun, not the Whisper maintainers, not an independent lab. StepFun's own documentation benchmarks against Doubao ASR 2.0, Seed 2.0 Lite and HY3.0 ASR Preview, all Chinese ASR products. The Whisper side has no vendor publishing comparisons at all, because Whisper Large v3 Turbo has been downloadable for years and its numbers depend on who is serving it.
So this page does the honest version: it reads the one board that measures both, separates what is comparable from what is not, and states plainly where the evidence runs out. The short answer is that the gap on accuracy is real and roughly three points wide, and the gap on everything else — price, licence, deployability — runs the other way and is larger.
What each one actually is
StepAudio 3 ASR is a hosted transcription model released by StepFun on September 15, 2026 as part of the five-model StepAudio 3 audio family. It is reached through a single streaming endpoint that returns incremental text during decoding and a final transcript at the end. StepFun describes it as transcription with a language model attached for context, tuned for medical, legal, finance, automotive and programming audio and for inputs that break conventional recognisers — whispered speech, fast speech, connected speech, singing, and audio with music underneath. There are no open weights, no on-premise path, and no self-hosting option at any tier. The published list rate is 2.8 yuan per hour, about $6.67 per 1,000 minutes on the leaderboard's own price axis.
Whisper Large v3 Turbo is an open-weights model published by OpenAI under the MIT licence: 809 million parameters, 99 languages, a pruned and fine-tuned derivative of Whisper large-v3 with the decoder layers reduced. It has been downloaded millions of times and is served commercially by a long list of platforms and runnable on your own hardware. That last property is the whole comparison, and it is the reason the accuracy gap is not the only number that matters.
The one board that measures both
Artificial Analysis's AA-WER index reports the percentage of words transcribed incorrectly, lower is better, and version 2 of it blends three datasets: AA-AgentTalk at 50%, VoxPopuli-Cleaned-AA at 25%, and Earnings22-Cleaned-AA at 25%. The non-streaming view shows 36 of the 71 models it tracks. Both models appear on it, which is more than most comparisons can claim.
Read as the board lists them:
• StepAudio 3 ASR — 1.7% AA-WER, roughly $6.67 per 1,000 minutes of audio
• Whisper Large v3 Turbo, one hosted endpoint — 4.6% AA-WER, about $0.67 per 1,000 minutes
• Whisper Large v3 Turbo, a second hosted endpoint — 5.5% AA-WER, about $15.00 per 1,000 minutes
• Whisper Large v3, a different open-weights generation — 4.1% on one endpoint, 10.1% on another
• StepAudio 2.5 ASR, StepFun's previous generation — 4.7%
The last two lines are the most instructive on the board, and they are not about StepFun at all. The same open Whisper weights score 4.1% on one endpoint and 10.1% on another. That is a 6-point spread on identical parameters, produced entirely by differences in serving: chunking strategy, hardware, decoding configuration, and how each host handles audio longer than its internal limits. The board's own footnote says as much, noting that for one of its datasets some models get their audio chunked to about nine minutes and others to about thirty seconds because of time limits.
Which means the comparison "StepAudio 3 ASR vs Whisper Large v3 Turbo" is really two comparisons stacked. The model against the weights, and the model against whichever endpoint you happened to benchmark. The second one varies by more than the first.

Where the accuracy gap comes from, and how much to trust it
Three points of word error rate is a large gap in a field where the top five systems sit within 0.6 points of each other. It is also the only number in this comparison that a third party measured, and it comes with real caveats that cut in both directions.
Against StepAudio 3 ASR: the figure is a blended index, and the blend is 50% AA-AgentTalk — an agent-conversation dataset. A model tuned for domain-specific transcription with an LLM attached for context is well suited to that shape of audio. Reweight the index toward read speech or broadcast news and the ordering could tighten. There is also still no published technical report for the ASR model, no released test set, no decoding configuration and no reproducible protocol from anyone outside StepFun.
Against Whisper Large v3 Turbo: the 4.6% is one hosted endpoint's number, not a property of the model. The weights are fixed and public; the score is not. A team that runs the same weights carefully, with domain-appropriate chunking and a good decoder configuration, can land meaningfully better than the board's row — and a team that runs them carelessly can land worse than the 10.1% the other open-weights generation posted. The honest summary is that the open-weights side of this comparison has a floor you can measure and a ceiling you have to earn.
What is missing is the number that would settle it: both systems, one audio set, one decoding protocol, published. Nobody has run it, and until someone does, "1.7% vs 4.6%" is a comparison of two measurements taken under different conditions rather than a measurement of two systems against each other.
The other four dimensions, where Whisper wins by more
Accuracy is the dimension StepFun wins. On the rest of the list the open-weights model is not close, and these are the ones that decide whether a deployment is possible at all:
• Licence and weights — MIT-licensed open weights with 809M parameters vs a hosted API with no weights published at any tier.
• Deployment — self-hosted, on-premise, air-gapped, or through any of dozens of commercial platforms vs StepFun's API only.
• Cost floor — about $0.67 per 1,000 minutes on one hosted endpoint, and lower still if you run it yourself, vs about $6.67 per 1,000 minutes on the vendor's published rate.
• Data residency — the audio never leaves infrastructure you control vs audio sent to a third-party endpoint.
• Integration risk — the model cannot be withdrawn, repriced or deprecated under you, because you hold the weights vs a hosted rate that can change with a price list.
• Long-audio behaviour — chunking is your decision and can be tuned per file vs whatever the vendor's endpoint does internally, which is not documented.
That price gap is roughly ten to one against the cheaper Whisper endpoint, and it is the mirror image of what happened inside StepFun's own line: the model this one replaces, StepAudio 2.5 ASR, listed at 0.15 yuan per hour, and the current tier lists at 2.8. StepFun raised the transcription rate by about nineteen times to buy accuracy, which puts it in a different market from a self-hosted open-weights model rather than a more expensive version of it.

Which one you should pick
The split is cleaner than most head-to-heads:
• Take Whisper Large v3 Turbo if the audio is ordinary, the volume is high, the data cannot leave your infrastructure, or you need the deployment to be permanent and price-stable. The MIT licence means the decision is yours to reverse and nobody can take it back.
• Take StepAudio 3 ASR if the audio is genuinely hard — accented, whispered, sung, or with music underneath — or if the domain is one of the five StepFun tuned for. That is what the premium buys, and on that audio a three-point accuracy gap is the difference between a usable transcript and a manual correction pass.
• Take neither, exclusively, if you do not know which your audio is. The cost of being wrong is asymmetric: the open-weights path can be tested on your own hardware for the price of a GPU hour, and the hosted path costs nothing to try but commits you to a rate you cannot control.
The case for a hybrid is stronger here than in most comparisons, and the reason is the endpoint spread on the board. Because the same Whisper weights score anywhere from 4.1% to 10.1% depending on who serves them, the practical move is to route the open-weights path across more than one host rather than committing to one — which is what automatic failover gives you, and it is the same mechanism that lets you put a hosted model in front of a production path without betting the path on it.
How this fits a stack, and what we do and do not serve
Whatever you pick for transcription, the transcript goes somewhere. A summariser, a structured extraction, a compliance check, a search index — those calls land on ordinary text models, and they are the part of the bill that scales with usage rather than with audio minutes. StepFun's own realtime model advertises tool calling and asynchronous execution of long tasks, which means even the audio layer is constantly handing work to something that is not an audio model.
To be explicit about scope, because it would be easy to imply otherwise: neither StepAudio 3 ASR nor Whisper Large v3 Turbo is on OrcaRouter. StepAudio is reached through StepFun's own API, and the Whisper weights are yours to run or to take to a hosting platform. What OrcaRouter carries is the delegation layer the transcript feeds — nearly 200 text models behind one API, with automatic failover so a downstream call does not die with a single provider, a routing DSL that composes several models into one call, and model fusion when one model's judgement is not enough. Where a provider publishes a rate, that list price is passed through with no markup, so a price change reaches your bill the day it is announced — which, given that this comparison contains a vendor that raised its transcription rate nineteen-fold in one release, is not a hypothetical benefit.
If you are about to spend real money on the accuracy question, the cheap sequence is this: run the open weights on your own audio first, because you can, and because the number you get will be more relevant than any board's. Then decide whether the gap that remains is worth ten times the rate. Most teams will find that it is worth it for a slice of their traffic and not for the rest, and that is a routing problem rather than a model choice.
What would settle it
One published test. Both systems, the same audio, the same decoding protocol, an honest split between read speech, conversational audio and the hard cases StepFun claims to have tuned for. Until that exists, the comparison is a third-party index on one side and a set of open weights with a variable score on the other, and the only responsible way to read it is as a starting point rather than an answer.

