
StepAudio 3 ASR vs StepAudio 3: There Is No Model by That Name
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Search for StepAudio 3 and you will find a family, not a model. StepFun's September 15, 2026 audio release put five products under that name — Realtime, ASR, text-to-speech, Gen and Music — and the transcription product among them, StepAudio 3 ASR, is the one that has been climbing leaderboards since. The tier that StepFun's own documentation calls its largest ASR model to date is StepAudio 3 ASR Max, and the conversational sibling it shares a backbone with is StepAudio 3 Realtime. If you are trying to compare "StepAudio 3 ASR" against "StepAudio 3", you are comparing a product against a product line, and the answer depends entirely on which of the three you actually meant.
This page resolves that. It is not a marketing comparison — it is the disambiguation you need before you can read any StepAudio 3 spec sheet correctly, because the three names above carry different prices, different endpoints and different failure modes.
Why the name is ambiguous
StepFun shipped the whole audio stack in one day, and the naming convention made the family name the stem of every product name. That is tidy on a launch slide and awkward everywhere else. It means a search for the family name returns a mixture of five different products, and a benchmark row labelled "StepAudio 3" could plausibly be any of them.
The practical consequence is that you cannot tell from a headline which thing was measured. A post that says "StepAudio 3 tops the transcription leaderboard" is describing StepAudio 3 ASR. A post that says "StepAudio 3 wins on conversational dynamics" is describing StepAudio 3 Realtime. Both are true, they are different products, and they are not interchangeable.
There is a second ambiguity inside the ASR line specifically. StepFun's documentation refers to an ASR Max tier as its largest ASR model to date, with its own pricing and its own endpoint, while the public leaderboard rows carry the plainer label. Working out whether those are one SKU under two names or two SKUs is the most useful thing this page can do, and the next section does it.
The three things the name can mean
• StepAudio 3 ASR — the transcription product. A streaming endpoint that returns incremental text as it decodes and a final transcript at the end. StepFun describes it as transcription with a language model attached for context, tuned for medical, legal, finance, automotive and programming audio, and for hard inputs like whispered speech, fast speech, connected speech, singing and background music. It is the model sitting at 1.7% word error rate on Artificial Analysis's AA-WER index for non-streaming speech to text.
• StepAudio 3 ASR Max — the tier StepFun's documentation calls its largest ASR model to date. It carries the published list rate of 2.8 yuan per hour. It is not a cheaper transcriber; it is the top of the ASR line.
• StepAudio 3 Realtime — the full-duplex conversational model. It reasons over speech directly rather than running a transcribe-then-reason chain, and it transcribes as a by-product of that. It is not an ASR product, and its accuracy on transcription tasks is not what it was tuned for.
Everything else in the family — TTS, Gen, Music — is a different job entirely and should not appear in a transcription comparison at all.
Are ASR and ASR Max the same model?
The evidence points to yes, and the check is arithmetic. StepFun lists the ASR Max tier at 2.8 yuan per hour. Converted at roughly 7.1 yuan to the dollar that is about $0.39 per hour, or roughly $6.6 per 1,000 minutes. Artificial Analysis lists the model on its leaderboard at $6.67 per 1,000 minutes. Two independently published figures, one from the vendor's price list and one from the third-party board, landing within a cent of each other. That is what you would expect if the leaderboard row and the ASR Max list rate describe the same SKU.
It is worth being precise about what that does and does not prove. It proves the leaderboard's price axis and the vendor's rate card are describing the same product at the same price. It does not prove the leaderboard's 1.7% was measured on the exact endpoint you would call, because the board does not publish its decoding configuration or the endpoint it hit. So: same product, very likely; same measurement conditions, unverified.
What is certain is the generational jump inside the line. StepAudio 2.5 ASR sits at 4.7% on the same board at 0.15 yuan per hour. The current tier is 1.7% at 2.8 yuan per hour. That is a 64% reduction in word error rate for roughly nineteen times the rate — the sharpest trade-off in the family and the one that decides whether the upgrade is worth it.

The dimensions that actually differ
Once the naming is settled, the comparison collapses to a short list. These are the ones that change a decision:
• Accuracy — StepAudio 3 ASR at 1.7% AA-WER non-streaming vs StepAudio 2.5 ASR at 4.7% on the same board. Measured by Artificial Analysis, not by StepFun.
• Price — 2.8 yuan per hour vs 0.15 yuan per hour. Both vendor list rates, both current. Roughly 19x.
• Weights — hosted API only for both. No open weights anywhere in the family, no on-premise path, no self-hosting option at any tier.
• Endpoint shape — a single streaming transcription endpoint returning incremental and final text vs the same shape in the previous generation. Nothing about the integration model changed, so a migration is a model identifier swap rather than a rewrite.
• Hard-audio tuning — StepAudio 3 ASR is explicitly tuned for whispered, fast, connected and sung speech and for audio with music underneath. That tuning is the product; the previous generation was not built for it.
• Vendor-claimed domain gains — StepFun reports Chinese down 9.3%, English down 25.0%, dialects down 22.3%, verticals down 42.1% and code-switching down 27.9% generation over generation. Vendor-reported and unreproduced, so treat them as directional.
Notably absent from that list: context length, audio format support, maximum audio duration and a model identifier. StepFun's public documentation for this model does not state them, and anyone quoting those numbers is quoting something else.
ASR vs Realtime is the comparison people actually need
If you are choosing between transcription products rather than between generations, the interesting matchup is not ASR against ASR Max — it is ASR against Realtime. StepFun's own technical material says the two share their pretraining and midtraining stages and diverge only during supervised fine-tuning. Same base, different fine-tune, different product.
That single fact has a practical reading. The 1.7% word error rate is a property of the ASR fine-tune. It is not a promise about the realtime model, which optimises for turn-taking, interruption handling and reasoning over speech rather than for transcript accuracy. If your architecture transcribes as a side effect of a live conversation, you are not buying the 1.7% — you are buying a conversational model that happens to emit text.
The reverse is also true and less obvious: because they share a backbone, the ASR model is not a small specialist model bolted onto the family. It is the same scale of system, fine-tuned differently. That is why the ASR rate is close to the rest of the family rather than near the cheap end of the market.
What to do with this if you are picking one
Three questions settle it. Do you need a transcript, or do you need a conversation? If a transcript, StepAudio 3 ASR is the product and the family name is noise. If a conversation, you want StepAudio 3 Realtime and you should not be reading ASR benchmarks at all. And if you need a transcript of genuinely hard audio — accented, whispered, sung, or with music under it — the 19x premium is the price of the tier that was tuned for it, and the honest test is your own audio rather than a blended index.
The decision that is easy to get wrong is treating the transcription layer as permanent. Whatever you pick, the transcript is an input to something else — a summariser, a structured extraction, a compliance check, a search index — and those calls land on ordinary text models. That is the layer that scales with usage rather than with audio minutes, and it is the layer worth keeping portable from the start. OrcaRouter puts nearly 200 text models behind one API with automatic failover across providers, a routing DSL that composes several into a single call, and model fusion when one model's judgement is not enough. Where a provider publishes a rate, that list price is passed through with no markup, so a price change on the text side reaches your bill the day it is announced.
Being clear about scope, since this page is about a family we do not serve: no StepAudio 3 endpoint is on OrcaRouter. The transcription and voice models are reached through StepFun's own API. What OrcaRouter carries is the reasoning layer downstream of the transcript, which is where a switching decision is cheap to make and expensive to defer.

What nobody has settled
The naming is resolvable. The performance claim is not, yet. There is still no published technical report for the ASR model, no released test set, no decoding configuration and no reproducible protocol from anyone outside StepFun, and the vendor's own comparisons are against other Chinese ASR products rather than against the open-weights incumbent. The 1.7% is a real third-party measurement on a blended three-dataset index; everything else about this model is a vendor claim.
Two things would change the picture. A head-to-head against Whisper Large v3 Turbo on identical audio, which nobody has published. And a rate for the rest of the family — four of the five StepAudio 3 models were still on limited-time free preview pricing at launch, and only transcription and synthesis carried published rates, which means the price list you can read today covers two of five products.

That is the layer that scales with usage rather than with audio minutes, and it is the layer worth keeping portable from the start. automatic failover across providers, a routing DSL that composes several into a single call, and model fusion when one model's judgement is not enough.
