
Step 5 Preview: StepFun's Unannounced 600B Flagship Is Already on the Leaderboard
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Step 5 Preview has a leaderboard rank, a published price, a measured output speed and an Intelligence Index score of 44 — the same score as Kimi K3, a model roughly five times its size. What it does not have is a launch. StepFun has not announced it, StepFun's own platform documentation does not list it, and there are no weights to download. Every figure below comes from a third party that got access before the vendor said anything publicly, which makes this a what-we-know-so-far piece rather than a review. Read the numbers as evidence of what is coming, not as a spec sheet.
The leak is the story
The first public trace of Step 5 Preview is an evaluation run. Artificial Analysis has a live model page for it, scored through StepFun's own API under the test tag step-5-preview-b, with the platform dating the model to September 18, 2026. That page is where the 44, the pricing, the speed measurement and the benchmark rows in this article come from.
Against that, the vendor side is empty. StepFun's developer documentation lists models up to Step 3.7 Flash and Step 3.5 Flash and stops — no step-5, no preview entry. The stepfun-ai Hugging Face organisation has no Step 5 repository, which is consistent with a closed-weights flagship rather than an oversight. Community reports on Chinese developer forums point to a possible unveiling at the AGI Frontier event in Shanghai on September 19, which would be roughly a day after the evaluations surfaced. As of writing, nobody outside StepFun has an official specification, an official price or an official name for the shipping version.
The naming itself is the first clue that this is a preview and not a launch. StepFun skipped the Step 4.x line entirely and went to Step 5. A model released under a Preview suffix, benchmarked before it has a product page, is a model the vendor is still calibrating — pricing, routing and limits can all move before general availability.
What the specification looks like, as far as anyone can tell
These are the figures circulating, and they are the leak's most repeated claims — which is not the same as its most confirmed ones:
• Total parameters — about 600B, against roughly 2.8T for Kimi K3
• Activated per inference — about 27B, roughly 4.5% of the total, an unusually sparse mixture-of-experts ratio
• Input modalities — text and images; output is text only
• Reasoning — yes, with extended thinking
• Context window — 1M tokens on Artificial Analysis; a separate third-party configuration circulating in developer tooling lists 350,000 with a 64,000 maximum output
• Weight availability — closed, no repository
That context-window contradiction is worth flagging plainly, because it has not been resolved by anyone. A 1M-token window and a 350k-token window are different products with different memory costs, and the second figure comes with a 64k output cap that the first says nothing about. If you are planning a long-document pipeline on the strength of the 1M number, the 350k configuration is the reason to wait for a vendor page. The parameter counts have a similar softness: DataLearner's tracker still records MoE architecture and parameter counts for Step 5 Preview as unavailable, meaning the 600B/27B pair — the single most quoted fact about this model — is also one of the least officially confirmed.
The interesting number is 27B, not 600B
Sparse mixture-of-experts models only spend compute on the experts a token actually routes to, so the activated figure is the one that governs how the model behaves in production. At roughly 27B active, Step 5 Preview is doing per-token work in the same band as models a fraction of its size, while the 600B total is largely a memory-footprint question.
Two consequences follow, and both show up in the third-party measurements. Serving cost tracks activated parameters, which is part of why the price is where it is. And per-token speed benefits too: Artificial Analysis measures 99.8 output tokens per second, ranked 42nd of 200 models, against a median of 64.6. Time to first token comes in at 2.96 seconds against a median of about 3.57 seconds. If the 600B/27B split is accurate, this is a model designed to be cheap to serve rather than cheap to host — an important distinction if you are the one paying for the GPUs rather than the tokens.
Where it lands on the Intelligence Index
Artificial Analysis scores Step 5 Preview at 44 on Intelligence Index v4.3, which incorporates ten evaluations. That is 25th of 200 models on the board and well above the median model the index scores, which sits at 25.
• Intelligence Index — Step 5 Preview 44 vs Kimi K3 44
• Directly above — GLM-5.3 and Claude Opus 5, both at 45
• Directly below — DeepSeek V4.1 Flash at 40 and DeepSeek V4 Pro at 36
• Terminal-Bench 4.0 — Step 5 Preview 33.3% vs Kimi K3 at roughly 12.6%, and vs DeepSeek V4.1 Flash at 26.8%
• SciCode — 59% for Step 5 Preview, level with Kimi K3
• Output speed — 99.8 tokens per second vs a field median of 64.6
The headline is the tie with Kimi K3, and the honest reading is narrower than the headlines suggest. Matching a 2.8T-parameter model on an aggregate index at 600B total is a real efficiency result, but the aggregate hides the shape of it: the Terminal-Bench 4.0 gap favouring Step 5 Preview is dramatic, while the SciCode result is a dead heat. Aggregate parity on an index is not parity everywhere, and a preview build is the least reliable moment to assume the gaps will hold.
One more discrepancy worth naming: Artificial Analysis reports 44, while DataLearner's independent tracker records the same run as 43.6. That is a rounding difference rather than a disagreement about capability, but it is the sort of thing that makes the point about this model's numbers generally — they are third-party readings of an unannounced model, taken at slightly different moments, and none of them have been confirmed by StepFun. None of these benchmarks are vendor-reported, which cuts both ways: nobody at StepFun has claimed a number here, and nobody at StepFun has stood behind one either.

The price, and the verbosity that eats into it
Pricing is the part of this leak that will affect a budget soonest. Through StepFun's API, Artificial Analysis records the following:
• Input — $1.00 per million tokens, against a median of $1.88 for comparable models
• Output — $2.70 per million tokens, against a median of $10.00
• Cache — a 95% cache discount, putting cache hits at roughly $0.05 per million tokens
• Blended — about $0.51 per million tokens at a 7:2:1 cache-hit-to-input-to-output mix
• Cost per Intelligence Index task — $0.71
• Cost of a full index run — $918.34 in total
Output at $2.70 is the line that matters most, because it is under a fifth of what Kimi K3 charges for the same million tokens at $15.00. But there is a catch that a per-token comparison hides, and Artificial Analysis measures it directly: verbosity. Step 5 Preview generated 160M output tokens to complete the Intelligence Index, against a median of 90M for the field and a rank of 65th of 200 on that measure.
Work that through. At $2.70 per million, 160M output tokens is about $432 — roughly half the $918.34 total cost of the whole index run, spent on the model's own verbosity rather than on the tasks. Push the same 160M tokens through Kimi K3's $15.00 output rate and you are at $2,400, which is where the price advantage comes from. But the practical warning is for anyone estimating costs by lifting a token count from a different model: Step 5 Preview writes about 1.8 times the median, so an output-heavy workload buys the cheaper rate and more tokens at the same time. The discount survives, but it is smaller than $1.00/$2.70 against $3.00/$15.00 makes it look, and the size of it depends entirely on how chatty your prompts make the model.
The 95% cache discount is the other lever, and it is unusually aggressive. A workload with heavy prompt reuse — a long system prompt, a fixed document, a shared codebase — lands much closer to the blended $0.51 than to the $1.00 input rate. That blended figure assumes a 7:2:1 ratio, which is a modelling assumption rather than a guarantee, but it is the right shape of arithmetic to run against your own traffic.

What early testers are running into
These are individual reports from developer forums and social posts in the days around the leak, not measurements, and they should be weighted accordingly:
• Tool calling is the most common complaint — wrong tool names selected, and edits attempted without reading the target file first
• Errors reported past roughly 340,000 tokens of context, which is the failure mode to watch if the 1M window turns out to be the real number
• Content filtering truncating outputs on some prompts
• Capability impressions split: some testers put it above GLM-5.3 Flash, others place it below DeepSeek V4.1 Flash
An unreleased preview build failing at tool use is not a scandal — it is what the preview label is for. But it does mean the 44 should not be read as a promise about agentic reliability, because the Intelligence Index does not test the thing the early testers are complaining about most.
How you would actually call it — and what to use in the meantime
OrcaRouter does not route Step 5 Preview. There is no public endpoint and no weights, so there is nothing to route, and no third-party platform can honestly claim otherwise yet. Anyone telling you where to call it today is describing an evaluation endpoint, not a product.
The models it is being measured against are a different matter. Kimi K3, GLM-5.3 and DeepSeek V4.1 Flash are all available through OrcaRouter, alongside 199 models from 15 providers behind a single API key and a single bill. Pricing is passed through at the provider's list price with 0% markup, which matters more than usual on a model like this one: if Step 5 Preview's $1.00/$2.70 holds at launch and then moves, a pass-through route moves the same day rather than on someone else's repricing schedule.
When it does become callable, the sensible way to run an unreleased-adjacent model is the way you would run any model you do not yet trust — behind a route with automatic failover, so a tool-call failure or a long-context error falls through to a model that works instead of taking your application down with it. That is the pattern the early reports argue for here specifically: a preview build with reported tool-calling problems and long-context errors is exactly the model you want a fallback behind, not one you want on a production path by itself.

What would turn this from a leak into a release
Three things to watch, in order of how much they would settle. First, a model page and pricing on StepFun's own developer platform — the moment step-5 appears in the model list, the vendor spec sheet replaces every third-party reading in this article, including the 600B/27B pair and the 1M context claim. Second, the resolution of the 1M-versus-350k context contradiction; a 64k output cap under a 1M window is a meaningful difference for anyone building long-document or agentic pipelines, and it is currently unresolved. Third, whether the pricing is a launch posture or a permanent line — preview pricing on Chinese flagship models has a history of moving once capacity gets tight, and $2.70 per million output tokens is aggressive enough to invite that question.
Until at least the first of those lands, the honest summary is that Step 5 Preview looks like a genuinely efficient model — 600B total doing 2.8T-class aggregate work, at a price that undercuts the field by several times — with an unverified specification, a real verbosity tax, and a tool-calling reputation that has not been earned yet. Worth planning around. Not yet worth betting a production path on.
Compared in this article4
Detected from this article · Benchmarks: Artificial Analysis · updated daily
