Generated title card for Step 5 Preview reading 'Step 5 Preview — the leak so far', listing six specifications as label-and-value rows: total parameters about 600B, activated per inference about 27B, input text and image, context window 1M tokens, AA Intelligence Index 44, and price $1.00 in and $2.70 out per 1M tokens. A footer reads 'All figures third-party and unaudited - StepFun has not announced this model.' The OrcaRouter logo sits in the bottom-right corner.
Guides & Insights

Step 5 Preview: StepFun's Unannounced 600B Flagship Is Already on the Leaderboard

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Step 5 Preview has a leaderboard rank, a published price, a measured output speed and an Intelligence Index score of 44 — the same score as Kimi K3, a model roughly five times its size. What it does not have is a launch. StepFun has not announced it, StepFun's own platform documentation does not list it, and there are no weights to download. Every figure below comes from a third party that got access before the vendor said anything publicly, which makes this a what-we-know-so-far piece rather than a review. Read the numbers as evidence of what is coming, not as a spec sheet.

The leak is the story

The first public trace of Step 5 Preview is an evaluation run. Artificial Analysis has a live model page for it, scored through StepFun's own API under the test tag step-5-preview-b, with the platform dating the model to September 18, 2026. That page is where the 44, the pricing, the speed measurement and the benchmark rows in this article come from.

Against that, the vendor side is empty. StepFun's developer documentation lists models up to Step 3.7 Flash and Step 3.5 Flash and stops — no step-5, no preview entry. The stepfun-ai Hugging Face organisation has no Step 5 repository, which is consistent with a closed-weights flagship rather than an oversight. Community reports on Chinese developer forums point to a possible unveiling at the AGI Frontier event in Shanghai on September 19, which would be roughly a day after the evaluations surfaced. As of writing, nobody outside StepFun has an official specification, an official price or an official name for the shipping version.

The naming itself is the first clue that this is a preview and not a launch. StepFun skipped the Step 4.x line entirely and went to Step 5. A model released under a Preview suffix, benchmarked before it has a product page, is a model the vendor is still calibrating — pricing, routing and limits can all move before general availability.

What the specification looks like, as far as anyone can tell

These are the figures circulating, and they are the leak's most repeated claims — which is not the same as its most confirmed ones:

• Total parameters — about 600B, against roughly 2.8T for Kimi K3

• Activated per inference — about 27B, roughly 4.5% of the total, an unusually sparse mixture-of-experts ratio

• Input modalities — text and images; output is text only

• Reasoning — yes, with extended thinking

• Context window — 1M tokens on Artificial Analysis; a separate third-party configuration circulating in developer tooling lists 350,000 with a 64,000 maximum output

• Weight availability — closed, no repository

That context-window contradiction is worth flagging plainly, because it has not been resolved by anyone. A 1M-token window and a 350k-token window are different products with different memory costs, and the second figure comes with a 64k output cap that the first says nothing about. If you are planning a long-document pipeline on the strength of the 1M number, the 350k configuration is the reason to wait for a vendor page. The parameter counts have a similar softness: DataLearner's tracker still records MoE architecture and parameter counts for Step 5 Preview as unavailable, meaning the 600B/27B pair — the single most quoted fact about this model — is also one of the least officially confirmed.

The interesting number is 27B, not 600B

Sparse mixture-of-experts models only spend compute on the experts a token actually routes to, so the activated figure is the one that governs how the model behaves in production. At roughly 27B active, Step 5 Preview is doing per-token work in the same band as models a fraction of its size, while the 600B total is largely a memory-footprint question.

Two consequences follow, and both show up in the third-party measurements. Serving cost tracks activated parameters, which is part of why the price is where it is. And per-token speed benefits too: Artificial Analysis measures 99.8 output tokens per second, ranked 42nd of 200 models, against a median of 64.6. Time to first token comes in at 2.96 seconds against a median of about 3.57 seconds. If the 600B/27B split is accurate, this is a model designed to be cheap to serve rather than cheap to host — an important distinction if you are the one paying for the GPUs rather than the tokens.

Where it lands on the Intelligence Index

Artificial Analysis scores Step 5 Preview at 44 on Intelligence Index v4.3, which incorporates ten evaluations. That is 25th of 200 models on the board and well above the median model the index scores, which sits at 25.

• Intelligence Index — Step 5 Preview 44 vs Kimi K3 44

• Directly above — GLM-5.3 and Claude Opus 5, both at 45

• Directly below — DeepSeek V4.1 Flash at 40 and DeepSeek V4 Pro at 36

• Terminal-Bench 4.0 — Step 5 Preview 33.3% vs Kimi K3 at roughly 12.6%, and vs DeepSeek V4.1 Flash at 26.8%

• SciCode — 59% for Step 5 Preview, level with Kimi K3

• Output speed — 99.8 tokens per second vs a field median of 64.6

The headline is the tie with Kimi K3, and the honest reading is narrower than the headlines suggest. Matching a 2.8T-parameter model on an aggregate index at 600B total is a real efficiency result, but the aggregate hides the shape of it: the Terminal-Bench 4.0 gap favouring Step 5 Preview is dramatic, while the SciCode result is a dead heat. Aggregate parity on an index is not parity everywhere, and a preview build is the least reliable moment to assume the gaps will hold.

One more discrepancy worth naming: Artificial Analysis reports 44, while DataLearner's independent tracker records the same run as 43.6. That is a rounding difference rather than a disagreement about capability, but it is the sort of thing that makes the point about this model's numbers generally — they are third-party readings of an unannounced model, taken at slightly different moments, and none of them have been confirmed by StepFun. None of these benchmarks are vendor-reported, which cuts both ways: nobody at StepFun has claimed a number here, and nobody at StepFun has stood behind one either.

Screenshot of the Artificial Analysis model page for Step 5 Preview, headed 'Step 5 Preview Intelligence, Performance & Price Analysis', showing StepFun as the creator, a release date of September 2026, an Intelligence Index of 44 at rank 25 of 200, an output speed of 99.8 tokens per second at rank 42 of 200, input pricing of $1.00 and output pricing of $2.70 per million tokens with a 95% cache discount, $0.71 per Intelligence Index task, verbosity of 160M output tokens, and technical specifications listing reasoning as supported, input modality as text plus image, and a 1M token context window.

The price, and the verbosity that eats into it

Pricing is the part of this leak that will affect a budget soonest. Through StepFun's API, Artificial Analysis records the following:

• Input — $1.00 per million tokens, against a median of $1.88 for comparable models

• Output — $2.70 per million tokens, against a median of $10.00

• Cache — a 95% cache discount, putting cache hits at roughly $0.05 per million tokens

• Blended — about $0.51 per million tokens at a 7:2:1 cache-hit-to-input-to-output mix

• Cost per Intelligence Index task — $0.71

• Cost of a full index run — $918.34 in total

Output at $2.70 is the line that matters most, because it is under a fifth of what Kimi K3 charges for the same million tokens at $15.00. But there is a catch that a per-token comparison hides, and Artificial Analysis measures it directly: verbosity. Step 5 Preview generated 160M output tokens to complete the Intelligence Index, against a median of 90M for the field and a rank of 65th of 200 on that measure.

Work that through. At $2.70 per million, 160M output tokens is about $432 — roughly half the $918.34 total cost of the whole index run, spent on the model's own verbosity rather than on the tasks. Push the same 160M tokens through Kimi K3's $15.00 output rate and you are at $2,400, which is where the price advantage comes from. But the practical warning is for anyone estimating costs by lifting a token count from a different model: Step 5 Preview writes about 1.8 times the median, so an output-heavy workload buys the cheaper rate and more tokens at the same time. The discount survives, but it is smaller than $1.00/$2.70 against $3.00/$15.00 makes it look, and the size of it depends entirely on how chatty your prompts make the model.

The 95% cache discount is the other lever, and it is unusually aggressive. A workload with heavy prompt reuse — a long system prompt, a fixed document, a shared codebase — lands much closer to the blended $0.51 than to the $1.00 input rate. That blended figure assumes a 7:2:1 ratio, which is a modelling assumption rather than a guarantee, but it is the right shape of arithmetic to run against your own traffic.

A generated two-column comparison scoreboard titled 'Step 5 Preview vs Kimi K3 — the scoreboard'. The left column gives Step 5 Preview the rows AA Index 44, total params about 600B, active params about 27B, Terminal-Bench 4.0 33.3%, SciCode 59%, and price $1.00 / $2.70. The right column gives Kimi K3 the rows AA Index 44, total params about 2.8T, active params not disclosed, Terminal-Bench 4.0 about 12.6%, SciCode 59%, and price $3.00 / $15.00. A footer reads 'Step 5 Preview figures third-party via Artificial Analysis and unaudited; Kimi K3 figures per Artificial Analysis. Prices per 1M tokens.' The OrcaRouter logo sits in the bottom-right corner.

What early testers are running into

These are individual reports from developer forums and social posts in the days around the leak, not measurements, and they should be weighted accordingly:

• Tool calling is the most common complaint — wrong tool names selected, and edits attempted without reading the target file first

• Errors reported past roughly 340,000 tokens of context, which is the failure mode to watch if the 1M window turns out to be the real number

• Content filtering truncating outputs on some prompts

• Capability impressions split: some testers put it above GLM-5.3 Flash, others place it below DeepSeek V4.1 Flash

An unreleased preview build failing at tool use is not a scandal — it is what the preview label is for. But it does mean the 44 should not be read as a promise about agentic reliability, because the Intelligence Index does not test the thing the early testers are complaining about most.

How you would actually call it — and what to use in the meantime

OrcaRouter does not route Step 5 Preview. There is no public endpoint and no weights, so there is nothing to route, and no third-party platform can honestly claim otherwise yet. Anyone telling you where to call it today is describing an evaluation endpoint, not a product.

The models it is being measured against are a different matter. Kimi K3, GLM-5.3 and DeepSeek V4.1 Flash are all available through OrcaRouter, alongside 199 models from 15 providers behind a single API key and a single bill. Pricing is passed through at the provider's list price with 0% markup, which matters more than usual on a model like this one: if Step 5 Preview's $1.00/$2.70 holds at launch and then moves, a pass-through route moves the same day rather than on someone else's repricing schedule.

When it does become callable, the sensible way to run an unreleased-adjacent model is the way you would run any model you do not yet trust — behind a route with automatic failover, so a tool-call failure or a long-context error falls through to a model that works instead of taking your application down with it. That is the pattern the early reports argue for here specifically: a preview build with reported tool-calling problems and long-context errors is exactly the model you want a fallback behind, not one you want on a production path by itself.

Screenshot of the OrcaRouter models page at www.orcarouter.ai/models showing 199 models from 15 providers behind one API key and one bill, with filters for input modalities, context length, input price, status, series and supported parameters, and a 'How to call any model' card showing a POST request to the OrcaRouter chat completions endpoint.

What would turn this from a leak into a release

Three things to watch, in order of how much they would settle. First, a model page and pricing on StepFun's own developer platform — the moment step-5 appears in the model list, the vendor spec sheet replaces every third-party reading in this article, including the 600B/27B pair and the 1M context claim. Second, the resolution of the 1M-versus-350k context contradiction; a 64k output cap under a 1M window is a meaningful difference for anyone building long-document or agentic pipelines, and it is currently unresolved. Third, whether the pricing is a launch posture or a permanent line — preview pricing on Chinese flagship models has a history of moving once capacity gets tight, and $2.70 per million output tokens is aggressive enough to invite that question.

Until at least the first of those lands, the honest summary is that Step 5 Preview looks like a genuinely efficient model — 600B total doing 2.8T-class aggregate work, at a price that undercuts the field by several times — with an unverified specification, a real verbosity tax, and a tool-calling reputation that has not been earned yet. Worth planning around. Not yet worth betting a production path on.

Compared in this article4

Detected from this article · Benchmarks: Artificial Analysis · updated daily