
Step 5 Preview vs DeepSeek V4 Pro: One Is a Promise, the Other Is an MIT License
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Everything in this matchup turns on a question that has nothing to do with benchmarks: what can you actually download? DeepSeek V4 Pro — the build the vendor shipped as DeepSeek-V4-Pro-0813 on August 13, 2026 — is a 1.6-trillion-parameter mixture-of-experts model under an MIT licence, with weights sitting on Hugging Face right now. Step 5 Preview, which StepFun announced on September 20, 2026, is a 600-billion-parameter sparse MoE that scores eight points higher on the same independent board and has a Hugging Face repository containing exactly one file, .gitattributes. StepFun says the BF16 checkpoint arrives on October 15. So the honest shape of this comparison is not "which model is better." It is "which of these two can you build on this quarter, and which one are you waiting on."
That framing is less exciting than a benchmark shootout and considerably more useful, because the eight-point gap and the twenty-five-day wait are not the same kind of fact. One of them you can plan around today.
Eight points, and the four that matter
Artificial Analysis runs both models through Intelligence Index v4.3.2 — ten evaluations — which is the only place these two have been measured by the same hand.
• Intelligence Index — Step 5 Preview 44 vs DeepSeek V4 Pro 36
• Rank on that board — Step 5 Preview 24th of 200 vs DeepSeek V4 Pro 7th of 113 in its class
• Terminal-Bench 4.0 — Step 5 Preview 33.3%; DeepSeek V4 Pro is not listed on the same board, so there is no like-for-like agentic row to quote
• Output speed — Step 5 Preview 99.8 tokens/sec vs DeepSeek V4 Pro 88.8 tokens/sec
• Time to first token — Step 5 Preview 2.96s vs DeepSeek V4 Pro 1.69s
• Price per 1M tokens — Step 5 Preview $1.00 in / $2.70 out vs DeepSeek V4 Pro $1.32 in / $3.96 out at DeepSeek's peak rate, halving to $0.66 / $1.98 off-peak
• Cache discount — Step 5 Preview 95% vs DeepSeek V4 Pro 97%
• Context — both 1M tokens; DeepSeek publishes a 384,000-token output ceiling, StepFun has not published one
• Weights — Step 5 Preview closed, BF16 checkpoint scheduled for October 15; DeepSeek V4 Pro MIT, downloadable now
Two rows cut against the headline. DeepSeek V4 Pro answers its first token in 1.69 seconds against Step 5 Preview's 2.96 — a 43% head start on every call, which on a long agent loop compounds into real wall-clock time. And its cache discount is two points better, which matters most for exactly the workload both of these models are sold for: an agent re-sending the same system prompt and repository context on every turn lives almost entirely inside that discount.
The eight-point aggregate gap is real, and it is also narrower than eight points sounds. An index is a composite of ten evaluations, and the composition is where the two models separate. Step 5 Preview's 33.3% on Terminal-Bench 4.0 is a genuine agentic result; DeepSeek V4 Pro's own headline strength is long-horizon reasoning at a price nothing else in the open-weights tier matches. Neither board publishes a row where the two are directly opposed, which is the honest reason this comparison cannot be settled by a single number.

What "open" buys you when the vendor changes its mind
Here is the part of this matchup that a spec sheet cannot carry, and it is worth stating with the actual dates because it happened eleven days before Step 5 Preview was announced.
DeepSeek told users it was retiring V4 Pro. On September 9 the company said requests would be routed to the newer DeepSeek V4.1 Flash; on September 10 it firmed the date to 12:00 Beijing time on September 14. On September 11 it reversed — V4 Pro API service continues past September 14, billing unchanged, in DeepSeek's own words. The model string, the 1M-token context and the 500-request concurrency cap are all still what they were, and DeepSeek's Models & Pricing page carries the reversal as a footnote on the deepseek-v4-pro row.
Read that as a demonstration of the thing open weights actually protect you from, and of the thing they do not. The API was almost taken away by a scheduling decision. The weights could not be. An MIT-licensed checkpoint on your own hardware has no vendor who can announce its retirement, and that is a different kind of guarantee from a price promise — it is not a promise at all.
Which is precisely the gap Step 5 Preview is sitting in. StepFun has committed to October 15 for the BF16 checkpoint, and the repository name it has already reserved — stepfun-ai/Step-5-Preview-BF16 — is the format you fine-tune and quantise from, named at creation and filled later. That is a scheduling signal, not an artefact. Until the repository holds something other than a .gitattributes, Step 5 Preview is a model you rent, on one vendor's terms, with no licence to read and no fallback if the terms change.
The two ways to be the cheap option
Both of these models undercut the closed frontier by a wide margin, and the mechanisms are not the same, which matters for how durable the price is.
DeepSeek V4 Pro is cheap because a lab published weights and a competitive serving market drove the margin out of them. That is a structural floor: anyone can serve the checkpoint, so the price is set by the cost of a GPU-hour rather than by a company's pricing strategy. DeepSeek's own API prices it in two windows — $1.32 and $3.96 during peak hours, half that outside them — and peak is narrow, a few hours on weekday mornings in UTC, so most traffic lands on the cheaper rate and weekends never see the higher one.
Step 5 Preview is cheap because StepFun decided to price it that way. There is one endpoint, one price list and no second source. That is not an accusation — a 600B/27B sparse model genuinely costs less to serve than its parameter count suggests, and the activated-parameter ratio is the reason. It is simply a different kind of price, and the difference shows up the day StepFun decides the preview has done its job.
The price gap between the two is small enough that it should not decide anything: $1.00 and $2.70 against $1.32 and $3.96 is a 24% difference on input and 32% on output, well inside the range that a cache discount, an output-length difference or a retry rate can erase. On the verbosity question the two are close — Step 5 Preview generated 160M output tokens on the index run against a 92M median, and DeepSeek V4 Pro's own run is in the same band. Neither is a frugal model, and both are sold on being cheap per token rather than cheap per task.

Where each one belongs in a real stack
If you need a model you can fine-tune, quantise, host in your own VPC, or hand to a legal team with a licence file they can read, this is not a close call and it is not about benchmarks. DeepSeek V4 Pro is available under MIT today. Step 5 Preview has no licence, no config and no parameter file, and the earliest you can plan around is mid-October. The eight-point index gap does not change that arithmetic, because a model you cannot download cannot be deployed.
If what you want is the strongest agentic result available at a dollar an input million, Step 5 Preview is ahead on the one board that has measured both, and it is the better default for the trial-and-error work that fills most of an agent loop — extraction, classification, test scaffolding, the routine calls between the two that matter. Its 99.8 tokens per second also finishes output faster than DeepSeek V4 Pro's 88.8, which shortens the loop itself rather than just the bill.
The awkward case is the one most teams are actually in: you want the Step 5 Preview number but you cannot put a preview endpoint that opened on launch day, with no published output ceiling, on a production path. That is a routing problem, and it is the one this comparison ends on. DeepSeek V4 Pro is reachable through OrcaRouter at $0.66 and $1.98 per million tokens, the off-peak half of DeepSeek's $1.32 and $3.96 rate card, and the surrounding text models alongside it — behind one key for more than 200 models with 0% markup, so a DeepSeek repricing lands on your bill the same day rather than at renewal. Point the exploratory slice of traffic at the preview and hold the production path on the model with a licence behind it, with automatic failover across providers doing the work that a preview endpoint's capacity curve makes necessary.

What would settle it
Three things, and all three are checkable rather than arguable.
The first is the licence. When the BF16 checkpoint lands, what does the repository say a business is allowed to do with it? A permissive licence closes most of this gap on its own, because it makes the eight-point lead something you can own rather than rent. A restrictive one leaves DeepSeek V4 Pro as the only model in the pair you can actually build a product on.
The second is the output ceiling. DeepSeek publishes 384,000 tokens. StepFun's announcement says 1M context and does not mention an output cap, and a separate third-party configuration that circulated during the leak listed 350,000 context with a 64,000 output maximum. For the long-horizon agentic work both models are sold for, that number is not a footnote.
The third is whether the price holds once the preview suffix comes off. A preview price is a customer-acquisition number; a shipping price is a business model. DeepSeek V4 Pro's floor is set by a competitive serving market and cannot move much. Step 5 Preview's can move on a Tuesday.
Until the first two are answered, the practical read is that DeepSeek V4 Pro is the model you deploy and Step 5 Preview is the model you benchmark against — with the caveat that a preview endpoint days old, sitting eight points above a mature MIT-licensed flagship, is a real result, and the reason October 15 is worth a calendar entry rather than a shrug.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
