
Step 5 Preview vs Nemotron 3 Ultra: Twenty-One Index Points for Twenty Cents
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 113 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 52 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 62 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 399 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
One of these models you can download this afternoon, and it is the one that loses by twenty-one points. Step 5 Preview, StepFun's closed flagship, opened its API on 20 September 2026 and has no licence file you can hold; NVIDIA Nemotron 3 Ultra 550B A55B has been on Hugging Face since 4 June 2026 under a permissive licence with its training recipes attached. On the Artificial Analysis Intelligence Index the pair reads 44 against 23. On output price it reads $2.70 against $2.50 per million tokens. Twenty-one points for twenty cents is not a comparison that resolves cleanly in either direction, which is why it is worth doing properly rather than skimming the two headline columns and picking a side.
Neither of these is a launch this week, and that matters for how you read the numbers below. Both have been callable for months, both have been through the same independent harness, and neither has a rate-card change in the window. What has changed is smaller and more interesting: the index they are scored on was rebuilt in September, and the two of them are now being judged against two different medians drawn from two different populations. That is where this comparison actually gets decided.
Two very different kinds of product
Read the spec sheets side by side and the first thing to notice is that they are barely the same category of thing.
• Parameters — Step 5 Preview is about 600 billion total with 27 billion active per token on StepFun's own documentation; Nemotron 3 Ultra is 550 billion total with 55 billion active, per the figure the independent board records against its configuration.
• Architecture — StepFun does not publish an internals description for Step 5 Preview. NVIDIA's card documents a hybrid: interleaved Mamba-2 layers, selected attention layers, a LatentMoE arrangement and Multi-Token Prediction heads.
• Context window — 1M tokens on Step 5 Preview's specification table, with a maximum output of 64,000 tokens, also documented by StepFun. Nemotron 3 Ultra is recorded at 262,144 tokens by the evaluator that ran it; the vendor card advertises up to 1M, reachable through a serving configuration rather than by default.
• Input — text, images and video on Step 5 Preview, up to 60 images per request and MP4 files under 128MB. Text only on Nemotron 3 Ultra.
• Reasoning control — a documented low, medium and high effort setting on the StepFun side; a chat-template flag on the NVIDIA side that turns the thinking trace on or off.
• Weights — none published for Step 5 Preview, which is proprietary and API-only, with a BF16 checkpoint StepFun has said will follow on 15 October 2026. Nemotron 3 Ultra ships BF16 and NVFP4 builds and a GenRM reward model on Hugging Face under OpenMDW-1.1.
• Deployment floor — not applicable for the API model; eight B200-class or GB200-class GPUs, or sixteen H100s, on the open one, per the vendor's minimum line.
• Rate card — $1.00 input, $0.10 on a cache hit and $2.70 output for Step 5 Preview, from StepFun's English pricing page. About $0.60 input, $0.16 on a cache hit and $2.50 output for Nemotron 3 Ultra, and note carefully where that pair comes from: NVIDIA publishes no rate card, so it is the observed provider pricing the independent board records, not a vendor number.
The row most people treat as decisive is the first one, and it is the least informative of the eight. The two totals sit within nine percent of each other while active parameters differ by a factor of two, and active parameters are what govern the compute cost of each generated token. Neither figure predicts a twenty-one-point gap.

The median is a choice made before you read the score
The independent board does not score models against a single field. It buckets them — and the bucket a model lands in changes the sentence printed under its score, without changing the score.
Step 5 Preview is placed in the general reasoning class, which the board counts at 227 models, and its comparable-model median is 26. Nemotron 3 Ultra is placed in the open-weights class, counted at 117 models, where the median is 18. So the same board describes 44 as "well above average" and 23 as "above average," and both descriptions are true, because 44 is eighteen points clear of its median and 23 is five points clear of its own.
This is not a trick and it is not the board being unfair. It is the correct way to present an open model — you compare a downloadable checkpoint against other downloadable checkpoints, because a team that can only take a download is not choosing from the 227. But it does mean that anyone quoting "above average" about Nemotron 3 Ultra and "above average" about a 44 is comparing two different sentences to each other. If you want one number for the pair, use the raw index and accept the twenty-one-point gap; if you want the decision, use the class and admit you have already chosen your field.

What one harness measured on both
Vendor tables cannot be subtracted from each other, because StepFun and NVIDIA ran different suites on different days and NVIDIA's materials do not carry every benchmark StepFun reports. The honest ground is the intersection: what a single evaluator ran against both.
On the Artificial Analysis Intelligence Index that intersection is 44 against 23. On cost per completed index task it is $1.03 against $0.60. On output speed it inverts, and sharply: 87 tokens per second for Step 5 Preview against 135.6 for Nemotron 3 Ultra, so the model that scores nearly twice as high is the one that answers roughly half a second of generation slower per token.
The sharpest single row is Terminal-Bench 4.0. Step 5 Preview scores 33.3 percent there. Nemotron 3 Ultra scores 0.5 percent. That is not a rounding artifact on either side and it is not a subtle gap — on that specific agentic-terminal suite the open-weight flagship effectively does not register. The trap is that the same board also lists the same model at 53.9 percent on Terminal-Bench 2.1. Two benchmarks with nearly the same name, two different generations of the same task family, and one model sitting at 53.9 on the older one and 0.5 on the newer one. If you have seen a Nemotron number in the fifties quoted at you, that is where it came from, and it is not the current suite.
One agreement deserves a mention precisely because it is rare. NVIDIA's own card claims 87.0 percent on GPQA without tools; the independent run measured 86.7 percent. A vendor number and an outside number landing three-tenths of a point apart is what a trustworthy evaluation table looks like, and it makes the 0.5 percent on Terminal-Bench 4.0 easier to accept as real rather than as an evaluator error.
Where the money actually goes on an index run
Per-token prices predict very little about what a workload costs, and this pair demonstrates the point better than most. The board publishes the cost decomposition for each model's full index run, and for both of these the dominant line item is not generation.
Step 5 Preview's $1.03 per task splits into $0.85 of input and $0.17 of output. Inside the input figure, $0.62 is cache reads, $0.22 is cache writes and only $0.014 is fresh, uncached input. Inside the output figure, $0.12 is the reasoning trace and $0.06 is the final answer.
Nemotron 3 Ultra's $0.60 per task splits into $0.53 of input and $0.07 of output, and the composition inside the input line is almost the mirror image — $0.48 of cache writes against just $0.04 of cache reads.
Two conclusions fall out. The first is that on a long-horizon evaluation with heavy prefix reuse, roughly eighty percent of what you pay sits on the input side of the ledger, which makes your cache hit rate and your context discipline the numbers to optimise rather than your output length. The second is that the two models arrive at their bills differently: Step 5 Preview is paying to re-read a large shared prefix, while Nemotron 3 Ultra is paying to write distinct content into the cache in the first place. Similar headline cost, two different invoices — and if your workload looks more like one of those patterns than the other, the ordering at $1.03 and $0.60 can flip.
The verbosity figures explain the rest of the output line. Step 5 Preview generated about 160 million output tokens across the index suite against a median of 82 million, which the board describes flatly as very verbose. Nemotron 3 Ultra generated 110 million against a median of 140 million, which it calls fairly concise. The cheaper model is the terse one, and it is terse in the direction that matters for a bill.

What the download buys, and what it costs to keep
The price of Nemotron 3 Ultra is not $2.50 per million output tokens if you take the checkpoint. That is the price of renting somebody else's deployment of it. Take the download and you have bought a different set of costs and a different set of rights.
The rights are concrete and they are the part an API-only model cannot match at any price. OpenMDW-1.1 comes with the weights, and NVIDIA publishes the pre-training and post-training data collections and the training recipes alongside them. There is an NVFP4 build that is roughly half the size for the same model, and the Ultra weights were pre-trained with an NVFP4 recipe rather than quantised afterwards — which is why the small build is a first-class artifact on the card rather than a community experiment. The commercial position is stated plainly: ready for commercial and non-commercial use.
The costs are equally concrete. The minimum deployment is eight B200-class GPUs or sixteen H100s, and that is the floor the card states rather than a suggestion. The context window you actually get depends on how you configure serving, with 256K as the default and 1M behind an explicit setting. A team that cannot commit a rack is not choosing between these two models at $1.00 and $0.60 input; it is choosing between one model and a purchase order.
There is a version of this comparison where the open model wins on the axis people say they care about and rarely price — dependency. Step 5 Preview is an endpoint you call and a vendor you trust, with a checkpoint promised rather than shipped. If StepFun revises its rate card, tightens its limits, deprecates the ID or has a bad week, your only lever is your own error handling. Nemotron 3 Ultra's weights are already on disk in somebody's cluster, and that is worth something real even while the same model is losing by twenty-one index points.
It is worth being precise about what "promised" means on the other side, because the picture is messier than either camp admits. Multiple third-party mirrors of a BF16 build appeared on Hugging Face on 20 September, the day the API opened, some of them well over a terabyte of safetensors. Those are community re-uploads, not a vendor release, and StepFun has published no open-weights announcement alongside them. What they establish is that the weights are circulating and that the promised 15 October checkpoint would be a formalisation rather than a disclosure.
A number that moved while we were reading it
One figure in this piece changed between two reads of the same page, and it is worth naming because that is exactly how a comparison quietly goes stale.
The independent board showed a cost of $0.71 per index task for Step 5 Preview in coverage dated 9 October 2026. Read again on 10 October, the same page says $1.03. Nothing changed about the model in that window, because nothing can change about an API-only model's weights overnight. What changed is the evaluation accounting: when a vendor's cache price moves, when the evaluator revises its cache-read assumptions, or when a run is recomputed against a different token mix, the per-task figure moves and the index score does not.
So treat cost per task as a live number with a date stamped on it rather than a property of the model. Every per-task figure in this article is the 10 October read and is labelled as one. If you are building a budget from this page, build it from a figure you pull yourself on the day you commit — and if you are comparing two write-ups from different weeks, check the capture dates before you conclude that a model got cheaper.
Which one you would actually call
Say the uncomfortable part first: OrcaRouter routes neither of these two models. Step 5 Preview is reachable through StepFun's own API and a handful of third-party platforms, and Nemotron 3 Ultra is served from NVIDIA's own endpoints and by several providers, and you can check both claims against our catalogue in the same minute you read them.
What that leaves is the shape of the dependency rather than the choice of model, and it is the one place a routing layer earns its keep here. An API-only preview on a single vendor path is precisely the arrangement where a bad afternoon at the provider becomes your outage, and the cheap insurance is a control plane that can fail a request over to a qualified fallback the moment a provider path degrades. That is what automatic failover is for: not as a substitute for Step 5 Preview, but as the thing you put in front of the models you can actually call, so that a vendor-specific endpoint is never the only road into production. The same catalogue that does that carries 200-plus models behind one key and one bill, with provider list price passed through at 0 percent markup — which is the other half of the argument, because a rate-card change anywhere upstream is live on our side the same day rather than at the next contract renewal.
If neither model is the one you will call, the short version runs like this. Take Step 5 Preview when you want the score, the video and image input and the 90 percent cache discount, and accept that you are renting a preview with no licence to the weights and an October checkpoint you have not seen yet. Take Nemotron 3 Ultra when the weights matter more than the score — for on-premise constraints, for fine-tuning, for a licence you can read — and budget the rack before you budget the tokens, because at $0.60 a million input tokens a 550-billion-parameter deployment is only cheap if somebody else is already paying for the GPUs.
The thing to watch is 15 October. If StepFun publishes the BF16 checkpoint it has promised, this article's central asymmetry — that only one of these two is a download — stops being true, and the twenty-one index points become considerably harder to argue away. If the date slips, the case for the open-weight model strengthens by default, which is a strange way for a capability gap to close and a very common one.
