Generated title card for Step 5 Preview vs Qwen 3.8 Max Prime reading 'one index point, four dollars and thirty-eight cents', with three stat chips: Artificial Analysis Intelligence Index 44 against 45, cost per index task $1.03 against $5.41, and output price $2.70 against $9.902 per million tokens. The OrcaRouter logo sits in a white strip at the bottom right.
Guides & Insights

Step 5 Preview vs Qwen 3.8 Max Prime: One Index Point, Four Dollars and Thirty-Eight Cents

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Start with the thing a search box will not tell you. Qwen 3.8 Max Prime is not a model. It is a serving lane — the same Qwen3.8-Max weights the vendor took to general availability on 3 August 2026, sold under a second model ID at exactly twice the international rate card. Step 5 Preview, StepFun's closed flagship, opened its API on 20 September 2026 and is a model in the ordinary sense: one endpoint, one price, one set of weights you cannot hold. So the honest version of this matchup pits a measured model against an unmeasured tier, and the arithmetic that falls out is blunter than either vendor would put on a slide. On the independent Artificial Analysis Intelligence Index the two sit one point apart, 44 against 45, and on what a completed task costs they sit $1.03 against $5.41 apart. One point, four dollars and thirty-eight cents, per task — and that is before you have bought any of the speed Prime actually charges for.

What follows is that number taken apart: where it comes from, why the standard Qwen ID is already the pricier half of the comparison before Prime doubles anything, and what the extra four dollars per task is and is not buying.

The name is doing work the product does not

Alibaba announced Prime — 优速模式, fast mode — at the Yunqi Conference on 22 September 2026 and published it as a row on the Model Studio rate card. Alibaba's own documentation is unambiguous about what changes: output speed rises to 1.5 to 2 times the standard API, and, in the vendor's words, the capabilities and usage restrictions are identical to the original model. No new parameter count, no new endpoint contract, no capability the standard ID lacks. You change the model name in the request body and you are on the fast lane.

Generated two-column comparison scoreboard for Step 5 Preview and Qwen 3.8 Max Prime across six shared rows: what the name refers to, 'a model' against 'a serving lane for Qwen3.8-Max'; Artificial Analysis Intelligence Index 44 against 45, measured on the standard Qwen ID with no Prime row; cost per index task $1.03 against $5.41; output price $2.70 against $9.902 per million tokens; output speed 87 against 35.4 tokens per second, again measured on the standard ID; and weights, 'proprietary, BF16 checkpoint promised 15 October' against 'proprietary, none'. A footer reads that index, speed and cost-per-task figures are per Artificial Analysis and were measured on the standard Qwen3.8-Max ID rather than on Prime, and that rate cards are Alibaba's international list price and StepFun's own. The OrcaRouter logo sits in a white strip at the bottom right.

That makes the search term a trap. Someone looking for a Qwen flagship newer than Qwen3.8-Max finds "Prime" and reads it as a version number. It is a price. Everything the name implies about capability is supplied by Qwen3.8-Max itself: a 2.4-trillion-parameter sparse mixture-of-experts model with roughly 95 billion parameters active per token, 512 experts per layer, and a 983,616-token context on the configuration the independent board tested.

Step 5 Preview has no equivalent ambiguity, and it has the opposite problem. StepFun lists it at about 600 billion total parameters with 27 billion active, text, image and video input, a 1M-token context window and a documented 64,000-token output ceiling, with reasoning effort selectable at low, medium or high. It is proprietary. A BF16 checkpoint has been promised for 15 October 2026 and has not shipped; community mirrors of a BF16 build have circulated on Hugging Face since the day the API opened, but re-uploads are not a release and StepFun has published no open-weights announcement.

So the pair is genuinely lopsided before a single benchmark lands: a previewed model that might become a download, against a repriced serving tier of a model that will not.

What one index point costs

Both models have been run by the same independent evaluator, and this is where the comparison turns uncomfortable for the cheaper-per-token story. Read on 10 October 2026:

• Index — Step 5 Preview 44, well above a comparable-model median of 26. Qwen3.8-Max on the 0902 build, 45. One point.

• Cost per completed index task — $1.03 against $5.41. Better than five times.

• Output price — $2.70 against $6.00 per million tokens on the standard ID, which becomes $9.902 once you move to Prime. Step 5 Preview is 2.2 times cheaper than standard and 3.7 times cheaper than Prime.

• Input price — $1.00 against $2.00 standard and $3.301 on Prime. The mirror image of the output column, and the one that matters more once you read the cost breakdown below.

• Cache hit — $0.10 per million tokens on Step 5 Preview, a 90 percent discount; $0.25 on Qwen3.8-Max, an 87.5 percent discount the vendor's own page rounds to 88 percent.

• Output tokens per second — 87 against 35.4. Here Qwen is the slow one, and by a factor of roughly two and a half.

The per-token columns and the per-task column disagree, and the per-task column is the one to believe, for a reason that shows up only if you open the cost breakdown. On Step 5 Preview's index run, $0.85 of the $1.03 is input-side and $0.17 is output. On Qwen3.8-Max's, $4.76 of the $5.41 is input-side and $0.65 is output. Sticker prices are roughly a factor of two apart; task bills are a factor of five apart, because the Qwen run pushed about 9.9 billion input tokens through the harness against Step 5 Preview's 5.5 billion, and because most of that volume is cache traffic being re-read at whatever discount the vendor grants rather than at the headline rate.

Step 5 Preview is described by the board as very verbose, at roughly 160 million output tokens across the suite against a median of 82 million — and Qwen3.8-Max is described in exactly the same words, at roughly 190 million against the same median. Neither of these is a model that answers briefly. If output length is what you were hoping to optimise, neither column helps you.

Screenshot of the Artificial Analysis model page for Qwen3.8 Max (0902), showing an Intelligence Index of 45 ranked 32 of 227 with a comparable-model median of 26, speed ranked 190 of 227, cost ranked 105 of 227 and verbosity ranked 103 of 227, pricing of $2.00 per million input tokens and $6.00 per million output tokens with an 88 percent cache discount, a measured cost of $5.41 per index task, 190 million output tokens generated across the run against an 82 million median, about 35 output tokens per second, and a 984k-token context.

The comparison has a hole in it, and it is Prime-shaped

Every figure in the previous section was measured on the standard Qwen3.8-Max ID. None of it was measured on Prime.

That is not a technicality. Prime's entire proposition is that tokens arrive faster, and the board's only speed number for this family is the standard ID's 35.4 output tokens per second — the slowest of the three IDs discussed here by a wide margin, and the one row where Qwen3.8-Max is not merely expensive but visibly underperforming. If Prime delivers the top of Alibaba's stated range, that 35.4 becomes something like 70, still below Step 5 Preview's 87 while costing 3.7 times as much per output token. If it delivers the bottom of the range, it becomes roughly 53.

Those are estimates built on a vendor-stated multiplier, not measurements, and they should be labelled as such. Nobody we can find has published an independent throughput test of Prime mode. The 1.5-to-2× figure is Alibaba's own, and it is the single load-bearing claim under the whole tier.

The one structural detail that argues in Prime's favour is the throttling rule, and it is easy to skim past. Alibaba's documentation states that under Prime, when call volume reaches the rate limit, if the platform still has spare resources the request is not throttled — so the throughput available to you does not fall below the stated limit. That is a soft ceiling with a hard floor: under contention you get the limit, under idle capacity you can get more. For an agent loop firing dozens of sequential calls it matters more than the median, because the slowest call decides when the loop finishes. It is also the clearest explanation of why Prime exists as a product at all. Alibaba is selling queue position out of capacity it has already built, and charging double for the right to be served first from it.

What the two rate cards say when you put them side by side

Alibaba publishes its Qwen rows adjacent to each other, which makes the tier's arithmetic unusually clean. On the international page the Prime row reads $3.301 input and $9.902 output per million tokens, against $1.65 and $4.951 for the standard ID and for its 0902 pin. That is exactly 2.0 on both columns — not a rounded approximation but the literal ratio of the published numbers. StepFun's English pricing page lists Step 5 Preview at $1.00 input, $0.10 on a cache hit and $2.70 output, with cache-miss input including the cost of writing new content into the cache. The vendor's published cache-hit figure is a 90 percent discount; the independent board records 95 percent from the same materials, and it is worth knowing which of the two you are modelling against.

Put the three cards in one column and the shape is clear. Prime output, $9.902. Standard Qwen output, $6.00 on the board's record and $4.951 on Alibaba's own international page. Step 5 Preview output, $2.70. And on the cheap lane nobody advertises, Step 5 Preview's cache read is $0.10 against Qwen's $0.25 — which matters more than the output column for any workload that repeats its prefix.

One warning about stale numbers, because this family has been repriced more than once in seven weeks. The widely quoted $2 input and $6 output for Qwen3.8-Max is the August launch headline, and it is still what Alibaba's Singapore tab shows. The international and global rows now read $1.65 and $4.951. If you are modelling spend, use the rate card for your region and re-read it on the day you commit.

Two ways to read the 2×

Because Prime is the same model at twice the price of the standard ID, the premium is easy to price and hard to justify. At the top of Alibaba's range you pay 2× for 2× the speed, which is cost-neutral per second of latency removed. At the bottom you pay 2× for 1.5×, a 33 percent premium per second saved. Neither end is unreasonable for interactive work; neither is defensible for an overnight batch. That is why the standard advice on the tier — latency-sensitive calls on Prime, everything else on the standard ID — is also the correct advice.

The comparison against Step 5 Preview is where the reasoning changes shape, and this is the part the "vs" framing makes visible. Prime is not competing with Step 5 Preview on price at all. The standard ID is already more expensive per output token than Step 5 Preview is, five times more expensive per completed task, and it scores one point higher. Prime multiplies the cost side of an equation that was already losing and buys back speed that Step 5 Preview has more of for free. If speed is what you need from this family, the honest statement is that the model on the other side of this page gives you 87 tokens per second at $2.70 per million output tokens, and the fast lane gives you a range that tops out, on Alibaba's own numbers, around 70 at $9.902.

That is not an argument that Step 5 Preview wins. It is an argument that Prime's value is entirely internal to Alibaba's own product line, and that the case for it rests on workloads where Qwen3.8-Max's one-point index lead, its 983,616-token context or its different task profile do work that a 1M-token Step 5 Preview context does not. Which is the comparison nobody can run yet, because the Prime row does not exist on any independent board.

Generated rate-card infographic comparing three published price lists in one column: Qwen 3.8 Max Prime at $3.301 input and $9.902 output per million tokens, the standard qwen3.8-max and qwen3.8-max-0902 IDs at $1.65 and $4.951 on Alibaba's international page against the $2 and $6 still shown on its Singapore tab, and Step 5 Preview at $1.00 input, $0.10 cached input and $2.70 output from StepFun's English pricing page. A footer reads that the Prime row is exactly 2.0 times the standard row on both columns and that all figures are vendor list prices read on 10 October 2026.

Where this lands for a buyer

Neither of these two is on OrcaRouter. Step 5 Preview is reachable through StepFun's own API and a handful of third-party platforms; Prime is an Alibaba Cloud Model Studio tier served on a workspace-scoped endpoint; and you can check both claims against our catalogue in the same minute you read them, which is a better test than taking our word for it.

What we do route is a different product in the same family, and it happens to be the one this article's arithmetic keeps pointing back to. The standard Qwen3.8-Max and its dated pin Qwen3.8-Max-0902 sit in the catalogue on the same key as their siblings Qwen3.8-Max-Preview, Qwen3.8-Flash and Qwen3.8-27B, alongside 200-plus other models, with provider list price passed through at 0 percent markup. Two things follow from that, and both bear on the decision above. The first is that a rate-card change at Alibaba appears on our side the same day Alibaba publishes it — precisely the property you want from a vendor that has already repriced this family twice in seven weeks. The second is that your baseline stops being a single-vendor bet: the standard ID is the tier you are most likely to keep on your default path, and keeping it behind a routing layer rather than a direct integration is what makes the Prime decision reversible per endpoint instead of per contract.

If you are choosing between the two names in this article's title, the split is straightforward. Reach for Step 5 Preview when you want the score, the video and image input and the 90 percent cache discount, and accept that you are renting a preview whose checkpoint is a promise for 15 October rather than a licence. Reach for Qwen 3.8 Max Prime only if you are already committed to Qwen3.8-Max and have measured, on your own prompts, that the standard ID's 35 tokens per second is what is costing you money — and price that decision against the 2× rather than against the 1.5×, because the top of a vendor's range is not the number you will see in production.

The measurement that would settle this does not exist, and it is worth saying so plainly rather than dressing it up. Somebody needs to run Prime through a harness that also runs Step 5 Preview, publish a throughput distribution rather than a multiplier, and put a cost-per-task figure next to the index point it buys. Until then the only numbers both sides share are the two rate cards, and they say the same thing from opposite directions: the fast lane is the most expensive way to buy a Qwen3.8-Max token, and the slowest way to buy a second of wall clock against what the alternative charges for the same second.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily