
Step 5 Preview vs GLM-5.5: One Model Ships Today, the Other Is Still a Forecast
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Step 5 Preview opened its API on September 20, 2026, the same day StepFun announced it, and you can call it this afternoon from a single endpoint at $1.00 per million input tokens and $2.70 per million output tokens. GLM-5.5 does not exist. There is no model card, no API identifier, no price, no checkpoint and no licence file, and there is no Z.ai page anywhere that uses the name — because GLM-5.5 is community shorthand that the company has never adopted. That asymmetry is the article. What follows is what StepFun actually shipped, what is genuinely established about Z.ai's next flagship, and the head-to-head you can run today in place of the one in the title.
What GLM-5.5 is, and what it is not
The name entered circulation in mid-2026 attached to a release window, and the window has a traceable origin: an analyst forecast relayed on June 25, 2026, that put Z.ai's next flagship in August. It was a watch window, not a commitment from the vendor, and August ended without it. What Z.ai shipped instead was GLM-5.3, announced August 14, with its API live around August 18 and open weights on August 28, followed by GLM-5.3-Flash on August 26. No GLM-5.5 checkpoint has appeared on the vendor's own Hugging Face organisation, where the newest repositories still top out at the GLM-5.3 series.
The naming is unsettled too. GLM-5.3, GLM-5.4, GLM-5.5 and GLM-6 have all been floated in the same window, and Z.ai has confirmed none of them. The only company-side signal is a co-founder's remark about an "epic plus" release, which is a mood rather than a specification. Analysts reading the vendor's cadence — roughly one flagship every 54 to 70 days — now project the next model into a window running October 8 to November 9, 2026, and expect it to be numbered GLM-5.4 rather than GLM-5.5.
Three claims circulate as though confirmed, and it is worth being exact about all three. The parameter count is the weakest: nothing official supports the "more than 1 trillion" figure, and later social posts inflating it past 3 trillion cite nothing at all. The 1M-token context is the safest assumption, because GLM-5.2 and GLM-5.3 both carry it. Open weights are an expectation rather than a promise — Z.ai has shipped weights for its last several flagships, which is a pattern, not a commitment.
What Step 5 Preview actually ships
StepFun's model is the concrete half of this comparison, and the specification is unusual. It is a sparse mixture-of-experts model at 600B total parameters with roughly 27B active per token — an activation ratio near 4.5% — built on a 92-layer narrow-and-deep transformer with sparse grouped-query attention and block-wise token merging. The context window is 1M tokens, input is text and images, output is text, and it reasons with extended thinking. StepFun skipped the entire Step 4.x line to get here, going straight from Step-3.7-Flash to Step 5.
• Intelligence Index — 44 on Artificial Analysis, 24th of 200 models scored, against a median of 24
• Output speed — 99.8 tokens per second, 45th of 200
• Time to first token — 2.96 seconds on the same independent run
• Price — $1.00 per 1M input tokens, $2.70 per 1M output tokens, with a 95% cache discount
• Cost per Intelligence Index task — $0.71, 27th of 200
• Weights — none today; a BF16 checkpoint is scheduled for October 15, 2026
• Licence — none published. The model is proprietary, with no commercial-use grant, redistribution right or fine-tuning permission stated
The price is the eye-catching row, and so is one that does StepFun less favours than it first appears: 160M output tokens on the Intelligence Index against a median of 92M, 64th of 200 on that measure. The model is verbose, and verbosity is a cost multiplier in exactly the agentic workloads it is sold for. The $0.71 cost-per-task figure already folds that in, which is why it is the more honest number to quote than the list price.
StepFun's own evaluations are unreproduced and should be read as vendor material until somebody independent runs them. They put the model at 29.5 on ALE-CLI, 66.4 on FrontierFinance, 83.3 on DRACO, 80.5 on ProgramBench, 67.7 on DeepSWE v1.1 and 49.0 on StepCodeBench, with 33.3% on Terminal-Bench 4.0 and 1571 on GDPval-AA v2. StepFun also reports a 508 TFLOPS MLA peak in a 24-hour GPU kernel optimisation task against 493 for Claude Opus 5. Those are StepFun's numbers about a model StepFun has not opened.

The head-to-head you can run instead of this one
If you want the comparison the title promises, the honest substitute is Step 5 Preview against GLM-5.3 — the model Z.ai actually shipped, still its current open flagship, and the one any GLM-5.5 would replace. On the independent board the two are one point apart. On almost everything else they are not.
• Intelligence Index — Step 5 Preview 44 vs GLM-5.3 45
• Parameters — 600B total / 27B active vs 753B total / 40B active
• Output speed — 99.8 tokens/sec vs 72.1 tokens/sec
• Time to first token — 2.96s vs 2.99s
• Price per 1M tokens — $1.00 in / $2.70 out vs $1.40 in / $4.40 out
• Cache discount — 95% vs 81%
• Cost per Intelligence Index task — $0.71 vs $2.01
• Input modality — text and images vs text only
• Licence — none published vs a custom Z.ai licence that is not MIT
The shape repeats what the Step 5 Preview launch established against every shipping model it has been scored against: a decisive efficiency lead and a statistical tie on capability. It generates tokens about 38% faster, costs roughly half as much per million tokens in both directions, and lands 2.8 times cheaper per index task against GLM-5.3 — while scoring a point lower. On the board, that 44 also puts it level with Kimi K3, which is worth knowing before treating either as a leader.
The one row where Step 5 Preview is unambiguously behind is the one you cannot benchmark. GLM-5.3 has a licence file and downloadable weights; Step 5 Preview has a Hugging Face repository containing a single .gitattributes and a stated date of October 15 for the BF16 checkpoint. If Z.ai's next flagship arrives with open weights under the same custom terms it has used all year, it lands in the only category where Step 5 Preview currently has no answer — which is the real reason a GLM-5.5 comparison is worth waiting for rather than running.

Why waiting for GLM-5.5 is the expensive option
The practical problem with a model that has been two months away for two months is that your evaluation work cannot wait with it. GLM-5.5 has no endpoint, so it cannot be benchmarked, price-compared, load-tested or put behind a fallback rule. Every week spent planning around it is a week of decisions deferred on the strength of a forecast — and the forecast has already missed one window.
GLM-5.3, by contrast, is live on OrcaRouter under z-ai/glm-5.3 at $1.26 input and $3.96 output against the $1.40 and $4.40 Z.ai lists on its own page. That is passed through at zero markup, which means the number you see is the provider's current rate rather than one we set, and a vendor price change is live on our side the same day rather than at the next billing cycle.
Step 5 Preview is not in our catalogue and we do not route it. It runs on StepFun's own API and Studio, and that is the only place it runs until the October 15 checkpoint lands. What OrcaRouter gives you in the meantime is everything on the other side of that comparison behind one API for more than 200 models: GLM-5.3 at the rate above, GPT-6 Astra, DeepSeek V4 Pro, Claude Opus 5 and the rest, with automatic failover across providers holding a production path steady while a preview's capacity curve is still unknown, and the routing DSL composing several of them into a single call rather than a second integration. Benchmark the preview against a production model on the same key; leave the production path where it is.

What would change the answer
Two dates decide this. October 15 is when StepFun has said the BF16 checkpoint for Step 5 Preview lands, and the licence that ships beside it is the single most consequential unknown in the comparison — a permissive licence would make a 2.8-times cost-per-task advantage something you own rather than rent, and would erase the only axis on which GLM-5.3 is clearly ahead. A restrictive licence leaves GLM-5.3 as the only one of the pair you can build a product on, and turns the one-point gap into trivia about two models you would use differently anyway.
The second is the analyst window running October 8 to November 9 for Z.ai's next flagship, which may or may not be called GLM-5.5 and may or may not resemble the rumour. If it arrives at all, it arrives as a model with a licence and a checkpoint — the two things Step 5 Preview is missing — and the comparison becomes a real one.
Until both resolve, the useful question is not which of the two is better, because one of them cannot be run. It is which of the models that exist you should be building on now, and whether the answer changes when the next checkpoint appears. That question has a testable answer this week, and it does not require waiting for a name to become a model.
Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
