Generated hero card headed 'Step 5 Preview vs Muse Glimmer - the API and the laptop'. Left card 'Step 5 Preview' reads: AA Index 44 against a class median of 26, StepFun, announced Sep 20 2026, about 600B total / 27B active, 1M-token context, $1.00 in / $2.70 out per 1M tokens, runs on the vendor's servers. Right card 'Muse Glimmer' reads: AA Index 17, Meta Superintelligence Labs, released Aug 10 2026, 30B parameters at 4-bit under 20 GB, 131K-token context extendable to 262K, Apache 2.0, runs on your own GPU offline. A footer reads 'Index figures per Artificial Analysis; Muse Glimmer specifications per Meta.'
Guides & Insights

Step 5 Preview vs Muse Glimmer: The Frontier API and the Laptop Model

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

On a leaderboard, Step 5 Preview and Muse Glimmer are not close: 44 against 17 on the Artificial Analysis Intelligence Index — a gap of twenty-seven points on a scale whose current leader sits in the high fifties — that no amount of tuning will close. On the question a team actually has to answer — where do my tokens run — they are direct competitors. Step 5 Preview is StepFun's flagship, announced 20 September 2026, roughly 600 billion total parameters with 27 billion active, a 1M-token context, priced at $1.00 per million input tokens and $2.70 per million output tokens, and reachable only over a network. Muse Glimmer is Meta Superintelligence Labs' 30-billion-parameter open-weight agentic model, released 10 August 2026 under Apache 2.0, shipped quantised to 4-bit so it fits under 20 GB of memory and runs on a single consumer GPU with the network unplugged. The right way to read this pairing is not "which is smarter". It is "which of the two constraints — capability or perimeter — is the one your workload cannot move."

The comparison is also a useful corrective to a habit this market has picked up: treating a composite index score as the whole of a model's value. The twenty-seven-point gap is real and it is not the only number in the file. On long-context retrieval, the same independent board puts Muse Glimmer within half a point of open-weight models from a much larger weight class, and Meta's own agentic benchmarks are respectable for a model that runs on a laptop. Meanwhile Step 5 Preview's most-cited weakness is not capability at all but verbosity, and it is closed until its promised weights arrive on 15 October.

Two models, two entirely different objects

Step 5 Preview is a mixture-of-experts system served from StepFun's infrastructure: about 600B total parameters, 27B active per token, text and image input, a 1M-token context window, and an independent Intelligence Index score of 44 against a comparable-model median of 26. Its output runs at 87 tokens per second against a 78 average on the same board's measurement, and it generated about 160 million output tokens across the index task suite where the median model emits about 82 million — roughly double the text per unit of work, billed at $2.70 per million. The BF16 weights were promised for 15 October 2026 and are not in the repository yet, so today the model exists for you only as an API.

Muse Glimmer is the opposite object. It is a 30B-parameter model trained by logit distillation from Meta's larger Muse Spark teacher, then fine-tuned on long-context agentic data and reinforced for tool use. Meta ships it quantised to 4-bit, which takes memory from over 55 GB at full precision to under 20 GB — inside a 24 GB or 32 GB VRAM envelope — and it was tested by Meta on an Nvidia RTX 5090 and on Apple M4 Max and M5 Max silicon. It takes text and image input across more than a hundred languages, supports a 131,072-token context extendable to 262,144, and is built to take a goal, break it into steps, call tools and retry a failed call rather than stopping. It is Apache 2.0, and it runs offline.

• Independent score — Step 5 Preview: 44 against a median of 26 in its class vs Muse Glimmer: 17 on the same index in its high configuration.

• Scale — Step 5 Preview: about 600B total, 27B active per StepFun vs Muse Glimmer: 30B total, quantised to 4-bit under 20 GB.

• Context window — Step 5 Preview: 1M tokens vs Muse Glimmer: 131,072 extendable to 262,144, per Meta.

• Price — Step 5 Preview: $1.00 in / $2.70 out per million tokens, published vs Muse Glimmer: no per-token charge; you supply the GPU.

• Where it runs — Step 5 Preview: the vendor's servers, over HTTP vs Muse Glimmer: your own hardware, offline.

• Licence — Step 5 Preview: closed preview, weights promised for 15 October 2026 vs Muse Glimmer: Apache 2.0, downloadable now.

• Long-context retrieval — Muse Glimmer scores 80.0 on the index board's long-context reasoning evaluation, within half a point of models several times its price class and close enough to Step 5 Preview's reading that the difference is not decision-relevant for find-the-fact work.

The twenty-seven points, and what they do not measure

A 30B model distilled to run in 20 GB of memory is going to lose to a 600B flagship on a general intelligence index, and Meta's own material does not claim otherwise — one of its framings is blunt that Glimmer is not designed to match the raw reasoning power of full-scale cloud models. So the score of 17 is not a verdict on the engineering; it is the expected result of weighing a different objective. The right question is which of your tasks the gap actually covers.

The long-context reading is where the two come closest and where the intuition breaks down. Muse Glimmer scores 80.0 on the board's long-context reasoning evaluation — a score that would be unremarkable for a frontier model and is remarkable for one running on a consumer GPU. If your workload is retrieval, summarisation or extraction over long documents, a local model at that level is not obviously the wrong tool, and the fact that the request never leaves your machine is a feature you cannot buy from an API at any price.

Then there is the part of the comparison that Meta's numbers do not show. A peer-reviewed study of multi-agent orchestration tested Muse-Glimmer-30B in a controlled setup — same task, same harness, same consultants — and found it recovered none of the candidates produced by larger frontier models, failing all three attempts in each setting. That is a single study of one configuration, and it is the kind of evidence a model page never surfaces. For agentic pipelines where a weaker orchestrator has to recover from a stronger model's failure, it is the most decision-relevant result in this piece, and it is worth reading before any of Meta's vendor-reported benchmarks.

Those benchmarks, for completeness, are Meta's own and unreproduced: 76.0% on SWE-Bench Verified, 51.2% on SWE-Bench Pro, 51.7% on TerminalBench 2.1, 65.9% on OSWorld-Verified, 83.5% on GPQA Diamond, 22% on Humanity's Last Exam and 75.5 on MCP-Atlas. They are measured against models in its own size class, and they describe a capable local agent rather than a frontier reasoner.

Generated two-column scoreboard card titled 'Step 5 Preview vs Muse Glimmer - the scoreboard'. The left column gives Step 5 Preview an independent index of 44, the vendor's servers as the runtime, closed preview weights, $1.00 / $2.70 per 1M tokens, about 160M output tokens against an 82M median, and wins on capability ceiling, 1M context and image input. The right column gives Muse Glimmer an independent index of 17, your own GPU offline as the runtime, Apache 2.0 weights, no per-token charge, a long-context retrieval score of 80.0 within half a point of much larger models, and wins on privacy perimeter, zero marginal cost and portability. A footer reads 'Index and long-context figures per Artificial Analysis; Muse Glimmer specifications per Meta, vendor-reported.'

Where the boundary actually falls

Step 5 Preview's structural advantage is that it does the work you cannot do locally. A 1M-token context, image input, a measured frontier-adjacent score and 87 output tokens per second with no hardware to buy: for drafting, planning, code review across a large repository, or anything where the ceiling of the model is the bottleneck, the API wins and the twenty-seven points are the whole argument. The cost of that is the opposite of Muse Glimmer's: prompts leave your perimeter, the price reapplies on every token, and the model's verbosity makes long-output workloads more expensive than the $2.70 rate suggests.

Screenshot of the Artificial Analysis model page for Muse Glimmer in its high configuration, headed 'Muse Glimmer (High) Intelligence, Performance & Price Analysis', showing Meta as the creator, an open-weights release date of August 2026, an Intelligence Index of 17, and a 131k-token context window with text and image input.

Muse Glimmer's advantage is that it does the work that cannot leave. If the data is regulated, if the latency budget is a few milliseconds, if the machine is offline, or if the marginal cost of the millionth request has to be zero, then no API is a candidate — not this one, not any. The price of that is capability: you are choosing a 17 where a 44 exists, and in agentic orchestration you may be choosing a coordinator that cannot recover from a stronger model's mistakes.

For most teams the answer is not a choice but a split, and the split needs somewhere to live. OrcaRouter routes more than 200 models behind one API key and one bill at 0% markup with provider list price passed through, plus automatic failover and a routing DSL for sending traffic by task, cost or latency. Neither Step 5 Preview nor Muse Glimmer is in that catalogue — StepFun's flagship is not routed by us and Glimmer is a local download — so the value on offer is not access to either model. It is the set you would benchmark them against, callable on one key today, so that the local-versus-remote decision is settled by your own task distribution rather than by whichever vendor's page you read last.

Screenshot of the OrcaRouter models page at www.orcarouter.ai/models, showing 207 models from 16 providers behind one API key and one bill, with filters for input modalities, context length, input price, status, series and supported parameters, and a credits panel.

How to decide

Choose Step 5 Preview when the ceiling is the constraint. If your tasks are open-ended, multimodal, long-context or agentic in the sense of needing a capable planner, the API model is the only one of the two that can do them, and its price is low enough for its class that piloting it is a budget line rather than a project. Budget for verbosity, plan for the 15 October weight date to slip, and keep the workloads that must not leave the building out of it.

Choose Muse Glimmer when the perimeter is the constraint. If your data cannot be sent, if you need to run on a plane or in a factory, or if your volumes make per-token pricing absurd, a 30B Apache-2.0 model that fits in 20 GB is the practical choice, and its long-context retrieval result is good enough that the capability gap will not show up in every workload. Test it in orchestration before you trust it there, because that is where the independent evidence is weakest.

If both constraints bind at once, run both: local for the data that cannot move, API for the tasks that need a bigger model, and let the routing decision be a config change rather than a migration. The comparison that matters is not 44 against 17. It is which of your requests can afford to leave the machine, and that is a question only your own traffic can answer.