Generated two-column comparison card headed 'Step 5 Preview vs Intern-S2 Mobius' with the subtitle 'A 600B flagship you rent, or a 35B reasoner you run'. Left card 'Step 5 Preview' reads: AA Index 44 (measured); about 600B total / 27B active; context 1M tokens; $1.00 in / $2.70 out per 1M; closed preview API; weights promised Oct 15 2026. Right card 'Intern-S2 Mobius' reads: no independent score; 35B parameters; context not published; no API price, self-host; Apache 2.0 available now; claimed about 4x speedup, unreproduced. A footer reads 'Step 5 Preview figures per StepFun and Artificial Analysis; Intern-S2 Mobius figures are InternLM claims, unreproduced.'
Engineering & Research

Step 5 Preview vs Intern-S2 Mobius: Do You Need a 600B Flagship, or a Fast 35B Reasoner?

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The honest reason to compare Step 5 Preview with Intern-S2 Mobius is not that they are similar. It is that they are answers to two different questions from the same buyer — the team that has a reasoning workload to run and no appetite for buying a cluster. StepFun's Step 5 Preview, announced on 20 September 2026, is the large answer: roughly 600 billion total parameters with 27 billion active, a 1M-token context, an independently measured Artificial Analysis Intelligence Index score of 44, and a published price of $1.00 per million input tokens and $2.70 per million output tokens. InternLM's Intern-S2 Mobius, released 29 July 2026, is the small answer: a 35-billion-parameter open-weight model built on the Mobius-v0 architecture, continual-pretrained from Qwen3.5-35B, whose central claim is not capability but speed — roughly a 4x end-to-end speedup on reasoning benchmarks against the baseline it was trained from, with no independent benchmark score of any kind. One is a flagship you rent; the other is a small model you run. The comparison is really about whether the workload in front of you justifies the flagship at all.

One piece of housekeeping before the numbers, because it costs people real time: InternLM has shipped several models under the Intern-S2 banner, and they are not the same thing. Intern-S2-397B is the scientific multimodal flagship; Intern-S2 Mobius is the 35B efficient reasoner discussed here; an earlier Intern-S2 release is a separate model again. If you are searching for "Intern-S2" and landing on benchmark tables for a 397B model, you are reading about a different checkpoint.

What each model actually claims

Step 5 Preview's claims are the conventional ones for a 2026 flagship, and one of them is independently checked. StepFun lists about 600B total parameters with 27B active, text and image input, a 1M-token context window, and a $1.00/$2.70 price per million input and output tokens. Artificial Analysis measures it at 44 on its Intelligence Index, above a comparable-model median of 26, with output at 87 tokens per second against a 78 average and about 160 million output tokens generated across the index task suite against a median of 82 million — roughly double the typical verbose footprint, which matters on an output-billed model.

Intern-S2 Mobius's claim is architectural and narrower. Its design separates knowledge from computation: a globally shared Memory holds knowledge vectors, and multiple Reasoners iteratively query and refine hidden states against it, instead of binding storage and computation layer by layer the way a conventional transformer does. InternLM's stated payoff is a roughly 4x end-to-end speedup on reasoning benchmarks versus the Qwen3.5-35B it descends from, with comparable or stronger reasoning scores, plus shorter visible chains of thought. The model is Apache 2.0, text and image input, and it is a download.

• Scale — Step 5 Preview: about 600B total, 27B active per StepFun vs Intern-S2 Mobius: 35B total.

• Independent score — Step 5 Preview: 44 on the Artificial Analysis Intelligence Index vs Intern-S2 Mobius: none published by any third party.

• Speed — Step 5 Preview: 87 output tokens per second against a 78 average, measured by the index board vs Intern-S2 Mobius: "nearly 4x" against its own Qwen3.5-35B baseline, vendor-reported and unreproduced.

• Context window — Step 5 Preview: 1M tokens per StepFun vs Intern-S2 Mobius: no independent context figure published.

• Price — Step 5 Preview: $1.00 in / $2.70 out per million tokens vs Intern-S2 Mobius: no API price; you pay for the hardware.

• Licence and access — Step 5 Preview: closed preview over the vendor's API, BF16 weights promised for 15 October 2026 vs Intern-S2 Mobius: Apache 2.0 weights available now.

• Verbosity — Step 5 Preview: about 160M output tokens across the index suite against an 82M median, per the index board vs Intern-S2 Mobius: shorter visible reasoning traces claimed by InternLM, unmeasured by anyone else.

The 4x claim, and why it is the interesting number here

Read the two vendors' numbers side by side and the mismatch is obvious: StepFun's headline is a capability score, InternLM's is a throughput multiplier. That is not a coincidence, and it is the most useful thing in this comparison. A 4x end-to-end speedup on reasoning tasks, if it survives contact with someone else's harness, is worth more to a high-volume deployment than a few points on an index — it changes the arithmetic of running the workload at all. Intern-S2 Mobius is aiming at teams whose bottleneck is tokens per second per dollar, not ceiling capability.

The caveat is that the multiplier is measured against Qwen3.5-35B, the model Intern-S2 Mobius was continual-pretrained from, which never had the Mobius training. That is a comparison against its own starting point rather than against a rival. There is no independent reproduction of the throughput claim, no third-party index score, and no published context window from anyone other than the vendor. Treat "nearly 4x" as an interesting architectural signal and a hypothesis to test on your own harness, not as a specification.

Step 5 Preview has the opposite evidentiary profile. Its capability number is independently measured; its efficiency story is the weak part. The index board's verbosity reading — about 160 million output tokens where the median model emits about 82 million — is the honest counterweight to a $2.70 output price, and it means a workload that leans on long structured generations will spend materially more than the sticker suggests. What it does have that Intern-S2 Mobius does not is a published API price at all, which is what makes it comparable in the first place.

The Hugging Face repository for the promised vendor build tells the other half of the story: this is what an open-weight release looks like when it has been announced but not yet shipped.

Generated two-column scoreboard card titled 'Step 5 Preview vs Intern-S2 Mobius - the scoreboard'. The left column gives Step 5 Preview a speed of 87 output tokens per second against a 78 average, about 160M output tokens against an 82M median, text and image input, closed preview weights, and best-for open-ended, long-context and multimodal work. The right column gives Intern-S2 Mobius a claimed speed of about 4x faster than its own Qwen3.5-35B baseline, shorter reasoning traces claimed, text and image input, Apache 2.0 weights, and best-for narrow high-volume reasoning inside your own perimeter. A footer reads 'Speed and verbosity on the left are third-party measurements; every figure on the right is the vendor's own and unreproduced.'

Where the footprint question actually bites

Intern-S2 Mobius's real advantage is not its score, which nobody has measured, but what it costs to be wrong about. Thirty-five billion parameters is a size a team can serve on hardware it already owns, evaluate without a purchase order, and abandon without writing anything off. That is a structural advantage the flagship cannot match no matter how many points it scores, and it is the reason this comparison is worth having rather than dismissing as a mismatch.

Screenshot of the Hugging Face repository page for stepfun-ai/Step-5-Preview-BF16, the vendor placeholder for the promised open-weight build of Step 5 Preview, showing the repository shell with no model card, no configuration, no licence terms and no tensor files.

The flagship's advantage is the mirror image. Step 5 Preview's 1M-token context, image input and measured general-reasoning capability cover work that a 35B model is unlikely to cover well, and it costs nothing to try for the length of a free window — the model has been free in three separate developer surfaces during October. Neither of those facts is available for a self-hosted model: there is no free week for running your own GPUs, and there is no price list that converts the decision into a line item.

If you want the comparison to survive past the promotional window, the thing to build is a harness rather than a preference. OrcaRouter routes more than 200 models behind one API key and one bill at 0% markup with provider list price passed through, with automatic failover and a routing DSL for steering traffic by cost, latency or capability. Neither Step 5 Preview nor Intern-S2 Mobius is in that catalogue — StepFun's flaghip is not routed by us and Intern-S2 Mobius is a self-host download — so the value here is not access to either. It is that the models you will benchmark them against are one key away, and a decision made by measurement moves with the evidence instead of with a launch blog post.

Screenshot of the OrcaRouter models page at www.orcarouter.ai/models, showing 207 models from 16 providers behind one API key and one bill, with filters for input modalities, context length, input price, status, series and supported parameters, and a credits panel.

How to decide

If your workload is agentic, multimodal, long-context or open-ended — the kind where you cannot enumerate the tasks in advance — the flagship is the answer, and Step 5 Preview is a reasonable instance of it: measured, priced, cheap to pilot, and reversible by the token. Just budget for the verbosity rather than the headline rate.

If your workload is narrow, repetitive, high-volume reasoning on data that has to stay inside your perimeter, the small model is the answer, and Intern-S2 Mobius is a genuine candidate precisely because it is small: it runs on hardware you have, it is Apache 2.0, and the speed claim, if it holds, is the kind of thing that decides a unit economics question. Go in knowing you are testing the vendor's claim rather than inheriting someone else's verdict.

The mistake is treating the choice as permanent. Buy the small model's evaluation first, because it is the cheap experiment; rent the flagship for the tasks the small model fails, because that is the cheap repair. Teams that keep both options alive on the same key and the same harness end up making this decision with their own data, and that is the only version of it that holds up when either vendor ships a successor.