Generated title card titled 'Step 5 Preview vs Claude Opus 5 — the price gap' with two columns: Step 5 Preview at AA Index 44, price $1.00 in / $2.70 out, index run cost $922.84 and output speed 99.8 tokens/sec; Claude Opus 5 at AA Index 51, price $5.00 in / $25.00 out, index run cost $7,274.74 and output speed 55.3 tokens/sec. A footer reads '7.9x the cost for 7 index points. Both per Artificial Analysis Intelligence Index v4.3.2 at maximum reasoning effort.'
Guides & Insights

Step 5 Preview vs Claude Opus 5: Seven Index Points for Seven Times the Bill

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

StepFun's launch material for Step 5 Preview leads with a single comparative claim: a single task on it costs one-eighth of the same task on Claude Opus 5. Both models now sit on the same independent board, so that claim can be checked rather than argued about. Artificial Analysis ran each through Intelligence Index v4.3.2 — ten evaluations — and published what each run cost, with Claude Opus 5 scored at its own maximum reasoning-effort setting. Step 5 Preview: 44 on the index, $922.84 to complete the run. Claude Opus 5: 51 on the index, $7,274.74. That is 7.9 times the money for 7 points of score, and the ratio StepFun quoted turns out to be the accurate one. What the ratio hides is the interesting part, because the gap is not mainly a price gap. It is a verbosity gap wearing a price gap's clothes, and separating the two changes which model you should put in front of which workload.

Two run costs, one board, and what is actually inside them

A total run cost is the cleanest single comparison available here because both models answered the same questions under the same harness, and neither vendor produced the figure. It is also the figure most likely to be misread, so it is worth opening up.

• Index score — Step 5 Preview 44 vs Claude Opus 5 51, Claude Opus 5 scored at its maximum reasoning-effort setting

• Total cost to complete the index — Step 5 Preview $922.84 vs Claude Opus 5 $7,274.74

• Blended price at a 7:2:1 cache-hit / input / output mix — Step 5 Preview $0.51 per 1M tokens vs Claude Opus 5 $3.85

• Output tokens generated during the index run — Step 5 Preview 160M vs Claude Opus 5 140M

• List price — Step 5 Preview $1.00 in / $2.70 out vs Claude Opus 5 $5.00 in / $25.00 out

Look at rows three and five together and the arithmetic stops being a straight price comparison. Step 5 Preview's blended rate is 7.5 times cheaper, and the list rate is 5 times cheaper on input and 9.3 times cheaper on output — so the blend is doing work the input price does not. That is because the blend is weighted toward output, and output is where the two models diverge most. Claude Opus 5 charges 9.3 times Step 5 Preview's output rate, and both models write a great deal: 140M output tokens for Claude Opus 5 against a 92M median across the models the index scores, and 160M for Step 5 Preview against the same 92M median.

So the honest reading is narrower than "the cheap model is a bargain." Step 5 Preview is cheaper on every axis, and it is not a frugal model — it writes 74% more output than the median model it is measured against, which is exactly the behaviour the price is supposed to punish. Claude Opus 5 writes 52% more than the median and charges 9.3 times as much for each of those tokens. A caller who constrains output length gets a different ratio than one who does not.

The question is not which is better. It is where the crossover sits

Suppose the index is a reasonable proxy for the work you actually send — long-horizon agentic tasks, code, multi-step tool use. Then the crossover is simple to state: Claude Opus 5 has to be at least 7.9 times more useful per task before it is worth its bill, and on the index it is 7 points better on a 100-point scale. Those two facts do not have to point the same way, because a 7-point index gap is not 16% better at everything. It is the aggregate of ten evaluations, and the composition matters more than the total.

That is the argument for caring about the shape of the two scores rather than the headline. Step 5 Preview's Terminal-Bench 4.0 reading is 33.3% — a real agentic result, ahead of DeepSeek V4.1 Flash at 26.8% — but nothing in the published record yet shows it matching Claude Opus 5 on the long-horizon evaluations that Anthropic's model is strongest at. StepFun's own materials concede this in the phrasing: on ALE-CLI, FrontierFinance and DRACO, the company places Step 5 Preview second, behind either GPT-6 Astra or Claude Opus 5, and ahead of the other open models in the comparison. Second place at a fifth of the price is a genuinely good result. It is also not first place, and the tasks where a model finishes second are usually the tasks that needed first.

The other asymmetry is what happens on a bad call. A wrong answer from a cheap model that took four seconds and 2,000 tokens costs you the retry; the same wrong answer from Claude Opus 5 at 25 dollars per million output tokens costs you the retry plus the deliberation. If your pipeline retries on failure — most agent loops do — the effective cost per successful task is the number that matters, and it is not the same ratio as the list price, because the two models will not fail at the same rate. Nobody has published that number for Step 5 Preview yet.

Screenshot of the Artificial Analysis model page for Step 5 Preview, showing StepFun as the creator and a September 2026 release date, an Intelligence Index of 44 at rank 24 of 200, output speed 99.8 tokens per second, cost per Intelligence Index task $0.71, verbosity 160M output tokens, pricing of $1.00 per million input tokens and $2.70 per million output tokens with a 95% cache discount, and technical specifications listing reasoning, text and image input, and a 1M-token context window.

The spec contrast, dimension by dimension

• Index score — Step 5 Preview 44 vs Claude Opus 5 51, both on Intelligence Index v4.3.2, Claude Opus 5 at its maximum reasoning-effort setting

• Price per 1M tokens — Step 5 Preview $1.00 in / $2.70 out vs Claude Opus 5 $5.00 in / $25.00 out

• Cache discount — Step 5 Preview 95% vs Claude Opus 5 90%

• Context — both 1M tokens; StepFun has not published an output ceiling and Anthropic's is 128K

• Input — both accept text and images; both output text only

• Output speed — Step 5 Preview 99.8 tokens/sec vs Claude Opus 5 55.3 tokens/sec

• Control surface — Step 5 Preview exposes extended thinking; Claude Opus 5 replaces temperature and top_p with an adaptive reasoning-effort dial

• Weights — Step 5 Preview closed today, BF16 checkpoint scheduled for October 15; Claude Opus 5 proprietary throughout

• Independent evidence — both scored by Artificial Analysis; StepFun's agentic benchmark rows are vendor-run and unreproduced

Two rows deserve a note. The speed gap is real and larger than the table suggests in practice: 99.8 tokens per second against 55.3 is a model that finishes roughly twice as fast on the same output, and Step 5 Preview's time to first token of 2.96 seconds against Claude Opus 5's 56.84 seconds on the same board is a different order of thing — though that second figure is measuring the max-effort reasoning configuration, where the model deliberates before it emits anything. It is not the latency a production call sees. Against the standard endpoint, our own catalogue reports a p50 time to first token of 6.73 seconds for Claude Opus 5, which is the number to plan against if you are not deliberately asking for maximum deliberation.

The cache discount row is the one that quietly decides repeated workloads, and it favours Step 5 Preview, 95% to 90%. An agent loop that re-sends the same system prompt and repository context on every call lives almost entirely in that discount, so a five-point difference compounds across a long session.

Generated two-column comparison scoreboard titled 'Step 5 Preview vs Claude Opus 5 — the scoreboard'. The left column gives Step 5 Preview the rows AA Index 44, price $1.00 / $2.70, cache discount 95%, output speed 99.8 tokens/sec, context 1M tokens, and weights due October 15. The right column gives Claude Opus 5 the rows AA Index 51, price $5.00 / $25.00, cache discount 90%, output speed 55.3 tokens/sec, context 1M tokens, and weights proprietary. A footer reads 'Both per Artificial Analysis Intelligence Index v4.3.2 at maximum reasoning effort. StepFun agentic rows are vendor-run.'

Where Claude Opus 5 is still the only answer

Nothing above argues against Anthropic's model. It costs what it costs because it does the thing the index is trying to measure, and it does it near the top of the board — 51 on Intelligence Index v4.3.2, seven points clear of Step 5 Preview and within reach of the leaders of the current generation. The workloads where a 7-point aggregate gap understates the difference are the ones with no tolerance for a second-place answer: a multi-hour refactor where the failure mode is a subtle behaviour change, a financial model where the reasoning chain has to be auditable, a migration where a wrong decision is discovered a week later. In those, the retry is not cheap and the deliberation is the product.

The workloads where the gap is irrelevant are the ones built out of many small decisions: classification and routing steps, extraction, summarisation, test scaffolding, the thousand calls an agent makes between the two calls that actually matter. On those, a model at 44 that answers in half the time at a fifth of the input price is not a compromise. It is the correct default, with the expensive model reserved for the step that has to be right.

Running the split without running two integrations

The practical version of this comparison is a tiered pipeline, and the reason teams do not build one is procurement rather than engineering — two vendors, two contracts, two SDKs, two keys to rotate. Claude Opus 5 is in our catalogue at $5.00 and $25.00, and the cheaper text models you would put in front of it are in the same catalogue, all behind one key for more than 200 models with provider list price passed through at 0% markup. Step 5 Preview is not on that list — it runs on StepFun's own API, and until the weights land that is the only place it runs. Composing the tiering across the models we do route is a routing-DSL expression rather than a code change — send the routine calls to the small model, escalate to the expensive one on a confidence or complexity condition, and keep both legs behind the same endpoint.

Automatic failover matters for a specific reason here rather than as a generic feature. Step 5 Preview opened its API on announcement day with no published weights and no disclosed serving fleet; it is a preview endpoint days old, and preview endpoints are capacity-constrained in a way that mature ones are not. Claude Opus 5 is served by many independent providers, which is the redundancy a preview cannot offer. Keeping a proven model as the fallback leg — on our side that means a production model from the catalogue, not the preview — is how you get a real quality signal on your own workload without making your production path depend on a fleet that is still being sized.

Screenshot of the OrcaRouter model page for Claude Opus 5 at www.orcarouter.ai/models/anthropic/claude-opus-5, showing the model id anthropic/claude-opus-5, the max output and context figures of 128K and 1M tokens, input pricing of $5.00 and output pricing of $25.00 per million tokens, a p50 time to first token of 6.73 seconds, 59.7M tokens of traffic over seven days, and the endpoints and SDK snippets behind the OrcaRouter API.

The decision rule

If your workload is a small number of very hard tasks, Claude Opus 5 is not the expensive option — it is the cheap one, because the second-best answer costs more than the difference. If your workload is a large number of ordinary tasks with a small number of hard ones inside it, Step 5 Preview at $1.00 and $2.70 with a 95% cache discount is the better default and the 7-point gap is mostly a rounding error on work that did not need it.

What would change that calculation is independent evidence on failure rates and on the long-horizon agentic rows. StepFun's own numbers put the model second on the evaluations it built its launch around, and no third party has reproduced them. Watch the October 15 checkpoint for two things: whether the license permits the fine-tuning that would let you close a 7-point gap on your own data, and whether the verbosity holds once the preview suffix comes off. The first would change the ratio; the second already moved it once.