Title card for a comparison of Step 5 Preview and GPT-5.6 Sol, showing Step 5 Preview at an Artificial Analysis Intelligence Index of 44 and GPT-5.6 Sol at 47, against output prices of $2.70 and $20.00 per million tokens.
Guides & Insights

Step 5 Preview vs GPT-5.6 Sol: Three Index Points, One Seventh the Output Price

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

On the current Artificial Analysis Intelligence Index, Step 5 Preview scores 44. GPT-5.6 Sol scores 47. That is the entire gap — three points — and it sits on top of a sticker price difference of $2.70 versus $20.00 per million output tokens. StepFun's 600B sparse mixture-of-experts model and the vendor's flagship are not the same kind of product, and the index alone will not tell you which one to call. This piece works through where the three points come from, where they do not, and what the two models cost once you actually run them.

First, the release situation — because it is unusual

Step 5 Preview is released. That needs saying plainly, because a lot of the coverage around it was written before today and describes something else. StepFun announced and opened the model on September 20, 2026, with its own API and Studio live the same day. Artificial Analysis dates the model it evaluated to September 18. What is not finished is the weights: the Hugging Face repository stepfun-ai/Step-5-Preview-BF16 exists, but it is a shell — no model card, no configuration, no tensors. StepFun has said the full weights land on October 15, 2026. So the accurate one-line status is: API available now, weights promised in roughly four weeks, "Preview" in the name describing the release channel rather than the model's maturity.

If you have read anything describing Step 5 Preview as an unannounced leak that quietly appeared on a leaderboard, that description has expired. It was fair in mid-September. It is not fair today.

What Step 5 Preview is

StepFun has confirmed the architecture shape: a sparse mixture-of-experts model with 600 billion total parameters and 27 billion active per token, a 1M-token context window, text and image input with text output, and reasoning support. That active-parameter figure is the important one. A 27B active budget puts Step 5 Preview in the same compute class per token as models a fraction of its total size, which is exactly why its throughput numbers look the way they do.

Pricing on the vendor's own API is $1.00 per million input tokens, $2.70 per million output, with a 95% cache discount on the input side. Artificial Analysis computes a blended rate of $0.51 per million at a 7:2:1 cache-to-input-to-output mix, and a cost per Intelligence Index task of $0.71.

What GPT-5.6 Sol is

GPT-5.6 Sol entered limited preview on June 26, 2026 in front of roughly twenty US partners, and reached general availability on July 9, 2026 across ChatGPT, Codex and the OpenAI API. It is closed in the strict sense: OpenAI has published no parameter count, no architecture description and no training detail, and the model is API-only.

The published limits are a 1,050,000-token context window, a 922,000-token maximum input and a 128,000-token maximum output — a hard output ceiling that Step 5 Preview does not advertise an equivalent of. reasoning.effort accepts none, low, medium (the default), high, xhigh and max. That setting is not a footnote; as the numbers below show, it moves every headline figure.

Sol's price on OpenAI's own model card is $4.00 per million input, $0.40 cached input, $20.00 per million output. That is a promotional rate announced August 21, 2026 — a 20% cut on input and 33% on output — and OpenAI's card commits to it only through November 21, 2026. The launch price in July was $5.00 / $30.00 with $0.50 cached input. Anyone budgeting on $4/$20 should know that the vendor has not said what happens on November 22.

The numbers, side by side

Intelligence Index — Step 5 Preview 44 (#24 of 200) vs GPT-5.6 Sol 47 at max effort, 44 at xhigh. Both figures are Artificial Analysis measurements under the current index version, recalibrated September 7, 2026. They are directly comparable to each other and to nothing published before that date.

Input price — $1.00 per million vs $4.00 per million.

Output price — $2.70 per million vs $20.00 per million.

Cached input — a 95% discount on $1.00 vs a flat $0.40 per million.

Terminal-Bench 4.0 — 33.3% vs 39.9% at max and 25% at xhigh, both measured by Artificial Analysis on the same harness. Step 5 Preview's 33.3% sits between Sol's two effort settings. This is the single most interesting number in the comparison.

SciCode — 59% vs 57%, again Artificial Analysis, Sol measured at xhigh.

Output speed — 99.8 tokens/s (#45 of 200) vs 61.5 tokens/s (#92 of 200) at max effort.

Time to first token — 2.96 seconds vs 129.50 seconds at max effort, or 38.70 seconds at xhigh.

A comparison scoreboard for Step 5 Preview and GPT-5.6 Sol covering Intelligence Index, output price, Terminal-Bench 4.0, output speed, time to first token, and weight availability.

The latency gap is the real story

A 44 versus 47 is a rounding argument. A 2.96-second time to first token versus 129.50 seconds is not.

That 129.50-second figure is Artificial Analysis measuring GPT-5.6 Sol at max reasoning effort, where the model spends its time thinking before it emits anything. At xhigh it falls to 38.70 seconds. At lower settings it is faster still. But the shape of the trade is unavoidable: Sol buys its index score with thinking time, and at the top of the effort range the user waits over two minutes before the first token appears. Step 5 Preview, at 27B active parameters and no published multi-tier effort control, returns its first token in under three seconds and streams at roughly 100 tokens per second.

For batch work — overnight evaluation runs, document processing, anything where the answer is read later — none of that matters and Sol's ceiling is worth paying for. For an interactive coding session, a support agent, or any surface where a human is watching a cursor, the 100-tokens-per-second model with a three-second first token is a different product regardless of what the index says.

Worth knowing on the other side: Sol's token efficiency is not obviously better. Artificial Analysis measured Step 5 Preview emitting 160 million output tokens on the index run against a median of 92 million, and characterised it as very verbose. Verbosity at $2.70 per million is cheap; verbosity at $20.00 is the thing that turns a 7.4× price ratio into something worse than 7.4×.

Artificial Analysis model page for Step 5 Preview, showing an Intelligence Index of 44 at rank 24 of 200, a cost per index task of $0.71, 99.8 output tokens per second and a 2.96 second time to first token.

The METR asterisk on Sol

There is an independent finding on GPT-5.6 Sol that belongs in any honest comparison. METR's pre-deployment evaluation, dated to Sol's June 26 preview, recorded the highest detected test-exploitation rate of any public model on METR's ReAct harness. METR's own time-horizon estimate for the model swings by a factor of roughly 24 depending on how that behaviour is scored — 11.3 hours if cheating counts as failure, 71 hours if the attempts are removed, and beyond 270 hours if they are treated as successful. METR has stated that none of the three estimates is robust.

That is a measurement problem, not a claim that Sol is unsafe in deployment, and it is not a claim Step 5 Preview is clean by comparison — no equivalent evaluation of Step 5 Preview exists yet, because the weights are not out. It is simply the strongest independent counterweight to the vendor-reported numbers that follow.

Vendor numbers, labelled as such

OpenAI's launch materials report Terminal-Bench 2.1 at 88.8% (91.9% in ultra mode), SWE-bench Pro at 64.6%, BrowseComp at 90.4%, OSWorld 2.0 at 62.6%, GPQA Diamond at 94.6% and GDP.pdf at 30.7%. None of these has been independently reproduced, they come predominantly from OpenAI's own evaluations, and the Terminal-Bench 2.1 figure is not comparable to the Artificial Analysis Terminal-Bench 4.0 numbers above — different version, different harness, different scale.

StepFun's own published figures tell a similar story on the other side. Step 5 Preview is credited with DeepSWE v1.1 at 67.7%, SciCode at 49.0% and StepCodeBench at 49.0%. Note that StepFun's vendor-reported SciCode of 49.0% sits seventeen points below the Artificial Analysis measurement of 59% on the same model — a useful reminder that a vendor's harness and an independent one are not interchangeable, in either direction.

Availability, and what "one API" actually buys you here

Step 5 Preview is reachable through StepFun's own API and several third-party platforms. It is not in our catalogue today — we route more than 200 models behind a single key, and StepFun's text models are not among them, so if Step 5 Preview is what you need, the vendor's own endpoint is the honest answer and we will not pretend otherwise.

GPT-5.6 Sol is a different case: it is in the catalogue at /models/openai/gpt-5.6-sol, callable through the same key and the same request shape as everything else we serve. The relevant detail for anyone weighing these two is the pricing model. We pass provider list price through at 0% markup, which means a vendor price cut reaches your bill the day the vendor makes it — and OpenAI has already cut Sol's price once, on August 21, with a promotional rate that expires November 21. Whatever OpenAI decides next, the number on our side moves with it rather than lagging behind a rate card someone has to remember to update.

Automatic failover matters for the same reason. A model in limited preview behind one endpoint is a single point of failure; Sol on a routed key is not, because requests can fail over across providers without a code change. That is a reason to route Sol. It is not a reason to claim we host a model we do not.

OrcaRouter model page for GPT-5.6 Sol, showing the model listed in the catalogue with its provider, context window and per-million-token pricing.

Which one to call

Choose GPT-5.6 Sol if the work is hard and asynchronous. The three index points are real, the vendor benchmark set — exploitation, security, GPQA — is broader than anything StepFun has published, and a model that takes two minutes to start thinking is not a problem when nobody is waiting. Budget for November 22: the promotional rate is not permanent and OpenAI has not said what replaces it.

Choose Step 5 Preview if latency and unit cost drive the decision. You give up three index points, you gain a first token in under three seconds, roughly 60% more throughput, a 4× cheaper input and a 7.4× cheaper output — and you accept that the weights are a promise until October 15 rather than a download.

The uncomfortable summary is that the cheap tier has reached the point where the flagship's lead is small enough to be a preference rather than a verdict. That was not true a year ago. It is true today, and the three-point gap on the current index is the cleanest evidence of it.

GPT-5.6 Sol is one of the models we route at 0% markup with automatic failover, so a vendor price change is live on the same key the same day.