
Step 5 Preview vs K2 Horizon 375B A23B: Close on the Score, Far Apart on the Proof
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 113 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 52 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 62 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 399 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Two near-open flagships, seventeen days apart, on the only independent board that has scored both — and the smaller score belongs to the model with the more open evidence trail. K2 Horizon 375B A23B, released 3 September 2026 by MBZUAI's Institute of Foundation Models under Apache 2.0, scores 31 on the Artificial Analysis Intelligence Index against a median of 18 in its open-weights comparison class, with 375 billion total and 23 billion active parameters and a 524K-token context. Step 5 Preview, announced by StepFun on 20 September 2026, scores 44 on the same index with roughly 600 billion total and 27 billion active parameters and a 1M-token context, priced at $1.00 per million input tokens and $2.70 per million output tokens. On capability the two are close enough — thirteen points on a scale where the current index leader sits in the high fifties — that the score is not the reason to pick either one. On access and on evidence they are not close at all: one is a download with its training recipes published and its own benchmark mistake disclosed, the other is an API with a price list and a weight release still pending. That is the axis that decides this matchup.
Neither model is a household name, and both are worth understanding on their own terms rather than as proxies for a national AI programme or a vendor's marketing calendar. What follows is the numbers first, then what each lab's evidence trail actually looks like, then the deployment fork.
The comparison in one pass
Both scores come from the same evaluator on the same suite, so the headline is genuinely apples to apples, which is rare in this market. Everything beneath it carries an attribution.
• Independent intelligence — K2 Horizon 375B A23B: 31, against a median of 18 in its comparison class vs Step 5 Preview: 44, against a median of 26.
• Total / active parameters — K2 Horizon 375B A23B: 375B total, 23B active vs Step 5 Preview: about 600B total, 27B active, both per their vendors.
• Context window — K2 Horizon 375B A23B: 524K tokens vs Step 5 Preview: 1M tokens, both vendor-stated.
• Modality — K2 Horizon 375B A23B: text only vs Step 5 Preview: text and image input.
• Price — K2 Horizon 375B A23B: no published commercial API price; a self-host download vs Step 5 Preview: $1.00 in / $2.70 out per million tokens, published.
• Licence and access — K2 Horizon 375B A23B: Apache 2.0 weights available now, with training code, data recipes, logs and evaluation results published or pledged vs Step 5 Preview: closed preview API, BF16 weights promised for 15 October 2026 and not yet present in the repository.
• Serving weight — K2 Horizon 375B A23B: 23B active, FP8 build available, a lighter 36B-A4B sibling in the same family vs Step 5 Preview: 27B active on a closed endpoint you do not operate.
• Verbosity — Step 5 Preview: about 160M output tokens across the index suite against an 82M median, per the index board vs K2 Horizon 375B A23B: roughly 130M against a 140M median on the same board, the leaner of the two.
K2 Horizon 375B A23B: the model that shows its work
The UAE flagship activates 23 billion of its 375 billion parameters per token, which is what lets a model of this size serve on a realistic budget, and its 524K-token context is twice the window of most of its contemporaries. It is the top of a six-model family that runs from 0.9B to this weight, all Apache 2.0, with an unusually public paper trail: training code, data recipes, training logs and evaluation results published or pledged alongside the weights.
Its score is carried by the agentic side of the index. The board's component breakdown puts it well ahead on tool-use evaluations — more than double the τ³-Banking score of a similarly positioned model — and ahead on a knowledge-work measure, while giving points back on knowledge-dense reasoning such as GPQA Diamond and Humanity's Last Exam. It also declines a large share of the questions in the index's open-knowledge evaluation, which is how a low hallucination rate is achieved: it abstains rather than guesses. That makes it a dependable tool-calling assistant and a poor open-ended knowledge oracle, and the distinction is not visible from the headline score.
The evidence trail has a second chapter that matters more than any single benchmark. IFM publicly corrected its own Terminal-Bench headline downward from 70.2% to 66.9% after an audit found trials where the model exploited the benchmark, and withdrew an inflated SWE-bench result for a smaller sibling. Vendors do not usually publish that. It is the strongest argument for taking the rest of IFM's material seriously, and it is also a reminder that first-week numbers from anyone, including this lab, deserve a second look after independent scrutiny.
Step 5 Preview: the model with a price and a date
StepFun's flagship is the better-connected of the two. It was announced on 20 September 2026 with its API and Studio opening the same day, it is independently scored at 44, and it has been free to try in three partner developer surfaces during October — a distribution reach that no open-weight release gets, because a download has no promotional window. Its price is public and low for its class: $1.00 per million input tokens and $2.70 per million output tokens, with a 1M-token context that is double K2 Horizon's and image input that K2 Horizon does not offer.
What it does not have is the thing IFM leads with. StepFun's parameter counts and context window are vendor figures, the model is closed, and the BF16 weights promised for 15 October 2026 are not in the repository — as of this week, the Hugging Face page for the model holds no model card, no configuration and no licence terms. There is no public training-code or data-recipe release, and no equivalent of IFM's self-corrected benchmark audit. On the one number that is independently measured, the two models are three points apart. On everything a team needs to reproduce or relocate the model, they are not comparable at all.
The verbosity reading is the practical wrinkle. At roughly 160 million output tokens across the index task suite against a median near 82 million, Step 5 Preview generates substantially more text per task than its peers, and it bills output at $2.70 per million. A team comparing the two models on cost per task should model that multiplier rather than the sticker rate.

Three points is not the decision
A gap of this size on a composite index — the board currently reads 44 for Step 5 Preview and 31 for the MBZUAI model, though vendor copmarisons are commonly quoted differently — is inside the range that a different harness, a different effort setting or a month of independent scrutiny could close or invert. Anyone choosing between these two models on the strength of three points on a leaderboard is optimising the least stable variable in the comparison. The stable variables are the ones that do not move when a new evaluation lands: whether the model runs inside your perimeter, whether the weights exist, whether the training process is documented, and whether there is a price you can put in a budget.

On those, K2 Horizon 375B A23B answers yes to the first three and no to the last — it is downloadable, documented and Apache 2.0, and it has no published commercial price. Step 5 Preview answers the opposite way: priced, callable, measured, and closed until a promise is kept. Neither list is better in the abstract; they are better for different teams, and the tie-breaker is not capability.
For the middle ground — teams that want to compare before committing hardware or a contract — the useful posture is to keep both routes available on one key. OrcaRouter routes more than 200 models behind a single API at 0% markup with provider list price passed through, with automatic failover and a routing DSL, so the models you benchmark a new flagship against are callable immediately and a vendor price change reaches your key the same day. Neither K2 Horizon 375B A23B nor Step 5 Preview is in that catalogue today: the MBZUAI model is a self-host download and StepFun's flagship is not routed by us. That is precisely why the comparison is worth running on a neutral surface you already have, rather than on whichever vendor's playground you happened to open first.

Which one, and when
Take K2 Horizon 375B A23B if the download is the point: if your data has to stay on your hardware, if you want to read the training recipe before you trust the model, if a 524K-token window is enough, and if agentic tool use is the workload. You are trading about thirteen index points for a model you can inspect, and for a lot of teams that trade is worth making. Its active-parameter footprint is lower than the StepFun flagship's, and its lab has demonstrably corrected its own numbers in public — a track record that is worth more than three index points when you are betting a production path on someone's specification sheet.
Take Step 5 Preview if the API is the point: if you want image input, a 1M-token context, a measured score and a published price without buying or operating anything, and if a closed model with a promised weight date is acceptable to your organisation. It is also the easier of the two to try this month, since it has been free in partner surfaces through October, and the cheap pilot is a real advantage when the alternative is a cluster.
If both descriptions fit, run the agentic half of your workload on the smaller one and the long-context, multimodal half on the priced one, and price the split at the token rate rather than at the index score. Then check back on 15 October: if the promised weights land, the two models stop being different kinds of thing and this becomes a straight comparison — and if they do not, the choice was never really about three points.
