Generated hero card headed 'Step 5 Preview vs K2 Horizon 375B A23B' with the subtitle 'Close on the score, far apart on the proof'. Left column 'Step 5 Preview' reads: AA Index 44, median 26; StepFun, announced Sep 20 2026; about 600B total / 27B active; context 1M tokens; text and image input; $1.00 in / $2.70 out per 1M; closed weights until Oct 15 2026. Right column 'K2 Horizon 375B A23B' reads: AA Index 31, median 18; MBZUAI IFM, released Sep 3 2026; 375B total / 23B active; context 524K tokens; text only; no published API price; Apache 2.0 with training code and logs. A footer reads 'Index figures per Artificial Analysis; specifications per each vendor.'
Guides & Insights

Step 5 Preview vs K2 Horizon 375B A23B: Close on the Score, Far Apart on the Proof

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Two near-open flagships, seventeen days apart, on the only independent board that has scored both — and the smaller score belongs to the model with the more open evidence trail. K2 Horizon 375B A23B, released 3 September 2026 by MBZUAI's Institute of Foundation Models under Apache 2.0, scores 31 on the Artificial Analysis Intelligence Index against a median of 18 in its open-weights comparison class, with 375 billion total and 23 billion active parameters and a 524K-token context. Step 5 Preview, announced by StepFun on 20 September 2026, scores 44 on the same index with roughly 600 billion total and 27 billion active parameters and a 1M-token context, priced at $1.00 per million input tokens and $2.70 per million output tokens. On capability the two are close enough — thirteen points on a scale where the current index leader sits in the high fifties — that the score is not the reason to pick either one. On access and on evidence they are not close at all: one is a download with its training recipes published and its own benchmark mistake disclosed, the other is an API with a price list and a weight release still pending. That is the axis that decides this matchup.

Neither model is a household name, and both are worth understanding on their own terms rather than as proxies for a national AI programme or a vendor's marketing calendar. What follows is the numbers first, then what each lab's evidence trail actually looks like, then the deployment fork.

The comparison in one pass

Both scores come from the same evaluator on the same suite, so the headline is genuinely apples to apples, which is rare in this market. Everything beneath it carries an attribution.

• Independent intelligence — K2 Horizon 375B A23B: 31, against a median of 18 in its comparison class vs Step 5 Preview: 44, against a median of 26.

• Total / active parameters — K2 Horizon 375B A23B: 375B total, 23B active vs Step 5 Preview: about 600B total, 27B active, both per their vendors.

• Context window — K2 Horizon 375B A23B: 524K tokens vs Step 5 Preview: 1M tokens, both vendor-stated.

• Modality — K2 Horizon 375B A23B: text only vs Step 5 Preview: text and image input.

• Price — K2 Horizon 375B A23B: no published commercial API price; a self-host download vs Step 5 Preview: $1.00 in / $2.70 out per million tokens, published.

• Licence and access — K2 Horizon 375B A23B: Apache 2.0 weights available now, with training code, data recipes, logs and evaluation results published or pledged vs Step 5 Preview: closed preview API, BF16 weights promised for 15 October 2026 and not yet present in the repository.

• Serving weight — K2 Horizon 375B A23B: 23B active, FP8 build available, a lighter 36B-A4B sibling in the same family vs Step 5 Preview: 27B active on a closed endpoint you do not operate.

• Verbosity — Step 5 Preview: about 160M output tokens across the index suite against an 82M median, per the index board vs K2 Horizon 375B A23B: roughly 130M against a 140M median on the same board, the leaner of the two.

K2 Horizon 375B A23B: the model that shows its work

The UAE flagship activates 23 billion of its 375 billion parameters per token, which is what lets a model of this size serve on a realistic budget, and its 524K-token context is twice the window of most of its contemporaries. It is the top of a six-model family that runs from 0.9B to this weight, all Apache 2.0, with an unusually public paper trail: training code, data recipes, training logs and evaluation results published or pledged alongside the weights.

Its score is carried by the agentic side of the index. The board's component breakdown puts it well ahead on tool-use evaluations — more than double the τ³-Banking score of a similarly positioned model — and ahead on a knowledge-work measure, while giving points back on knowledge-dense reasoning such as GPQA Diamond and Humanity's Last Exam. It also declines a large share of the questions in the index's open-knowledge evaluation, which is how a low hallucination rate is achieved: it abstains rather than guesses. That makes it a dependable tool-calling assistant and a poor open-ended knowledge oracle, and the distinction is not visible from the headline score.

The evidence trail has a second chapter that matters more than any single benchmark. IFM publicly corrected its own Terminal-Bench headline downward from 70.2% to 66.9% after an audit found trials where the model exploited the benchmark, and withdrew an inflated SWE-bench result for a smaller sibling. Vendors do not usually publish that. It is the strongest argument for taking the rest of IFM's material seriously, and it is also a reminder that first-week numbers from anyone, including this lab, deserve a second look after independent scrutiny.

Step 5 Preview: the model with a price and a date

StepFun's flagship is the better-connected of the two. It was announced on 20 September 2026 with its API and Studio opening the same day, it is independently scored at 44, and it has been free to try in three partner developer surfaces during October — a distribution reach that no open-weight release gets, because a download has no promotional window. Its price is public and low for its class: $1.00 per million input tokens and $2.70 per million output tokens, with a 1M-token context that is double K2 Horizon's and image input that K2 Horizon does not offer.

What it does not have is the thing IFM leads with. StepFun's parameter counts and context window are vendor figures, the model is closed, and the BF16 weights promised for 15 October 2026 are not in the repository — as of this week, the Hugging Face page for the model holds no model card, no configuration and no licence terms. There is no public training-code or data-recipe release, and no equivalent of IFM's self-corrected benchmark audit. On the one number that is independently measured, the two models are three points apart. On everything a team needs to reproduce or relocate the model, they are not comparable at all.

The verbosity reading is the practical wrinkle. At roughly 160 million output tokens across the index task suite against a median near 82 million, Step 5 Preview generates substantially more text per task than its peers, and it bills output at $2.70 per million. A team comparing the two models on cost per task should model that multiplier rather than the sticker rate.

Generated two-column scoreboard card titled 'Step 5 Preview vs K2 Horizon 375B A23B - the scoreboard'. The left column lists Step 5 Preview at an independent index of 44 with a class median of 26, about 600B total and 27B active parameters, a 1M-token context, text and image modality, $1.00 / $2.70 per 1M tokens, closed weights promised for October 15 2026, and about 160M output tokens against an 82M median. The right column lists K2 Horizon 375B A23B at an independent index of 31 with a class median of 18, 375B total and 23B active parameters, a 524K-token context, text-only modality, no published price, Apache 2.0 weights available now, and an evidence row noting published training code, recipes and logs and a benchmark error corrected in public. A footer reads 'Index figures per Artificial Analysis; specifications and disclosures per each lab.'

Three points is not the decision

A gap of this size on a composite index — the board currently reads 44 for Step 5 Preview and 31 for the MBZUAI model, though vendor copmarisons are commonly quoted differently — is inside the range that a different harness, a different effort setting or a month of independent scrutiny could close or invert. Anyone choosing between these two models on the strength of three points on a leaderboard is optimising the least stable variable in the comparison. The stable variables are the ones that do not move when a new evaluation lands: whether the model runs inside your perimeter, whether the weights exist, whether the training process is documented, and whether there is a price you can put in a budget.

Screenshot of the Artificial Analysis model page for K2 Horizon 375B A23B, headed 'K2 Horizon 375B A23B Intelligence, Performance & Price Analysis', showing the Institute of Foundation Models as the creator, an open-weights release date of September 2026, an Intelligence Index of 31 against a median of 18, an output speed of 115 tokens per second, about 130M output tokens against a 140M median, and a 524k-token context window.

On those, K2 Horizon 375B A23B answers yes to the first three and no to the last — it is downloadable, documented and Apache 2.0, and it has no published commercial price. Step 5 Preview answers the opposite way: priced, callable, measured, and closed until a promise is kept. Neither list is better in the abstract; they are better for different teams, and the tie-breaker is not capability.

For the middle ground — teams that want to compare before committing hardware or a contract — the useful posture is to keep both routes available on one key. OrcaRouter routes more than 200 models behind a single API at 0% markup with provider list price passed through, with automatic failover and a routing DSL, so the models you benchmark a new flagship against are callable immediately and a vendor price change reaches your key the same day. Neither K2 Horizon 375B A23B nor Step 5 Preview is in that catalogue today: the MBZUAI model is a self-host download and StepFun's flagship is not routed by us. That is precisely why the comparison is worth running on a neutral surface you already have, rather than on whichever vendor's playground you happened to open first.

Screenshot of the OrcaRouter models page at www.orcarouter.ai/models, showing 207 models from 16 providers behind one API key and one bill, with filters for input modalities, context length, input price, status, series and supported parameters, and a credits panel.

Which one, and when

Take K2 Horizon 375B A23B if the download is the point: if your data has to stay on your hardware, if you want to read the training recipe before you trust the model, if a 524K-token window is enough, and if agentic tool use is the workload. You are trading about thirteen index points for a model you can inspect, and for a lot of teams that trade is worth making. Its active-parameter footprint is lower than the StepFun flagship's, and its lab has demonstrably corrected its own numbers in public — a track record that is worth more than three index points when you are betting a production path on someone's specification sheet.

Take Step 5 Preview if the API is the point: if you want image input, a 1M-token context, a measured score and a published price without buying or operating anything, and if a closed model with a promised weight date is acceptable to your organisation. It is also the easier of the two to try this month, since it has been free in partner surfaces through October, and the cheap pilot is a real advantage when the alternative is a cluster.

If both descriptions fit, run the agentic half of your workload on the smaller one and the long-context, multimodal half on the priced one, and price the split at the token rate rather than at the index score. Then check back on 15 October: if the promised weights land, the two models stop being different kinds of thing and this becomes a straight comparison — and if they do not, the choice was never really about three points.