
Step 5 Preview vs A.X-K2: A 600B API Model Against a 690B Download
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 113 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 52 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 62 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 399 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Step 5 Preview and A.X-K2 are the two halves of a decision that keeps coming up this autumn: rent a frontier-class model from an API, or take an open-weight flagship home and run it yourself. StepFun's Step 5 Preview is the rental — announced on 20 September 2026, callable the same day at $1.00 per million input tokens and $2.70 per million output tokens, roughly 600 billion total parameters with 27 billion active, a 1M-token context, and an independent Artificial Analysis Intelligence Index score of 44. SK Telecom's A.X-K2 is the download — released 12 August 2026 under Apache 2.0, approximately 688 to 692 billion total parameters with 33 billion active, a 262K-token context, and an index score of 21. One is bigger, one is priced, and the one that is priced is the one that scores much higher. That inversion is the whole comparison, and it is worth understanding before you pick a side on principle.
The usual framing for these two models is national: Korea's sovereign-AI flagship against one of China's most aggressive labs. That framing is not useful for a deployment decision. What is useful is that A.X-K2 is the only one of the pair you can hold, and Step 5 Preview is the only one of the pair you can buy by the token. Everything else follows from that split.
The numbers, held to their sources
The one apples-to-apples measurement is the Intelligence Index, and it is independent on both sides: 44 for Step 5 Preview against a median of 26 among comparable models, and 21 for A.X-K2 against a median of 18 in its open-weights comparison class, both measured by the same third party on the same scale. Read at the time of writing, those are today’s figures — scores on this index move when the evaluator recalibrates, and an earlier reading of the same A.X-K2 page showed a higher number. Everything else in this matchup comes with an attribution you have to carry around with it.
• Independent intelligence — Step 5 Preview: 44 on the Artificial Analysis Intelligence Index vs A.X-K2: 21 on the same index, independently measured.
• Total / active parameters — Step 5 Preview: about 600B total, 27B active per StepFun vs A.X-K2: about 688–692B total, 33B active per SK Telecom materials and the index card.
• Context window — Step 5 Preview: 1M tokens per StepFun vs A.X-K2: 262K tokens.
• Modality — Step 5 Preview accepts text and image; A.X-K2 is text-only.
• Price — Step 5 Preview: $1.00 per million input and $2.70 per million output, published by the vendor vs A.X-K2: no commercial API price published at all.
• Where it runs — Step 5 Preview: StepFun's own API and Studio, plus partner surfaces during promotional windows vs A.X-K2: your hardware, from a download.
• Release — Step 5 Preview announced 20 September 2026 vs A.X-K2 released 12 August 2026.
• Licence — Step 5 Preview: closed preview with BF16 weights promised for 15 October 2026 vs A.X-K2: Apache 2.0.
A.X-K2: the heavier instrument, with the thinner evidence
SK Telecom built A.X-K2 as the successor to A.X K1, and the architecture is aimed squarely at long-context cost: a mixture-of-experts layout with 256 experts and 8 active per token, trained natively in FP8, with a mechanism the company calls Sparse Gated Attention. That is a serious piece of engineering, and it is the reason a model with nearly 700 billion total parameters can be served at all by an organisation that is not a hyperscaler.
What it does not have is independent confirmation. SK Telecom reports an average gain of 32.2 points across fourteen benchmarks over A.X K1, a jump of about 83.9 points on long-context and agent evaluations, 97.1 on AIME 2026, 80.5 on the Korean knowledge benchmark KMMLU-Pro and 91.6 on CLIcK, and says it matches or beats recent Qwen and DeepSeek releases on mathematics and Korean-language work. Every one of those is the vendor's own figure. The independent score — 21 on the Intelligence Index — clears the 18 median of its open-weights comparison class but sits twenty-three points below Step 5 Preview, and A.X-K2's strength set is where that particular index does not weigh much: its nine evaluations are largely English-language, reasoning- and agent-centric, and not one of them is a Korean-language benchmark.
None of that makes A.X-K2 a worse model than its score. It makes its score a lower bound on what it does for a Korean-language or mathematics-heavy workload, on the vendor's word, until an independent lab publishes those evals. The cost side is more concrete: a reference deployment of this weight class needs four B300-class GPUs and roughly 656 GiB in FP8. That is the price of the download, and unlike a token bill it does not scale back down when the experiment ends.
Step 5 Preview: the one you can try over HTTP today
Step 5 Preview's advantage is not that it is bigger or more open — it is neither — but that it is priced, callable and independently scored, and that it has been free to try in three partner surfaces during October. For a team evaluating rather than self-hosting, that combination is worth more than an Apache 2.0 licence, because it converts a hardware decision into an API call. StepFun's journal on the model is thin in the usual places, though: the 600B/27B parameter pair and the 1M-token context are the vendor's own numbers, its promised weights have not shipped, and the independent board's most useful figure about it is arguably not the score of 44 but the verbosity — roughly 160 million output tokens generated across the index's task suite against a median near 82 million, and A.X-K2 is further out still at about 230 million against a 140 million median. On a model that bills $2.70 per million output tokens, that gap is the difference between a cheap evaluation and an expensive production habit.
Why the bigger model scores lower
Parameter counts are not points, and this pairing demonstrates it cleanly. A.X-K2's extra ~300 billion total parameters buy density that shows up in the places its own evaluation suite rewards — Korean-language depth, mathematics, broad knowledge — while the index that both models appear on is mostly English, reasoning-heavy and agent-centric, where Step 5 Preview's builders appear to have tuned deliberately. The 27B-active design also makes the StepFun model markedly cheaper to serve at any given throughput, which is why a hosted price exists at all.
There is a second, less flattering reading that cuts the other way. A.X-K2 has been out since 12 August and Step 5 Preview since 20 September; the newer model has fewer weeks of independent scrutiny, and its 44 may yet move as more evaluations land. Both numbers are provisional in the way that first-quarter scores on any model are provisional. The difference is that A.X-K2's provisional score is the best evidence anyone outside SK Telecom has, while Step 5 Preview's provisional score sits on top of a vendor specification sheet that a third party has at least partly reproduced.
Serving economics, which is where the decision actually lives
If you can call it, you pay $1.00 in and $2.70 out per million tokens and you are done — no procurement, no GPUs, no ops. If you download it, you pay for the hardware, the power, the quantisation work and the evaluation harness, and you get to keep your prompts and your data inside your own perimeter. A.X-K2 at 33B active is not a laptop model; the reference configuration is a small cluster. Step 5 Preview's closed preview offers no perimeter at all until 15 October, when the BF16 weights are promised — and until those files appear in the repository, that date is a commitment rather than an option.

This is the point at which a routing layer earns its place rather than merely advertising itself. Most teams do not want to answer "rent or own" once and forever; they want to know what a model does on their traffic before they buy hardware for it, and then keep the ability to move if the answer changes. OrcaRouter does not route Step 5 Preview, and it does not route A.X-K2 — both of these are outside our catalogue today. What it does is put more than 200 models behind a single API key and one bill at 0% markup with provider list price passed through, so the comparison set you run either new flagship against is the set you can call this afternoon, and automatic failover keeps a production path alive while you are still deciding.


Which one to pick
Choose Step 5 Preview if your workload is general reasoning, agentic coding or anything with long context, if you want to pay per token rather than buy silicon, and if image input matters — it wins on the independent score, on context length, on latency to a first answer and on variety of modalities, and it is the only one of the two you can put behind a procurement-free pilot this month.
Choose A.X-K2 if your language is Korean, if your workload is mathematics-dense, or if the data simply cannot leave your infrastructure — and go in knowing that you are accepting SK Telecom's benchmark table on trust, that the hardware bill is real, and that the independent score you can point at is twenty-three points lower. If that trade looks wrong, the honest answer is not to pick the other flagship; it is to benchmark both on your own traffic, which is the only comparison either vendor's numbers cannot substitute for.
Neither model is the safe default. Step 5 Preview is the cheap, measured, callable option with its weights still on the way; A.X-K2 is the heavier, sovereign, downloadable option with its evidence still owed. Pick the constraint that actually binds your team — the perimeter or the bill — and let the other one go.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
