A generated hero title card for 'Nex-N2.5 Pro vs Kimi K3' with the subtitle 'The matchup where the newcomer's own benchmark card referees', a large rounded document labeled 'VENDOR BENCHMARK CARD' with two model columns and the line 'Pro trails K3 on most shared rows' above it, with the OrcaRouter logo composited bottom-right.
Guides & Insights

Nex-N2.5 Pro vs Kimi K3: the Matchup Where the Newcomer's Own Benchmark Card Referees

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Check the numbers Nex-N2.5 Pro and Kimi K3 have in common and you will notice who supplied them. Nex-AGI's Nex-N2.5 Pro — the 397-billion-parameter, ~17-billion-active multimodal agent introduced September 8, 2026 with weights still marked "coming soon" — publishes a benchmark table that lists Moonshot AI's Kimi K3 directly in the same columns. Kimi K3 is the 2.8-trillion-parameter open-weight flagship Moonshot shipped on July 16, 2026, with a million-token context, full weights under a Modified MIT license eleven days later, and a 57 on the Artificial Analysis Intelligence Index that makes it the top open-weight model on the board. And on row after row of that table, Nex-N2.5 Pro trails Kimi K3: Terminal-Bench 2.1 at 82.7 against K3's 88.3, BrowseComp 89.7 against 91.2, SWE-Bench Pro 61.2 against 63.3, AutomationBench 44.2 against 46.7. When a vendor prints its own model's numbers next to the current open-weights leader and its own model loses most of the rows, the most interesting thing in the matchup is that the table is probably telling the truth.

That is the unusual shape of this comparison. Usually a newcomer's card is the place to look for spin. Here, the card is the closest thing either model has to an honest referee, because Nex-AGI had no reason to publish a table that demotes its own Pro tier — and the rows where Pro does beat K3 are exactly the rows where a small, purpose-built computer-use model should beat a much larger generalist. So the real question this page answers is narrower and more useful than "which is better": does Nex-N2.5 Pro's claimed niche — near-K3 agentic performance at roughly a sixth of the active parameters — survive contact with the fact that Kimi K3 already exists, already has its weights out, and is already the model to beat?

The opponent, in full

Kimi K3 is not a niche model, which is precisely why it is the right yardstick. It is Moonshot's flagship: 2.8T total parameters with 104B active under a Stable LatentMoE design, a 1M-token context with output budget to match, native text, image and video understanding, always-on reasoning, and open weights that anyone with enough iron can actually run. Its independent standing is strong — the 57 Intelligence Index (Artificial Analysis) tops every other open-weight model, and it posted 1,679 Elo on the Frontend Code Arena ahead of Claude Fable 5 and GPT-5.6 Sol — while its price, $3.00 per million input and $15.00 per million output with $0.30 cached input, sits at the premium end of the open-weights tier. Kimi K3 is what "frontier and open" looks like when the vendor does not compromise on scale: it costs more per token than closed rivals like GPT-5.6 Sol on output, it is expensive to serve, and it is very hard to argue with on the leaderboards.

The challenger's arithmetic

Nex-N2.5 Pro is not trying to out-scale Kimi K3. It is trying to out-serve it. Nex-AGI post-trained the Pro tier from a Qwen3.5-lineage base into a vision-native agent for operating computers and browsers, and its pitch is efficiency: 397B total / ~17B active, a 262K context, and a single-node 8×H100 deployment recipe, against K3's 2.8T / 104B active and the kind of serving fleet a model that size demands. The implied argument is that a 17B-active specialist, trained specifically for the visual-feedback agent loop, can deliver most of a 104B-active generalist's agentic performance at a fraction of the serving cost. On Nex-AGI's own card the claim is plausible but partial: Pro beats or matches K3 on OSWorld-G (87.4 vs 79.6) and OmniDoc (92.2 vs 91.1), and loses the rest of the shared rows by single digits — a 397B model losing to a 2.8T model by 3 to 6 points on coding and browsing is, if true, exactly the efficiency story Nex wants to tell.

If true. Every one of those numbers is Nex-AGI's own, produced at least in part through its internal NexCUA harness, and Nex-N2.5 Pro cannot be independently checked today because its weights are not downloadable — the Hugging Face repository holds a README, a figures folder and a "weights are coming soon" banner, nothing more. The same caveat applies to the K3-side figures the card prints, most of which Nex compiled from Moonshot's published reports rather than measured itself. What Nex-AGI's table establishes is not who wins; it is that Nex-AGI's own engineers, choosing the benchmarks and the harness, could not make their Pro tier beat Kimi K3 on most shared ground. That is a more useful datum than any single score, because it sets an upper bound on how much the marketing can be exaggerating.

A comparison scoreboard for Nex-N2.5 Pro and Kimi K3: Nex-N2.5 Pro rows 'Total / active: 397B / ~17B', 'Weights: coming soon', 'Context: 262K', 'Price: no rate card (free preview)', 'Terminal-Bench 2.1: 82.7 (Nex-reported)', 'OSWorld-2: 56.4 (Nex-reported)'; Kimi K3 rows 'Total / active: 2.8T / 104B', 'Weights: open, Modified MIT', 'Context: 1M', 'Price: $3 / $15, cache $0.30', 'Terminal-Bench 2.1: 88.3 (Moonshot)', 'OSWorld-2: 58.3 (Nex-compiled)'; footer 'All figures vendor-reported; Kimi K3 AA Index 57 per Artificial Analysis.'

The rows that matter, side by side

Scale — Nex-N2.5 Pro: 397B total / ~17B active, Qwen3.5-lineage multimodal. Kimi K3: 2.8T total / 104B active, Stable LatentMoE, multimodal.

Weights — Nex-N2.5 Pro: promised, not shipped ("coming soon" on the card). Kimi K3: published July 27 under Modified MIT — downloadable today.

Context — Nex-N2.5 Pro: 262K. Kimi K3: 1M tokens, flat-priced.

Price — Nex-N2.5 Pro: no rate card yet; hosted preview free. Kimi K3: $3.00 / $15.00 per million, cached input $0.30.

Independent standing — Nex-N2.5 Pro: none yet. Kimi K3: AA Intelligence Index 57, top open-weight model on the board.

Coding (vendor cards) — Nex-N2.5 Pro: Terminal-Bench 2.1 82.7, SWE-Bench Pro 61.2. Kimi K3: 88.3 and 63.3 on the same rows.

GUI / computer use — Nex-N2.5 Pro: OSWorld-G 87.4 (beats K3 on its own card), OSWorld-2 56.4. Kimi K3: OSWorld-G 79.6, OSWorld-2 58.3 (Nex-compiled).

Self-host footprint — Nex-N2.5 Pro: single-node 8×H100 recipe once weights land. Kimi K3: 2.8T — a serving fleet, not a single node.

What "open" means differently in each column

The word "open" is doing very different work on the two sides, and it decides this matchup more than any benchmark. Kimi K3 is open in the sense that matters operationally: the weights are on the hub under a permissive-ish Modified MIT license, the reasoning traces are exposed in the streaming API, and a company with data-governance constraints can run it on hardware it controls and audit what the model was thinking. It is open in a way that costs you — a 2.8T model needs serious infrastructure — but the option exists today. Nex-N2.5 Pro is open in the sense of intent: Nex-AGI says the family weights will be released as open source, and the smaller Nex-N2.5 Mini sibling did ship Apache-2.0 weights on launch day. Pro's are not out, its license is announced on the card but not yet binding on anything downloadable, and until a checkpoint exists, "open" is a roadmap item rather than a property of the model. For a team choosing between the two, that is not a footnote — it is the difference between a model you can own this month and a model you are promised you can own.

Evaluating without waiting for a download button

You do not have to freeze this decision until Nex-N2.5 Pro's weights drop, and the routing layer is where you run the experiment cheaply. Kimi K3 is on OrcaRouter at Moonshot's list price — $3.00 / $15.00, 0% markup, passed through and updated the day Moonshot changes it — so the open-weights leader is one key away next to the other 200+ models in the catalog. Nex-N2.5 Pro is not offered through us, so trialing it today means Nex-AGI's hosted preview and later your own eight-H100 stack. The pattern that fits: put the long-context and audit-heavy work on Kimi K3 through the single endpoint you already use, point a test branch at Nex-N2.5 Pro's preview, and let automatic failover route around whichever one degrades — which is exactly how you compare a shipped frontier model against a weights-pending challenger without betting your production path on either. When Pro's weights do land, the verification is simple: run the shared rows an independent lab can reproduce and see whether the 17B-active efficiency story survives outside Nex-AGI's own harness.

A screenshot of the Hugging Face 'Files and versions' tab for nex-agi/Nex-N2.5-Pro (captured September 9, 2026), showing the model header with Text Generation and Transformers tags, 'License: apache-2.0', and a file tree listing only figures, .gitattributes and README.md - no model weight shards, confirming the Pro checkpoint is not yet published.A screenshot of the OrcaRouter model page for kimi/kimi-k3 (captured September 9, 2026), showing the Kimi K3 card under vendor MoonshotAI with model id kimi/kimi-k3, a 1M-token context, text + image input and text output, vision/tools/reasoning chips, and a 'Public benchmarks by MoonshotAI' note dated 2026-07-15.

Where the matchup lands

Against the largest open-weights model ever shipped, Nex-N2.5 Pro does not need to win the generalist war to be worth watching — it needs to win the serving-cost argument, and its own card suggests that argument is real but unproven. Kimi K3 is the safe and genuinely strong choice today: open weights, top of the open-weights index, a million-token context, and a price you can budget against, at the cost of infrastructure that only gets harder as the model grows. Nex-N2.5 Pro is the efficiency bet: near-K3 agentic rows at a sixth of the active parameters on a single node, per a vendor's table about a model that has not shipped. Take Kimi K3 when you want the frontier model you can actually own and audit right now. Watch Nex-N2.5 Pro for the day its weights land and someone independent runs the shared rows — because if a 17B-active agent really does stay within a few points of a 104B-active flagship on computer use, that is the result that changes the serving math for every open agent after it.