
Nex-N2.5 Pro vs Claude Opus 5: Nex-AGI Chose the Yardstick, and the Yardstick Won
- openaiNEWOpenAI: GPT-6 Astra2026-09-0455Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0247Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0247Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0157Intelligence82Coding
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2646Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1849Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1541Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1242Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1251Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0547Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0347Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3141Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2454Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2140Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2128Intelligence49Coding
Nex-AGI picked the yardstick, and the yardstick won. When the Shanghai-based open-model alliance introduced Nex-N2.5 Pro on September 8, 2026, the model card's computer-use table put Claude Opus 5 in the comparison column and Nex-N2.5 Pro in the contender slot: 68.3 for Claude Opus 5 versus 56.4 for Nex-N2.5 Pro on OSWorld-2, both figures exactly as Nex-AGI printed them on its own card. The independent leaderboards have since widened rather than narrowed the gap — BenchLM's OSWorld 2.0 snapshot of August 29 ranks Claude Opus 5 first among the fifteen models it lists, at 70.6. Conceding twelve points to the field leader on the benchmark you chose to publish is not the usual launch choreography, which is why the interesting part of Nex-N2.5 Pro vs Claude Opus 5 is not the score. It is what the two sides of that score actually are: a 397-billion-parameter open computer-use model that nobody outside the alliance has been able to run or verify yet, and the closed Anthropic flagship that leads the benchmark and has carried production traffic since late July.
The honest summary up front: this is not a rivalry between two measured quantities, because only one side of it has been measured by anyone other than its maker. Nex-N2.5 Pro is a sparse mixture-of-experts model, roughly 17 billion active parameters under a 397B total, post-trained from Qwen3.5 lineage, aimed squarely at long-horizon computer and browser operation with a visual feedback loop. It is being served today through free hosted endpoints Nex-AGI links from its model card, while the Apache-2.0 weights the family is premised on remain marked "coming soon" with no date. Claude Opus 5 is Anthropic's production flagship for demanding reasoning, coding and agentic work — a closed model with a 1M-token context, an adaptive-thinking dial, a documented price, and weeks of public leaderboard entries and production traffic behind it. Every advantage in this matchup belongs to one of those two facts: one model is runnable and verifiable, and the other is neither.
The benchmark both sides agree on
Computer-use benchmarks reward the same behavior these models are built to sell: an agent that looks at a screen, decides on an action, takes it, and reads the changed screen to check whether the action worked. Nex-N2.5 Pro is architected around exactly that loop — vision is not an input modality bolted onto it, it is the feedback channel the agent verifies against. Claude Opus 5 treats the same loop as one workload among many, but it happens to be very good at it, which is why Nex-AGI used it as the reference point in the first place.
The two numbers are not, strictly, from the same harness. Nex-AGI produced its OSWorld-2 figures through its own NexCUA evaluation harness, which it says it will open-source but has not yet released, and the 56.4 for Nex-N2.5 Pro is unreproduced vendor output. Claude Opus 5's 70.6 on BenchLM is a third-party snapshot of OSWorld 2.0, dated August 29, 2026, and it is consistent with the 68.3 that Nex-AGI itself cites from leaderboard sources. Different harnesses and slightly different benchmark versions mean the exact gap — twelve points on Nex-AGI's own card, closer to fourteen on BenchLM — should be read as directional. But it is not in dispute that Nex-N2.5 Pro, by the best numbers available on either side, sits below the current frontier computer-use agent.

Two answers to "can I run it today"

This is the fork that actually decides the matchup for a working team, and it does not favor the newcomer. Nex-N2.5 Pro's Hugging Face card still banners that the weights are coming soon under an Apache-2.0 license, with no release date attached. The only live builds are the free hosted endpoints Nex-AGI links from the card — callable, but owned by someone else, with no list price, no service terms a team can plan around, and no way to reproduce a single benchmark figure on your own hardware. When the weights do land, the documented serving recipe is a single node of eight H100s on a customized SGLang image — a real but routine footprint for a 397B MoE, and exactly why the 17B-active design is interesting. Until that happens, Nex-N2.5 Pro is a preview you can query, not a model you can build on.
Claude Opus 5 is the opposite on every one of those rows. It has been served through Anthropic's API and on routed platforms since July 24, its price is public and stable, its weights are closed and will stay closed, and any of its numbers can be checked today by running it. The surrounding ecosystem — Claude Code, MCP tooling, and the swarm of production agents that have hammered it for six weeks — is exactly the accumulated record a day-old model cannot claim. For a team whose requirement is "the model has to work next quarter," that record is the entire ballgame.
The spec sheet, such as it is
Put the two envelopes side by side:
• Released — Nex-N2.5 Pro: September 8, 2026 (hosted endpoints live, weights pending). Claude Opus 5: July 24, 2026, generally available.
• Scale — Nex-N2.5 Pro: 397B total / ~17B active MoE. Claude Opus 5: parameter count undisclosed.
• Modality — Nex-N2.5 Pro: text and image in, text out, vision central to its computer-use loop. Claude Opus 5: text and image in, text out, with strong visual analysis.
• Context / output — Nex-N2.5 Pro: 262,144-token context, documented serving length. Claude Opus 5: 1M-token context, 128K-token output ceiling.
• Reasoning control — Nex-N2.5 Pro: a reasoning_effort switch with none / medium (adaptive, default) / high. Claude Opus 5: adaptive thinking by default with an explicit low-to-max effort dial.
• Weights — Nex-N2.5 Pro: Apache-2.0, coming soon, no date. Claude Opus 5: closed.
• Price per 1M tokens — Nex-N2.5 Pro: free hosted tier, no list price published. Claude Opus 5: $5.00 input / $25.00 output, $0.50 cached reads, $6.25 cache writes.
The envelopes are different enough that the spec sheet does real work: Opus 5 brings a million-token window and a 128K output ceiling to long agent sessions, while Nex-N2.5 Pro's 262K context and free tier suit smaller-scope, screen-driven automation. Neither spec makes the other redundant; the models disagree about what an agent spends its context on.
Reading Nex-AGI's own scoreboard honestly
The rest of the card is worth reading with the same filter. Nex-N2.5 Pro's other computer-use rows — OSWorld-G 87.4, OSWorld-Verified 82.2, WebArena-Verified 67.6, WebTest 52.8, Vision2Web 68.2 — and its coding rows — Terminal-Bench 2.1 at 82.7, SWE-Bench Pro at 61.2, BrowseComp at 89.7 — are all vendor measurements through the alliance's own harness, unreproduced in the first 48 hours. The only conclusion a neutral reader can draw from the card is that Nex-AGI believes Nex-N2.5 Pro is a strong mid-tier computer-use model that trails the frontier agent on the row that matters most. That is a coherent, even conservative, positioning for a 397B open model. It is also a claim waiting for a second laboratory.
The money, when one side has no price
There is no honest price comparison to run here, because only one side has a price. Nex-N2.5 Pro's free hosted tier tells you nothing about what the model is worth — free preview endpoints are how vendors collect traffic and feedback, and Nex-AGI has published no paid per-token rate for the family. Claude Opus 5 costs $5.00 per million input and $25.00 per million output at Anthropic's list price, with cache reads at $0.50, and that is the number a budget has to absorb. If you are paying that bill today for computer-use agents and want to know whether a 17B-active open model could do the same job cheaper, the answer is: possibly, but you cannot find out until the weights drop and someone independent runs it. The price comparison is deferred, not settled, and treating a free endpoint as evidence of cheapness would be a category error.

What you can do today is set up the A/B for the day it becomes possible. Claude Opus 5 is on OrcaRouter at Anthropic's list price — the same $5.00 / $25.00, passed through with no markup, so a bill here matches Anthropic's and moves the day Anthropic moves it. One API key puts it behind the same endpoint as 200+ other models, with automatic failover if a provider hiccups and the routing DSL if you want Opus on the hard calls and a cheaper model on the routine ones. Nex-N2.5 Pro is not on OrcaRouter — we checked our model directory before writing, and no nex-agi model is listed there yet. The moment a provider serves it at list, it will appear here the same day, and the honest comparison this article cannot run yet becomes a routing rule instead of a migration.
Who should act, and who should wait
If your team runs production computer-use or agentic-coding workloads today and needs a model with a known price, a known failure profile and an independent record, Claude Opus 5 is the rational pick in this matchup, and nothing in Nex-N2.5 Pro's first 48 hours changes that. If your team is evaluating open computer-use models for a future build, Nex-N2.5 Pro belongs on the watch list — a 397B multimodal MoE with a sane self-hosting footprint and a plausible mid-tier benchmark story is exactly the kind of open model that reshapes a cost curve, provided the weights actually arrive and the card's numbers survive an independent run. The rational position is both at once: keep the proven leader on the critical path, and let the day the weights land decide whether the newcomer earns a trial. That is the one comparison in this matchup that has not happened yet — and it is the one that will actually matter.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
