A generated hero title card for 'Nex-N2.5 Pro vs MiniMax M3' with the subtitle 'The claimed 34-point gap and the model you can actually download', two score tiles reading 'OSWorld-2 56.4' under 'Nex-N2.5 Pro (claimed)' and 'OSWorld-2 22.3' under 'MiniMax M3', with the OrcaRouter logo composited bottom-right.
Guides & Insights

Nex-N2.5 Pro vs MiniMax M3: the Claimed 34-Point Gap and the Model You Can Actually Download

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The most striking number in this matchup is 56.4 to 22.3 — and the least verifiable. Those are the OSWorld-2 scores Nex-AGI's Nex-N2.5 Pro and MiniMax M3 carry on the new model's own benchmark card: Pro at 56.4, MiniMax's M3 at 22.3, a 34-point gap on the long-horizon computer-use benchmark that every serious GUI agent now gets measured against. Nex-N2.5 Pro, introduced September 8, 2026, is the 397B-total / ~17B-active multimodal agent whose weights are still "coming soon." MiniMax M3, released June 1, 2026, is the 428B-total / ~23B-active open-weight multimodal flagship that also accepts video input, serves a 1M context on MiniMax Sparse Attention, and has been downloadable since early June. One of those two models you can pull from the hub and run tonight. The other is a hosted preview with a promise. And the 34-point gap that makes the matchup worth reading is printed by the vendor of the model you cannot run, about a benchmark where even the incumbent's own number is vendor-reported.

So this page is really about how you weigh a dramatic, unverifiable lead against a modest, partially verifiable one — and about what the cheap model's existence does to the expensive model's pitch. MiniMax M3 is the counterweight that every weights-pending specialist has to argue against: it already does most of the job, it costs almost nothing, and its files are on your disk. For Nex-N2.5 Pro to matter, the 34-point claim has to survive the three checks that follow.

First check: are the numbers even measuring the same thing?

OSWorld-2 is the long-horizon, real-world extension of the original OSWorld benchmark — an agent operates actual desktop software across multi-step tasks and is scored on whether it gets real work done, not whether it clicks the right button. It is dramatically harder than its predecessor, and it is the benchmark where Nex-AGI claims its Pro tier separates itself from the multimodal pack. But look at where the two rows come from. Nex-N2.5 Pro's 56.4 is Nex-AGI's own measurement through its internal NexCUA harness, which the card says will be open-sourced "soon." MiniMax M3's 22.3 is a number Nex-AGI printed for a rival model — MiniMax has not published an OSWorld-2 result of its own, and M3's public computer-use figure is the older, easier OSWorld-Verified, which Nex's card also lists at 75.2 for M3 against 82.2 for Pro. Both rows in that 56.4-to-22.3 headline come from the same harness, run by the same vendor, on the same day, about one model it owns and one model it does not. That does not make the gap false — harnesses exist precisely to compare models under identical conditions. It does mean the number is a single vendor's controlled comparison, not a result any outside lab has confirmed, and the gap is exactly as trustworthy as the harness that produced it.

Second check: what does the model you can actually run deliver?

MiniMax M3 does not need its OSWorld-2 row to be useful, because its case rests on things you can verify by downloading it. It is 428B total / ~23B active, built on MiniMax Sparse Attention, with a 1M-token context and a 512K output budget; it takes text, image and video in and returns text; its weights have been on Hugging Face since the first week of June; and its list price on MiniMax's API is $0.30 per million input tokens and $1.20 per million output with $0.06 cached input — about a fifth of the output price of the cheapest closed frontier models, with video billed separately. Its independent standing is real but modest: a 44 on the Artificial Analysis Intelligence Index places it mid-tier among flagships, below the 56-61 range the closed leaders occupy. Its headline agentic numbers — 75.2 OSWorld-Verified as compiled on Nex's card, ~70 on MiniMax's own reporting, 83.5 BrowseComp — are vendor-reported and unreproduced, the same asterisk that follows Nex. The difference is that M3's claims point at a model you can run and re-test yourself, which is why "vendor-reported" stings less on the side with downloadable weights.

A comparison scoreboard for Nex-N2.5 Pro and MiniMax M3: Nex-N2.5 Pro rows 'Total / active: 397B / ~17B', 'Weights: coming soon', 'Context: 262K', 'Price: no rate card (free preview)', 'OSWorld-Verified: 82.2 (Nex-reported)', 'OSWorld-2: 56.4 (Nex-reported)'; MiniMax M3 rows 'Total / active: 428B / ~23B', 'Weights: open (Hugging Face)', 'Context: 1M / 512K out', 'Price: $0.30 / $1.20, cache $0.06', 'OSWorld-Verified: 75.2 (Nex-compiled)', 'OSWorld-2: 22.3 (Nex-compiled)'; footer 'Computer-use rows come from Nex-AGI's own card and harness; MiniMax M3 AA Index 44 per Artificial Analysis.'

Third check: what would you actually do with each one?

Buying today — MiniMax M3: real weights, a rate card, a working API. Nex-N2.5 Pro: a free hosted preview and a roadmap.

Computer use — MiniMax M3: a generalist that handles screenshots and video and posts mid-pack GUI scores. Nex-N2.5 Pro: a vision-native agent built for the visual-feedback loop, claiming the class-leading OSWorld-2 row.

Price — MiniMax M3: $0.30 / $1.20, cache $0.06 — cheap enough to throw at long agent loops. Nex-N2.5 Pro: no published price yet; the efficiency argument (17B active on a single 8×H100 node) is a projection, not a bill.

Context — MiniMax M3: 1M tokens, with requests above 512K billed at 2×. Nex-N2.5 Pro: 262K family config.

Inputs — MiniMax M3: text, image and video, natively. Nex-N2.5 Pro: image and text, oriented to screen-level perception.

Verifiability — MiniMax M3: self-hostable, so any claim is testable in-house. Nex-N2.5 Pro: no weights, so every claim is Nex-AGI's word.

Independent index — MiniMax M3: 44 (AA). Nex-N2.5 Pro: none yet.

Why the cheap incumbent is the harder opponent

There is a reason Nex-AGI chose to print MiniMax M3 in its multimodal table rather than ignore it: M3 is the price anchor that makes a computer-use specialist's premium hard to justify. If a generalist that costs $0.30/$1.20 and accepts video can do even 70-75% of what a dedicated GUI agent claims, the specialist has to be dramatically better at the actual task to earn its place — and Nex-N2.5 Pro's card says it is, on OSWorld-2, by 34 points. But a 34-point lead printed by the challenger about its unshipped model is a hypothesis, while M3's 44 Index, its open weights and its price are facts you can build on. The economically rational posture is not to dismiss the hypothesis — it is to insist that it survive an independent run before it costs anyone money.

Running the comparison on your own key

The routing layer is where a price-anchored comparison like this gets tested without committing to either side. MiniMax M3 is on OrcaRouter at MiniMax's list price — $0.30 / $1.20, passed through at 0% markup, so the day MiniMax changes a rate the change is live on the same key you already use for the rest of the catalog. Nex-N2.5 Pro is not offered through us, so you reach it through Nex-AGI's hosted preview until the promised weights appear. The experiment that actually answers this matchup: point your high-volume, long-context agent traffic at MiniMax M3 through the one endpoint, run a smaller evaluation branch against Nex-N2.5 Pro's preview, and compare them on your own task mix while automatic failover keeps both honest — a 34-point vendor claim is worth exactly as much as your own reproduction of it, and M3 is cheap enough that reproducing the test costs almost nothing. When Pro's weights land, the eight-H100 recipe is the real deliverable to check: whether the 17B-active specialist holds its OSWorld-2 lead when someone other than Nex-AGI runs it.

A screenshot of the Hugging Face 'Files and versions' tab for nex-agi/Nex-N2.5-Pro (captured September 9, 2026), showing the model header with Text Generation and Transformers tags, 'License: apache-2.0', and a file tree listing only figures, .gitattributes and README.md - no model weight shards, confirming the Pro checkpoint is not yet published.A screenshot of the OrcaRouter model page for minimax/minimax-m3 (captured September 9, 2026), showing the MiniMax M3 card under vendor MiniMax with model id minimax/minimax-m3, max output 512K, text + image + video input and text output, vision/tools/reasoning chips, and a 'Public benchmarks by MiniMax' note.

The verdict the evidence supports

On everything that is currently actionable, MiniMax M3 wins this matchup by default: it is downloadable, it is priced at a fifth of the frontier norm, it takes video, it has an independent Index, and its computer-use claims — while vendor-reported — point at a model you can re-test in your own stack this week. Nex-N2.5 Pro is the more interesting model and the worse purchase: a vision-native computer-use specialist whose own card claims a class-leading OSWorld-2 score and whose weights do not exist yet. Treat the 56.4-to-22.3 gap as the best available hypothesis about where computer use is heading, not as a result you can act on — and treat MiniMax M3 as the reason the specialist, when it finally ships, will have to prove that gap on someone else's harness to justify any price above $0.30/$1.20.