A hero title card for the comparison article 'NVIDIA-Nemotron-Labs-Teacher-General-Reasoning vs Qwen3.8-Max' reading 'The week open weights got serious', with an open-box graduation-cap icon on the left, a globe-and-chip icon on the right, a MODEL COMPARISON badge, and the OrcaRouter logo composited bottom-right.
Guides & Insights

NVIDIA-Nemotron-Labs-Teacher-General-Reasoning vs Qwen3.8-Max: The Week Open Weights Got Serious

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Open weights just had a week. On or around August 12, 2026, Alibaba released the first open-weight Max-class Qwe​n ever — Qwen3.8 Max, a 2.4-trillion-parameter flagship, published as Qwen3.8-2.4T-A95B on Hugging Face and ModelScope. Two days later, on August 14, NVIDIA published NVIDIA-Nemotron-Labs-Teacher-General-Reasoning, a 550B-parameter, 55B-active reasoning teacher from the Nemotron 3 Ultra training pipeline, on Hugging Face as well. Both are enormous, freshly-opened checkpoints from frontier labs, and the resemblance stops there. Qwen3.8 Max is open so you can run a frontier model yourself; NVIDIA-Nemotron-Labs-Teacher-General-Reasoning is open so you can train with it. Comparing them is really asking which kind of team you are — but the fact that both landed inside 72 hours is a signal about where the industry is heading.

What each release actually contains

Qwen3.8 Max is the deployable one. It is a sparse mixture-of-experts model with 2.4 trillion total parameters, about 95 billion active per token across 512 experts, built on the Qwen3.5-family hybrid architecture. In its open-weights form it is text-only with a 262K-token native context (extendable past 1M), and — notably — thinking is required: the model emits reasoning content between <think> tags and cannot be turned off. The hosted API version Alibaba launched on August 3 adds image and video input, a 1M default context, and optional thinking. The open-weights license is a custom Qwen3.8 Max License, permissive but with revenue-based conditions for very large commercial deployments. If you have the multi-node hardware, you can run a frontier-class model on your own infrastructure.

NVIDIA-Nemotron-Labs-Teacher-General-Reasoning is the training-time one. It is one of the more than ten domain-specialized teachers from the Multi-Teacher On-Policy Distillation (MOPD) recipe behind Nemotron 3 Ultra, derived from the post-trained Ultra student through an extra reasoning-specialized SFT and reinforcement-learning stage. Its role in that pipeline is to generate long reasoning traces on the hardest math, logic, and abstract-reasoning problems and to act as the judge that grades free-form answers. It ships under the commercial-friendly OpenMDW-1.1 license with the disclosed post-training data, and it carries zero benchmark numbers — NVIDIA points to an accuracy plot and the Ultra technical report instead. Its customers are distillation runs, not production traffic.

Qwen3.8 Max, the first open Max-class flagship

Alibaba's release matters because Max-class Qwe​n models have historically been API-only. Qwen3.8 Max breaks that: the weights are downloadable, and the model is genuinely frontier — Artificial Analysis gives it an Intelligence Index of 58 and an Agentic Index of 58, tying it with the top closed generalists in agentic settings. Its vendor-reported highlights are strong: 93.0 on PaperBench (ahead of GPT-5.6 Sol and Fable 5), 92.6 on GPQA Diamond, 86.1 on OSWorld-Verified. Its weak spots are equally real — 67.7 on SWE-bench Pro and 43.6 on Humanity's Last Exam trail the leaders, and independent testing found it expensive per finished task, averaging 64 turns per task on agentic runs. On Alibaba's API it prices at $2.00 per million input tokens, $6.00 per million output, and $0.25 per million cached input, a 20% cut from its preview pricing.

The teacher: training infrastructure with a model card

NVIDIA-Nemotron-Labs-Teacher-General-Reasoning is not competing on any of those numbers because it has not published any. What it offers instead is the same LatentMoE Mamba-2 + Transformer hybrid architecture as the Ultra student, up to 1M tokens of context, and a reasoning dial — enable_thinking true/false, a medium_effort mode, and a hard reasoning_budget ceiling — that a distillation pipeline needs in order to spend deep reasoning on hard samples and skip it on easy ones. Minimum hardware is 4×B200 (or 8×H100), serving through vLLM, SGLang, or TensorRT-LLM, and the checkpoint is a 1.12-terabyte download. It is not deployed by any inference provider, and NVIDIA has not announced NIM hosting. This is a weights-only release aimed at labs that already run large GPU fleets.

Specs, side by side

• Purpose — Teacher: training-time reasoning specialist and judge. Qwen3.8 Max: deployable frontier flagship.

• Scale — Teacher: 550B total / 55B active, LatentMoE with MTP. Qwen3.8 Max: 2.4T total / ~95B active, 512 experts, Qwen3.5 hybrid.

• Context — Teacher: up to 1M tokens, 256K default. Qwen3.8 Max: 262K native open-weights context, extendable past 1M; API version 1M default.

• Weights — Teacher: OpenMDW-1.1, released Aug 14, 2026. Qwen3.8 Max: Qwen3.8 Max License, released ~Aug 12, 2026.

• Independent evidence — Teacher: none yet. Qwen3.8 Max: AA Intelligence Index 58, Agentic Index 58, Coding Index 71.8.

• Thinking — Teacher: enable_thinking toggle, medium_effort, reasoning_budget ceiling. Qwen3.8 Max: required-on in open weights, cannot be disabled.

• Price — Teacher: no list price, self-hosted capex. Qwen3.8 Max: $2.00 / $6.00 per million tokens, $0.25 cached input.

A two-column scoreboard for NVIDIA-Nemotron-Labs-Teacher-General-Reasoning vs Qwen3.8-Max. Left column (Nemotron Teacher): training-time reasoning teacher, 55B active (550B total), up to 1M context, OpenMDW-1.1 weights, no independent score yet, self-hosted capex. Right column (Qwen3.8-Max): deployable frontier, 95B active (2.4T total), 262K native context extendable past 1M, Qwen3.8-Max License weights, AA Intelligence Index 58, $2.00/$6.00 per million tokens. Footer notes Qwen figures per Alibaba (vendor-reported) + Artificial Analysis and teacher specs unaudited. OrcaRouter logo composited bottom-right.A screenshot of the Hugging Face model page for nvidia/NVIDIA-Nemotron-Labs-Teacher-General-Reasoning showing the model card: 550B total / 55B active parameters, LatentMoE hybrid Mamba-2 + MoE + Attention with MTP, up to 1M token context, 10 languages, the OpenMDW-1.1 license tag, and the files and versions listing.

The evidence asymmetry is the whole story

Qwen3.8 Max has been tested; the teacher has not. That is partly age — Qwen3.8 Max launched August 3 and has had two weeks of leaderboard runs, while the teacher shipped yesterday — and partly intention, because the teacher's card was never meant to compete on scores. But the asymmetry has a real consequence for the comparison: every concrete number on this page with a source attached belongs to Qwen3.8 Max, and every number attached to the teacher is an architecture spec. The teacher's reasoning quality is unverified by anyone outside NVIDIA, and its family-level figures (the Ultra student's vendor-reported SWE-Bench Verified 71.9, MMLU-Pro 86.8, LiveCodeBench v6 89.0) describe a different model. Anyone building a decision on "which scores better" has to wait for independent runs of the teacher — which is exactly the gap the next few weeks should close.

Two thinking dials, set differently

The most interesting technical contrast between these two releases is what happens when you ask them to reason. Qwen3.8 Max in its open-weights form has no choice: thinking is always on, and the reasoning trace is part of every answer. That is great for answer quality and transparency and expensive for bulk generation. The teacher treats reasoning as a dial: enable_thinking flips it, medium_effort throttles it, and a reasoning_budget ceiling caps it. For a distillation pipeline the difference is decisive — generating millions of reasoning traces with an always-on-thinking model multiplies your token bill, while the teacher's switch lets you pay for deep reasoning only where a hard sample demands it. These are not competing design philosophies; they are two tools optimized for opposite ends of the same problem, and each is the right tool for its end.

A screenshot of the OrcaRouter model page for Qwen3.8-Max (Qwen) showing the model header, the qwen/qwen3.8-max slug, and the pricing, performance, and public-benchmarks tabs.

The bottom line

If you want to run a frontier model on your own hardware — or call one at a reasonable price — Qwen3.8 Max is the release that matters, and it is on OrcaRouter at Alibaba's list price, passed through at 0% markup, so the price cut Alibaba announced at launch is live on our side the same day. If you want to train or distill a reasoning model — or audit the signal that shaped Nemotron 3 Ultra — NVIDIA-Nemotron-Labs-Teacher-General-Reasoning is the release that matters, and it is self-host-only for now. The bigger story is that both landed in the same week: open weights just moved from "open enough to look at" to "open enough to build on," at two different points in the model supply chain. Teams that keep their model layer swappable behind one interface — self-hosted training on one side, managed frontier inference on the other — are the ones positioned to use both.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily