
NVIDIA-Nemotron-Labs-Teacher-General-Reasoning vs Qwen3.8-Max: The Week Open Weights Got Serious
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 114 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 53 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 59 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 367 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Open weights just had a week. On or around August 12, 2026, Alibaba released the first open-weight Max-class Qwen ever — Qwen3.8 Max, a 2.4-trillion-parameter flagship, published as Qwen3.8-2.4T-A95B on Hugging Face and ModelScope. Two days later, on August 14, NVIDIA published NVIDIA-Nemotron-Labs-Teacher-General-Reasoning, a 550B-parameter, 55B-active reasoning teacher from the Nemotron 3 Ultra training pipeline, on Hugging Face as well. Both are enormous, freshly-opened checkpoints from frontier labs, and the resemblance stops there. Qwen3.8 Max is open so you can run a frontier model yourself; NVIDIA-Nemotron-Labs-Teacher-General-Reasoning is open so you can train with it. Comparing them is really asking which kind of team you are — but the fact that both landed inside 72 hours is a signal about where the industry is heading.
What each release actually contains
Qwen3.8 Max is the deployable one. It is a sparse mixture-of-experts model with 2.4 trillion total parameters, about 95 billion active per token across 512 experts, built on the Qwen3.5-family hybrid architecture. In its open-weights form it is text-only with a 262K-token native context (extendable past 1M), and — notably — thinking is required: the model emits reasoning content between <think> tags and cannot be turned off. The hosted API version Alibaba launched on August 3 adds image and video input, a 1M default context, and optional thinking. The open-weights license is a custom Qwen3.8 Max License, permissive but with revenue-based conditions for very large commercial deployments. If you have the multi-node hardware, you can run a frontier-class model on your own infrastructure.
NVIDIA-Nemotron-Labs-Teacher-General-Reasoning is the training-time one. It is one of the more than ten domain-specialized teachers from the Multi-Teacher On-Policy Distillation (MOPD) recipe behind Nemotron 3 Ultra, derived from the post-trained Ultra student through an extra reasoning-specialized SFT and reinforcement-learning stage. Its role in that pipeline is to generate long reasoning traces on the hardest math, logic, and abstract-reasoning problems and to act as the judge that grades free-form answers. It ships under the commercial-friendly OpenMDW-1.1 license with the disclosed post-training data, and it carries zero benchmark numbers — NVIDIA points to an accuracy plot and the Ultra technical report instead. Its customers are distillation runs, not production traffic.
Qwen3.8 Max, the first open Max-class flagship
Alibaba's release matters because Max-class Qwen models have historically been API-only. Qwen3.8 Max breaks that: the weights are downloadable, and the model is genuinely frontier — Artificial Analysis gives it an Intelligence Index of 58 and an Agentic Index of 58, tying it with the top closed generalists in agentic settings. Its vendor-reported highlights are strong: 93.0 on PaperBench (ahead of GPT-5.6 Sol and Fable 5), 92.6 on GPQA Diamond, 86.1 on OSWorld-Verified. Its weak spots are equally real — 67.7 on SWE-bench Pro and 43.6 on Humanity's Last Exam trail the leaders, and independent testing found it expensive per finished task, averaging 64 turns per task on agentic runs. On Alibaba's API it prices at $2.00 per million input tokens, $6.00 per million output, and $0.25 per million cached input, a 20% cut from its preview pricing.
The teacher: training infrastructure with a model card
NVIDIA-Nemotron-Labs-Teacher-General-Reasoning is not competing on any of those numbers because it has not published any. What it offers instead is the same LatentMoE Mamba-2 + Transformer hybrid architecture as the Ultra student, up to 1M tokens of context, and a reasoning dial — enable_thinking true/false, a medium_effort mode, and a hard reasoning_budget ceiling — that a distillation pipeline needs in order to spend deep reasoning on hard samples and skip it on easy ones. Minimum hardware is 4×B200 (or 8×H100), serving through vLLM, SGLang, or TensorRT-LLM, and the checkpoint is a 1.12-terabyte download. It is not deployed by any inference provider, and NVIDIA has not announced NIM hosting. This is a weights-only release aimed at labs that already run large GPU fleets.
Specs, side by side
• Purpose — Teacher: training-time reasoning specialist and judge. Qwen3.8 Max: deployable frontier flagship.
• Scale — Teacher: 550B total / 55B active, LatentMoE with MTP. Qwen3.8 Max: 2.4T total / ~95B active, 512 experts, Qwen3.5 hybrid.
• Context — Teacher: up to 1M tokens, 256K default. Qwen3.8 Max: 262K native open-weights context, extendable past 1M; API version 1M default.
• Weights — Teacher: OpenMDW-1.1, released Aug 14, 2026. Qwen3.8 Max: Qwen3.8 Max License, released ~Aug 12, 2026.
• Independent evidence — Teacher: none yet. Qwen3.8 Max: AA Intelligence Index 58, Agentic Index 58, Coding Index 71.8.
• Thinking — Teacher: enable_thinking toggle, medium_effort, reasoning_budget ceiling. Qwen3.8 Max: required-on in open weights, cannot be disabled.
• Price — Teacher: no list price, self-hosted capex. Qwen3.8 Max: $2.00 / $6.00 per million tokens, $0.25 cached input.


The evidence asymmetry is the whole story
Qwen3.8 Max has been tested; the teacher has not. That is partly age — Qwen3.8 Max launched August 3 and has had two weeks of leaderboard runs, while the teacher shipped yesterday — and partly intention, because the teacher's card was never meant to compete on scores. But the asymmetry has a real consequence for the comparison: every concrete number on this page with a source attached belongs to Qwen3.8 Max, and every number attached to the teacher is an architecture spec. The teacher's reasoning quality is unverified by anyone outside NVIDIA, and its family-level figures (the Ultra student's vendor-reported SWE-Bench Verified 71.9, MMLU-Pro 86.8, LiveCodeBench v6 89.0) describe a different model. Anyone building a decision on "which scores better" has to wait for independent runs of the teacher — which is exactly the gap the next few weeks should close.
Two thinking dials, set differently
The most interesting technical contrast between these two releases is what happens when you ask them to reason. Qwen3.8 Max in its open-weights form has no choice: thinking is always on, and the reasoning trace is part of every answer. That is great for answer quality and transparency and expensive for bulk generation. The teacher treats reasoning as a dial: enable_thinking flips it, medium_effort throttles it, and a reasoning_budget ceiling caps it. For a distillation pipeline the difference is decisive — generating millions of reasoning traces with an always-on-thinking model multiplies your token bill, while the teacher's switch lets you pay for deep reasoning only where a hard sample demands it. These are not competing design philosophies; they are two tools optimized for opposite ends of the same problem, and each is the right tool for its end.

The bottom line
If you want to run a frontier model on your own hardware — or call one at a reasonable price — Qwen3.8 Max is the release that matters, and it is on OrcaRouter at Alibaba's list price, passed through at 0% markup, so the price cut Alibaba announced at launch is live on our side the same day. If you want to train or distill a reasoning model — or audit the signal that shaped Nemotron 3 Ultra — NVIDIA-Nemotron-Labs-Teacher-General-Reasoning is the release that matters, and it is self-host-only for now. The bigger story is that both landed in the same week: open weights just moved from "open enough to look at" to "open enough to build on," at two different points in the model supply chain. Teams that keep their model layer swappable behind one interface — self-hosted training on one side, managed frontier inference on the other — are the ones positioned to use both.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
