
Nemotron 4 Leak: NVIDIA's 1-Trillion-Parameter Open-Weights Answer to the Frontier
- metaNEWMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenNEWQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekNEWDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxNEWMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens · 2278 tok/s
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Intelligence77Coding
- grokxAI: Grok 4.52026-07-0856Intelligence72Coding
- tencentTencent: Hy32026-07-0642Intelligence59Coding
- obsidianQwen3.6 35B A3B Uncensored (Aggressive)2026-07-0232Intelligence42Coding
- obsidianGemma4 26B A4B Uncensored (Balanced)2026-07-0226Intelligence39Coding
- anthropicAnthropic: Claude Sonnet 52026-06-3055Intelligence72Coding
- klingKling: Kling 3.0 Turbo2026-06-1757Intelligence52Coding57Math
1 trillion parameters. That's the reported floor for the largest Nemotron 4 model, the flagship of NVIDIA's next open-weights family — and the number only sounds big until you put it next to what already shipped. In the past month, Moonshot AI shipped Kimi K3 at 2.8 trillion total parameters, Alibaba unveiled Qwen3.8-Max at 2.4 trillion, and NVIDIA's own current flagship, Nemotron 3 Ultra, sits at 550 billion. The Information reported on August 11, 2026 — citing multiple employees working on the project — that Nemotron 4 is NVIDIA's answer: an open-weights family whose largest model is expected to start at roughly 1 trillion parameters, about double Nemotron 3 Ultra.
Nothing here has shipped. NVIDIA has acknowledged it is building a Nemotron 4 family, but it has not confirmed the parameter target, the budget, the timeline, or the alliance details The Information attributes to unnamed employees. Treat every number in this piece as "reported," not "specified." The point of a leak write-up is to separate the small confirmed core from the larger reported halo — and to tell you which parts matter when the real spec sheet lands.
The one-line leak, unpacked
First, the name, because it will confuse you in search results: NVIDIA already shipped a model called Nemotron-4-340B in June 2024. The new Nemotron 4 is a different, larger family — the successor to the Nemotron 3 line (Nano, Super, Ultra) that shipped this year — reusing the "Nemotron 4" name for the generation after it. This article is about the new one.
What The Information actually reported: NVIDIA is developing Nemotron 4 as a family of open-source models aimed at the top of the open-weight leaderboard, and the flagship is expected to have "at least 1 trillion parameters." NVIDIA has not set a release date, has not begun the final training run — a process employees expect to take months — and has settled only the pre-training data and the architecture; the specification is still moving. Two employees said the family could be ready as early as late fall; others expect later.
The same report attaches a budget: roughly $28 billion in multi-year cloud-service agreements running through early 2031, about three times what NVIDIA disclosed a year earlier, with around $7 billion landing in the current fiscal year. It also notes that the prior flagship's research paper listed 570 authors and Nemotron 4 involves more — one former employee told The Information that "everyone wants to be involved." All of it unreported by NVIDIA.

Why the size number is the least interesting part
The obvious story is "NVIDIA is going trillion-parameter." The less obvious one is that the open-weight frontier got there first. Kimi K3 opened its weights on July 27 with 2.8 trillion total parameters — about 104 billion active per token in a mixture-of-experts configuration — and Alibaba's Qwen3.8-Max, unveiled July 19, is a 2.4-trillion-total model with roughly 95 billion active. Against those, a Nemotron 4 flagship at "at least 1 trillion" would double Nemotron 3 Ultra but still trail the size leaders, and both of those leaders are Chinese.

That makes total parameter count the wrong lens entirely. In a mixture-of-experts model, what determines cost and speed is the active-parameter count per token, not the headline total. A 1-trillion-total Nemotron 4 could land at roughly 100 billion active — squarely in Kimi K3's and Qwen3.8-Max's tier — or at half that. The number that leaks next is the number that matters, and it hasn't leaked.

The other direction is worth noting too: DeepSeek V4-Flash, out July 31 under an MIT license, went the opposite way — 284 billion total, positioned on price rather than scale, at an estimated cost around $0.14 per million input tokens. The market is rewarding both extremes. Scale is not the strategy; it is one strategy.
NVIDIA's position explains the move. The Information reports Nemotron 3 Ultra ranks second among US open-weights models — behind Thinking Machines' Inkling — and outside the global open leaderboard's top tier. Nemotron 4 is the attempt to get back into the frontier conversation, and NVIDIA's own executives frame it as a sovereignty play: VP of generative AI Kari Briski said NVIDIA invests in Nemotron because "every company and every country needs accessible frontier open-source models to strengthen safety and security, accelerate innovation, and provide a foundation they can rely on from one generation to the next." Jensen Huang has been making the same case on X, arguing open models "strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty." The structural tension is real — NVIDIA reportedly owns roughly $30 billion of OpenAI — but the bet is that open models expand the GPU pie rather than shrink it. As LMArena CEO Anastasios Angelopoulos put it: "No matter which company makes a great open-source model, NVIDIA wins."
The Nemotron Coalition is the part to watch
Nemotron 4 is not a solo project, and that is the detail most coverage misses. At GTC on March 16, 2026, NVIDIA announced the Nemotron Coalition — eight founding members: Black Forest Labs, Cursor, LangChain, Mistral AI, Perplexity, Reflection AI, Sarvam, and Thinking Machines Lab — organized to pool research, data, and compute behind "frontier open models." Its first project is a base model co-developed by Mistral AI and NVIDIA, trained on NVIDIA DGX Cloud, which the coalition says will be open-sourced and will underpin the Nemotron 4 family.
The Information adds the working details: Reflection, Cursor, Thinking Machines, and Mistral are confirmed contributors; Prime Intellect contributed 300,000 simulation environments for training; Cognition has discussed supplying code training data. Prime Intellect's CEO, Vincent Weisser, framed the alliance as collective action against what he called a single "god-like" model monopoly. What this means operationally: Nemotron 4 may ship as a consortium foundation that each member post-trains for its own product, rather than one vendor's checkpoint. That would make "open-weights NVIDIA model" the wrong category — the base is the interesting artifact.
Confirmed, reported, unknown
• Confirmed — NVIDIA is working on a Nemotron 4 family (company statement). The Nemotron Coalition exists and is building an open base model with Mistral AI (NVIDIA, March 2026).
• Reported, unconfirmed — the "at least 1 trillion" flagship target; the late-fall timeline; the ~$28 billion cloud commitment with ~$7 billion this fiscal year; which coalition members contribute what.
• Unknown — active-parameter count, context length, modality mix, license terms, price, which checkpoints ship first, and whether the Mistral-co-developed base actually becomes the family's foundation.
Timeline and what to watch
There is no release date. The final training run has not started; employees describe it as months-long, and the estimates range from "late fall" (two sources) to later. NVIDIA has kept the family moving in the meantime: it shipped Nemotron 3.5 Lightning, a 31.6B-parameter agent worker released August 11 — the same day the Nemotron 4 report broke — alongside NeMo Switchyard, NVIDIA's own open-source routing software. In other words, NVIDIA is actively iterating the lineup while the flagship is still on the drawing board.
What to watch, in rough order of importance:
• Active parameters — total is the headline; active count decides cost and speed. A ~1T-total MoE at ~100B active changes the whole picture versus one at ~50B.
• Independent scores — no third-party evaluation exists for a model that has not been trained. Watch the Artificial Analysis Intelligence Index when a hosted checkpoint appears.
• License terms — Nemotron 3 Ultra ships under OpenMDW-1.1, permissive but not MIT or Apache 2.0. Whether Nemotron 4 follows or tightens is undecided, and for a company building on the weights it decides everything.
• Which sizes ship first — Nemotron 3 arrived as a family (Nano, Super, Ultra). Expect a range of Nemotron 4 sizes, and expect the small ones before the flagship.
• The Mistral base — whether the coalition's base model actually becomes the family's foundation is the structural question behind all the parameter talk.
• Price — nothing is hosted, so nothing is priced. When a hosting partner lists Nemotron 4, the pass-through price is what you will actually pay.
What it means for your inference layer
You cannot call Nemotron 4 today — nobody can — so the practical question is how much of your stack should change to accommodate a model whose spec sheet is still moving. The answer is the same as for any unshipped frontier model: keep the inference layer model-agnostic. If you route through a single API that fronts 200+ models, a new Nemotron family is a config change, not a rewrite. OrcaRouter passes through the provider's list price with 0% markup, so whatever a host eventually charges for Nemotron 4 is exactly what you pay, and a vendor price cut lands the same day it is made. Automatic failover is the tool for an unproven first release: point early traffic at the new model, fall back to a stable alternative when it stalls or errors, and promote it to production only after it earns your workloads. None of that requires betting a production path on a leak — which is exactly the right posture when the spec sheet is this provisional.
FAQ
When will Nemotron 4 be released?
No date has been set. Employees told The Information the final training run has not started and will take months; two said late fall 2026 is possible, and others expect later. Treat "late fall" as a hopeful estimate from inside the project, not a roadmap.
Will Nemotron 4 be open-source?
NVIDIA's stated intent is open weights, and the Nemotron Coalition's first project — the Mistral-co-developed base model — is pledged to be open-sourced. But "open weights" is not "open source": Nemotron 3 Ultra ships under the OpenMDW-1.1 license, which is more permissive than NVIDIA's older terms but still not MIT or Apache 2.0. Nemotron 4's specific license has not been announced.
Is "at least 1 trillion parameters" a real spec?
It is an employee-sourced target reported by The Information, not a confirmed specification, and the same report says the count could change. Even at 1 trillion total, it would trail Kimi K3 (2.8T) and Qwen3.8-Max (2.4T) in raw size; the active-parameter count — which has not leaked — is the number that will determine cost and speed.
Should I wait for Nemotron 4 before choosing a model?
No. It is months away at best and unproven. Pick the best available model for the job now and keep the routing layer flexible; when Nemotron 4 actually ships, evaluate it the same way you would any new model — active parameters, independent benchmarks, license, and price — rather than on the headline total.
Bottom line
The family is real; the numbers are estimates. NVIDIA has confirmed it is building Nemotron 4, and the coalition structure around it is the most interesting part of the story — but every headline figure in the leak, from the 1-trillion floor to the late-fall timeline to the $28 billion budget, comes from unnamed employees and could move. The open-weight frontier has already passed the trillion-parameter mark, led by Chinese labs, and NVIDIA's response is a consortium base model plus a promise. For a reader choosing models today, that means one thing: do not plan around the leak. Keep the inference layer flexible, evaluate whatever ships on its real numbers, and let routing do what routing is for — trying the new thing without betting the production path on it.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
