
Nemotron 4 Leak: NVIDIA's 1-Trillion-Parameter Open-Weights Answer to the Frontier
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 144 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 125 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 933 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 50 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 106 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 217 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Nemotron 4 is still unreleased, but it stopped being a rumor with just a parameter count this week. The Wall Street Journal reported on August 22, 2026, citing people familiar with the matter, that NVIDIA is spending roughly $6 billion to build Nemotron 4 into one of the world's most powerful open-weight AI models — and that the plan includes a non-exclusive license to AI startup Poolside's model-training technology plus the transfer of more than 100 of Poolside's engineers into the Nemotron 4 project. That sits on top of the earlier leak from The Information, on August 11, that the family's flagship will carry "at least 1 trillion parameters." Combined, the two reports describe NVIDIA's most serious attempt yet to reach the top of the open-weight frontier.
Nothing here has shipped, and none of it is confirmed. NVIDIA has acknowledged it is building a Nemotron 4 family, but it has not confirmed the parameter target, the timeline, the licensing-and-hiring plan the WSJ attributes to people familiar with the deal, or the alliance details The Information attributes to unnamed employees. Treat every number in this piece as "reported," not "specified." The point of a leak write-up is to separate the small confirmed core from the larger reported halo — and to tell you which parts matter when the real spec sheet lands.
The one-line leak, unpacked
First, the name, because it will confuse you in search results: NVIDIA already shipped a model called Nemotron-4-340B in June 2024. The new Nemotron 4 is a different, larger family — the successor to the Nemotron 3 line (Nano, Super, Ultra) that shipped this year — reusing the "Nemotron 4" name for the generation after it. This article is about the new one.
What The Information actually reported: NVIDIA is developing Nemotron 4 as a family of open-source models aimed at the top of the open-weight leaderboard, and the flagship is expected to have "at least 1 trillion parameters." NVIDIA has not set a release date, has not begun the final training run — a process employees expect to take months — and has settled only the pre-training data and the architecture; the specification is still moving. Two employees said the family could be ready as early as late fall; others expect later.
The same report attaches a separate budget: roughly $28 billion in multi-year cloud-service agreements running through early 2031, about three times what NVIDIA disclosed a year earlier, with around $7 billion landing in the current fiscal year. That figure covers compute commitments, not the new deal reported this week. The report also notes that the prior flagship's research paper listed 570 authors and Nemotron 4 involves more — one former employee told The Information that "everyone wants to be involved." All of it unreported by NVIDIA.

This week: a $6 billion Poolside deal to make Nemotron 4 real
The Wall Street Journal reported this week that NVIDIA's plan for Nemotron 4 is not just a research project — it is a funded build. Per the WSJ, citing people familiar with the matter, NVIDIA will pay roughly $6 billion for a non-exclusive license to Poolside's "Model Factory," the training-and-evaluation system behind Poolside's open-weight Laguna coding models, and make job offers to more than 100 of Poolside's staff — reported as 109, most of them engineers — to work on the Nemotron project. Bloomberg separately reported that NVIDIA will invest an additional $1 billion in Poolside at a reported $12 billion pre-money valuation, bringing NVIDIA's total reported commitment to roughly $7 billion.
It is not an acquisition. The license is non-exclusive, Poolside's founders remain, and Poolside keeps operating as an independent company; per the reporting, Poolside plans to distribute the $6 billion license payment to its investors. The story first surfaced through Newcomer on August 20 and was confirmed by Bloomberg on August 21 and the WSJ on August 22 — all still "reported," none of it confirmed by NVIDIA or Poolside.
Why the deal matters more than the headline number: it gives the Nemotron 4 project a team that has actually shipped open-weight models. Poolside's Laguna family is a proven open-weight coding lineup, and the reported goal — a Nemotron release within roughly a year that can compete with the strongest frontier models — is the first concrete time frame anyone has attached to the project. The reports frame the move as NVIDIA's answer to the open-weight leaders in China, notably DeepSeek and Moonshot AI's Kimi K3, and, in the open-versus-closed debate, as a hedge against the US closed labs.
The structure matters too. A non-exclusive license means Poolside keeps its intellectual property and can license the same technology to others — NVIDIA is buying capability and talent, not ownership. Analysts read it as NVIDIA's effort to keep the model layer commoditized: if open-weight models stay strong and cheap, demand keeps flowing to the layer NVIDIA controls — the GPUs. As LMArena CEO Anastasios Angelopoulos put it in the earlier reporting: "No matter which company makes a great open-source model, NVIDIA wins."
Why the size number is the least interesting part
The obvious story is "NVIDIA is going trillion-parameter." The less obvious one is that the open-weight frontier got there first. Kimi K3 opened its weights on July 27 with 2.8 trillion total parameters — about 104 billion active per token in a mixture-of-experts configuration — and Alibaba's Qwen3.8 Max, unveiled July 19, is a 2.4-trillion-total model with roughly 95 billion active. Against those, a Nemotron 4 flagship at "at least 1 trillion" would double Nemotron 3 Ultra but still trail the size leaders, and both of those leaders are Chinese.

That makes total parameter count the wrong lens entirely. In a mixture-of-experts model, what determines cost and speed is the active-parameter count per token, not the headline total. A 1-trillion-total Nemotron 4 could land at roughly 100 billion active — squarely in Kimi K3's and Qwen3.8 Max's tier — or at half that. The number that leaks next is the number that matters, and it hasn't leaked.

The other direction is worth noting too: DeepSeek V4-Flash, out July 31 under an MIT license, went the opposite way — 284 billion total, positioned on price rather than scale, at an estimated cost around $0.14 per million input tokens. The market is rewarding both extremes. Scale is not the strategy; it is one strategy.
NVIDIA's position explains the move. The Information reports Nemotron 3 Ultra ranks second among US open-weights models — behind Thinking Machines' Inkling — and outside the global open leaderboard's top tier. Nemotron 4 is the attempt to get back into the frontier conversation, and NVIDIA's own executives frame it as a sovereignty play: VP of generative AI Kari Briski said NVIDIA invests in Nemotron because "every company and every country needs accessible frontier open-source models to strengthen safety and security, accelerate innovation, and provide a foundation they can rely on from one generation to the next." Jensen Huang has been making the same case on X, arguing open models "strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty." The structural tension is real — NVIDIA reportedly owns roughly $30 billion of OpenAI — but the bet is that open models expand the GPU pie rather than shrink it.
The Nemotron Coalition is the part to watch
Nemotron 4 is not a solo project, and the Poolside deal is not the only structure around it. At GTC on March 16, 2026, NVIDIA announced the Nemotron Coalition — eight founding members: Black Forest Labs, Cursor, LangChain, Mistral AI, Perplexity, Reflection AI, Sarvam, and Thinking Machines Lab — organized to pool research, data, and compute behind "frontier open models." Its first project is a base model co-developed by Mistral AI and NVIDIA, trained on NVIDIA DGX Cloud, which the coalition says will be open-sourced and will underpin the Nemotron 4 family.
The Information adds the working details: Reflection, Cursor, Thinking Machines, and Mistral are confirmed contributors; Prime Intellect contributed 300,000 simulation environments for training; Cognition has discussed supplying code training data. Prime Intellect's CEO, Vincent Weisser, framed the alliance as collective action against what he called a single "god-like" model monopoly. What this means operationally: Nemotron 4 may ship as a consortium foundation that each member post-trains for its own product, rather than one vendor's checkpoint. That would make "open-weights NVIDIA model" the wrong category — the base is the interesting artifact.
The Poolside deal slots in alongside all of that. The coalition pools research and compute; the license-and-hire brings in a working model factory and a staff that has shipped open weights. If the reports hold, NVIDIA is assembling Nemotron 4 from three sources at once — its own training operation, a consortium base, and Poolside's production system — which is a different shape than any single-checkpoint release.
Confirmed, reported, unknown
• Confirmed — NVIDIA is working on a Nemotron 4 family (company statement). The Nemotron Coalition exists and is building an open base model with Mistral AI (NVIDIA, March 2026).
• Reported, unconfirmed — the "at least 1 trillion" flagship target and the late-fall timeline, plus the ~$28 billion cloud commitment with ~$7 billion this fiscal year (The Information); the WSJ-reported ~$6 billion build budget, non-exclusive Poolside "Model Factory" license, and more than 100 Poolside engineers joining the Nemotron project; the Bloomberg-reported $1 billion NVIDIA investment in Poolside at a ~$12 billion pre-money valuation; which coalition members contribute what.
• Unknown — active-parameter count, context length, modality mix, license terms, price, which checkpoints ship first, whether the Mistral-co-developed base actually becomes the family's foundation, and whether the Poolside deal is structured as reported (neither company has commented).
Timeline and what to watch
There is still no release date. The final training run has not started; employees describe it as months-long, and the estimates range from "late fall" (two sources) to later. The Poolside reporting adds the first concrete time frame anyone has attached to the project: per the coverage, the goal is a release within roughly a year that can compete with the strongest frontier models. NVIDIA has kept the family moving in the meantime: it shipped Nemotron 3.5 Lightning, a 31.6B-parameter agent worker released August 11 — the same day the Nemotron 4 report broke — alongside NeMo Switchyard, NVIDIA's own open-source routing software.
What to watch, in rough order of importance:
• Active parameters — total is the headline; active count decides cost and speed. A ~1T-total MoE at ~100B active changes the whole picture versus one at ~50B.
• The Poolside integration — whether the Model Factory license and its new team actually reshape the training plan and compress the reported timeline, or whether it is talent acquisition that leaves the plan standing.
• Independent scores — no third-party evaluation exists for a model that has not been trained. Watch the Artificial Analysis Intelligence Index when a hosted checkpoint appears.
• License terms — Nemotron 3 Ultra ships under OpenMDW-1.1, permissive but not MIT or Apache 2.0. Whether Nemotron 4 follows or tightens is undecided, and for a company building on the weights it decides everything.
• Which sizes ship first — Nemotron 3 arrived as a family (Nano, Super, Ultra). Expect a range of Nemotron 4 sizes, and expect the small ones before the flagship.
• The Mistral base — whether the coalition's base model actually becomes the family's foundation is the structural question behind all the parameter talk.
• Price — nothing is hosted, so nothing is priced. When a hosting partner lists Nemotron 4, the pass-through price is what you will actually pay.
What it means for your inference layer
You cannot call Nemotron 4 today — nobody can — and the Poolside reports, however dramatic, do not change that. The practical question is still how much of your stack should change to accommodate a model whose spec sheet is still moving. The answer is the same as for any unshipped frontier model: keep the inference layer model-agnostic. If you route through a single API that fronts 200+ models, a new Nemotron family is a config change, not a rewrite. OrcaRouter passes through the provider's list price with 0% markup, so whatever a host eventually charges for Nemotron 4 is exactly what you pay, and a vendor price cut lands the same day it is made. Automatic failover is the tool for an unproven first release: point early traffic at the new model, fall back to a stable alternative when it stalls or errors, and promote it to production only after it earns your workloads. None of that requires betting a production path on a leak — which is exactly the right posture when the spec sheet is this provisional.
FAQ
When will Nemotron 4 be released?
No date has been set. Employees told The Information the final training run has not started and will take months; two said late fall 2026 is possible, and others expect later. This week's reporting adds a reported goal of a frontier-competitive release within roughly a year, but still no firm date. Treat both as estimates from inside the project, not a roadmap.
Will Nemotron 4 be open-source?
NVIDIA's stated intent is open weights, and the Nemotron Coalition's first project — the Mistral-co-developed base model — is pledged to be open-sourced. But "open weights" is not "open source": Nemotron 3 Ultra ships under the OpenMDW-1.1 license, which is more permissive than NVIDIA's older terms but still not MIT or Apache 2.0. Nemotron 4's specific license has not been announced.
Is the reported $6 billion Poolside deal confirmed?
No. The WSJ reported the ~$6 billion license and the more than 100 Poolside engineers joining the Nemotron project on August 22, citing people familiar with the matter; Bloomberg reported the additional $1 billion investment. Newcomer first reported it on August 20. Neither NVIDIA nor Poolside has confirmed any of it. The license is reported as non-exclusive, and Poolside would remain an independent company.
Is "at least 1 trillion parameters" a real spec?
It is an employee-sourced target reported by The Information, not a confirmed specification, and the same report says the count could change. Even at 1 trillion total, it would trail Kimi K3 (2.8T) and Qwen3.8 Max (2.4T) in raw size; the active-parameter count — which has not leaked — is the number that will determine cost and speed.
Should I wait for Nemotron 4 before choosing a model?
No. It is months away at best and unproven. Pick the best available model for the job now and keep the routing layer flexible; when Nemotron 4 actually ships, evaluate it the same way you would any new model — active parameters, independent benchmarks, license, and price — rather than on the headline total or the size of its reported budget.
Bottom line
The family is real; the numbers are estimates. This week the story changed shape: Nemotron 4 is no longer just a parameter leak, it is a funded build. NVIDIA is reported to be spending ~$6 billion, licensing Poolside's Model Factory, hiring more than 100 of its engineers, and investing another $1 billion in the company — all to make an open-weight flagship that can stand next to DeepSeek and Kimi K3. Every one of those numbers is press-reported, none is confirmed, and nothing has shipped. The open-weight frontier has already passed the trillion-parameter mark, and NVIDIA's answer is a consortium base model, a licensed model factory, and a promise. For a reader choosing models today, that still means one thing: do not plan around the leak. Keep the inference layer flexible, evaluate whatever ships on its real numbers, and let routing do what routing is for — trying the new thing without betting the production path on it.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
