
DeepSeek V4.1 Pro Is in the Works: A Routing Notice, an Architecture Promise, and What's Still Unknown
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2123Intelligence49Coding
DeepSeek V4.1 Pro has not been announced, listed, benchmarked, or priced. It appears on no models page, has no card, and no endpoint answers to its name — yet on September 9, 2026, a member of DeepSeek's technical team publicly put it on the record, almost in passing, inside a two-sentence notice about what will happen to DeepSeek V4 Pro traffic in the meantime. Tianyi Cui, who posts under @tianyi, wrote that once DeepSeek V4.1 Flash officially launches — expected around September 10, Beijing time — all API requests aimed at DeepSeek V4 Pro will be rerouted to DeepSeek V4.1 Flash and billed at Flash's lower rate, and that the arrangement holds until DeepSeek V4.1 Pro itself goes live.
That trailing clause is the news. Until now the V4.1 line had surfaced only through the two-day "intermediate version" beta that began serving on September 8 under the model ID deepseek-v4.1-flash-expires-on-0910, plus a feedback form asking testers whether V4.1 Flash could replace DeepSeek V4 Pro in production. Naming DeepSeek V4.1 Pro as the endpoint of the routing window is the first time the flagship has been acknowledged from inside the company. The pseudonymous DeepSeek watcher teortaxesTex, who quoted the notice, read it the same way — and then added a stronger claim: that the existence of DeepSeek V4.1 Pro is confirmation DeepSeek has "solved architectural issues of the V4 family, as they committed to do in the V4 paper."
One sourcing note before the details. Everything below about DeepSeek V4.1 Pro rests on two things: a single dated social post by a DeepSeek technical-team member, corroborated within hours by several Chinese outlets, and one tracker's interpretation of it. There is no technical report, no parameter count, no release date, and no way to call the model today. The routing notice is a statement; the "architecture solved" reading is inference. This piece keeps the two apart, because they are different kinds of evidence.
What the September 9 notice actually said
Around 15:14 Beijing time on September 9, Tianyi Cui wrote on X that continuing to sell the old DeepSeek V4 Pro at a premium no longer made sense. His reasoning, in substance: internal and external testing shows the V4.1 Flash model comprehensively surpasses DeepSeek V4 Pro on performance, cost, speed, and total processing time, so charging users more money, slower speeds, and greater compute for a weaker model would be inappropriate. Therefore, once V4.1 Flash officially goes live and until V4.1 Pro goes live, requests for the V4 Pro model will be uniformly routed to V4.1 Flash and billed at Flash pricing.
That "comprehensively surpasses" claim is DeepSeek-team-reported, not independently verified: no benchmark table accompanied the post, and the V4.1 Flash beta has no model card either. What is verifiable is the shape of the plan. DeepSeek is treating its current Pro tier as obsolete the moment Flash is official, and parking Pro traffic on the cheaper model until the next Pro exists.
What it means if you call DeepSeek V4 Pro today
For anyone running the current flagship — DeepSeek V4 Pro, the 1.6-trillion-parameter MoE that reached GA in mid-August — the change is automatic and server-side. Once the reroute takes effect, a request that names the V4 Pro model is answered by the V4.1 Flash build and billed at the Flash rate card, with no code change on the caller's side. The billing switch is the sharpest part. Chinese media covering the notice put the new Flash card's peak output at ¥8 per million tokens from September 10, against ¥27 per million for the Pro tier it replaces — a roughly 70% cut for traffic that used to hit Pro, on a model DeepSeek's team says is faster. As a simultaneous upgrade-and-price-cut announcement, it is aggressive even by DeepSeek's standards.

A timing caveat, because the framing matters: if you hold a DeepSeek V4 Pro key today, nothing changes until V4.1 Flash actually launches, and the reroute description comes from the team's post and the coverage of it — not yet from a changelog entry. As of the September 9 capture above, DeepSeek's official models page still listed the plain V4 lineup — DeepSeek V4 Flash, DeepSeek V4 Pro and DeepSeek-V4-Flash-Vision-Exp — with no V4.1 entry in the table.
Why "the V4 paper" is the key to the leak
teortaxesTex's specific claim is that DeepSeek V4.1 Pro being in the works confirms DeepSeek has solved architectural issues of the V4 family "as they committed to do in the V4 paper." That reference is checkable. DeepSeek's V4 technical report — DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence — is unusually candid about what its design does not yet do well, and about what future iterations were meant to address:
• The report calls its own architecture "relatively complex" and commits that "in future iterations, we will carry out more comprehensive and principled investigations to distill the architecture down to its most essential designs."
• Its two training-stabilization mechanisms — Anticipatory Routing and SwiGLU Clamping — worked in practice, the report says, but their underlying principles remained "insufficiently understood."
• Long-context retrieval still decays past 128K tokens. On an 8-needle retrieval test cited in the report, V4-Pro-Max scores roughly 0.92 at 128K and falls to about 0.59 at 1M — the gap between "can load a million tokens" and "can actually use them."
• And the V4 line shipped text-only; vision arrived later as a bolted-on experimental encoder, DeepSeek-V4-Flash-Vision-Exp, rather than as part of the base model.
Read against that list, the V4.1 story snaps into focus. DeepSeek's own description of the V4.1 Flash beta uses exactly the language those issues implied — a "new model structure" and native multimodal support from the ground up. A V4.1 Pro built on the same resolved base is the natural next step, which is precisely the inference teortaxesTex is drawing. It is an inference, and it should stay labelled as one: DeepSeek has published no V4.1 technical report, and "in the works" is not confirmed beyond a team member's reference to Pro's eventual launch.
The evidence ledger — what is confirmed and what is not
Layering the September 9 signal onto the September 8 beta gives a short list of confirmed facts and a long list of open ones.

• Vendor statements (corroborated by several Chinese outlets within hours) — DeepSeek has a V4.1 Flash model it plans to launch around September 10, Beijing time; DeepSeek V4 Pro API traffic will be routed to it and billed at Flash rates until V4.1 Pro launches; and DeepSeek's team believes the Flash build beats the current Pro on cost, speed, and total time.
• Plausible but unverified — that DeepSeek V4.1 Pro is currently in training or otherwise imminent. teortaxesTex's "in the works" carries no date and no internal source; the only corroboration is the routing window itself, which presumes a Pro will eventually exist.
• Interpretation — that Pro's existence proves the V4 paper's architectural promises were kept. The V4.1 Flash beta's "new structure, native multimodal" claims point the same direction, but no V4.1 architecture has been published to check them against.
• Unknown — V4.1 Pro's specs, parameter count, context window, pricing, launch date, and whether it ships open weights the way DeepSeek V4 Flash and DeepSeek V4 Pro did. The current V4.1 Flash beta is API-only and set to expire, so the open-weights question is genuinely open, not just unanswered.
What to watch
• September 10, Beijing time — the V4.1 Flash official launch, its new rate card, and the start of the V4 Pro reroute. If it happens as described, the "DeepSeek V4 Pro" model ID begins returning a V4.1 architecture at Flash prices from that day.
• A V4.1 technical report or model card — the beta has neither, and the first card is what would turn the architecture claims into something a reviewer can check.
• First independent benchmarks — whether the V4.1 base actually fixes the V4 report's known gaps: retrieval past 128K, and the reasoning deficit the report itself put at roughly three to six months behind frontier models.
• V4.1 Pro naming and timing — whether the model teortaxesTex calls DeepSeek V4.1 Pro appears as a distinct card, and how quickly after Flash it lands.
What to do meanwhile
There is nothing to call yet. DeepSeek V4.1 Pro is not on any API, and wiring production to a two-day beta model ID would be its own mistake. The reroute news does change one calculation for current DeepSeek V4 Pro users, though: the model they are paying flagship rates for is about to be officially declared superseded, which argues against signing a long commitment to it now and in favour of watching the Flash launch before the next buying decision.
That wait-and-see posture is exactly where a routing layer earns its keep. On OrcaRouter, DeepSeek V4 Flash, DeepSeek V4 Pro and DeepSeek-V4-Flash-Vision-Exp are reachable through one API at DeepSeek's list price passed through with zero markup — so a vendor-side price change or reroute shows up on our side the same day it ships, not whenever a rate sheet gets updated. When V4.1 Flash goes official and, later, if V4.1 Pro appears, the same setup is the low-switch-cost way to adopt them: send a slice of traffic at the new model, keep a stable model as automatic failover, and let the routing DSL decide which calls deserve an unproven model and which should stay on the workhorse.

The bottom line on a two-sentence roadmap
The cleanest way to read the September 9 signal is that DeepSeek has told the market, in the most offhand way possible, that its current flagship is a stopgap. V4.1 Flash is the proof of launch, the reroute is the admission that DeepSeek V4 Pro as sold today is already superseded, and the single phrase "until V4.1 Pro launches" is the roadmap. What is missing from that sentence is everything that would turn DeepSeek V4.1 Pro from a leak into a product decision: a date, a spec, a price, a benchmark, and any evidence beyond one tracker's read that the architectural promises in the V4 paper have actually been kept. Until DeepSeek publishes something, the rational position is the one teortaxesTex holds — treat V4.1 Pro as real, treat every detail about it as unconfirmed, and treat the Flash launch around September 10 as the first verifiable instalment of the V4.1 architecture.
While DeepSeek V4.1 Pro has nothing to call yet, the current flagship you can reach through one API today is DeepSeek V4 Pro on OrcaRouter — at DeepSeek's list price, passed through with 0% markup.
And the model DeepSeek says V4 Pro traffic will reroute to, if you want to compare against it directly: DeepSeek V4 Flash on OrcaRouter.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
