A hero title card for DeepSeek V4.1 Flash with the subtitle 'Two-day API beta: what we know so far', pill badges reading 'INTERMEDIATE BUILD' and 'API TEST - EXPIRES 09-10', a footer line 'Official release post expected very soon', and the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

DeepSeek V4.1 Flash Hits a Two-Day API Beta: New Architecture, Native Multimodal, and Pro-Level Ambition

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

DeepSeek V4.1 Flash showed up exactly where Deep​Seek tests things before it announces them: inside the live API, under a model ID that carries its own expiry date. The ID — deepseek-v4.1-flash-expires-on-0910 — began serving on September 8, 2026, after Deep​Seek's team posted an "intermediate version" internal test in its official community groups, and that trailing 0910 is the story in miniature. This is a build dropped into the real API environment for a limited test that is scheduled to go offline around September 10, ahead of the official release post that the tipsters flagging the model expect "very soon". It is not a launch. There is no model card, no technical report, no changelog entry, and as of this evening Deep​Seek's public Models & Pricing page still lists only DeepSeek V4 Flash, DeepSeek V4 Pro and DeepSeek-V4-Flash-Vision-Exp.

Everything below should be read with that asterisk in mind. The official channel here is a community-group announcement and an accompanying feedback form, not a release note. What follows is the union of that announcement, the coverage from Chinese media within hours of it, and first-day measurements from developers who pointed their own keys at the endpoint — with the vendor's claims and the community's numbers labelled as what they are.

What actually happened today

Around 15:00 Beijing time on September 8, Deep​Seek's team told its official community groups that an intermediate version — described in Chinese as a 中间版本, an in-between checkpoint — was being opened for limited testing inside the real API environment before a final version ships. Anyone with a Deep​Seek API key can call it: keep your existing base URL and key, and set the model name to deepseek-v4.1-flash-expires-on-0910. That is the entire access story — there is no separate endpoint, no waitlist, and no new console.

The practical envelope is tight. The suffix encodes the plan: the test version "will automatically expire and go offline on September 10", leaving roughly two days of uptime. Each account is capped at 20 concurrent requests — versus the 2,500-concurrency limit Deep​Seek publishes for deepseek-v4-flash — which is a strong hint about intent: this is for functional testing and evaluation, not for pointing production traffic at.

• What it is — an "intermediate version" checkpoint placed in the live API for short-term testing ahead of an official build.

• Where it runs — Deep​Seek's own API; base URL and key unchanged, model name swapped to deepseek-v4.1-flash-expires-on-0910.

• How long — the name says expires-on-0910; the test is expected to go offline around September 10.

• Who can call it — any account with a Deep​Seek key, at 20 concurrent requests per account.

A single-column scoreboard titled 'DeepSeek V4.1 Flash - the scoreboard' listing Model ID deepseek-v4.1-flash-expires-on-0910, Status 2-day API beta - expires 09-10, Concurrency 20 req/account (V4 Flash: 2,500), Billing same rate card as DeepSeek V4 Flash, Speed (community) ~420 tok/s peaks 500+, Native multimodal text + image (official claim), with a footer reading 'Speed is community-measured; no official benchmarks or specs published yet.'

What the beta bills: the DeepSeek V4 Flash rate card

Deep​Seek said the test is billed at the same rates as deepseek-v4-flash, and that model's current rate card is the one that applies. Since the August 16, 2026 repricing, Deep​Seek runs peak/off-peak pricing; the exact figures are the official ones for the Flash tier:

• Input, cache hit — $0.007 per 1M tokens off-peak, $0.014 peak

• Input, cache miss — $0.22 per 1M tokens off-peak, $0.44 peak

• Output — $0.66 per 1M tokens off-peak, $1.32 peak

• Peak hours — 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday; all other hours are off-peak at half the peak rate

A screenshot of the DeepSeek API docs Models & Pricing page (captured September 8, 2026) showing the current public lineup — deepseek-v4-flash, deepseek-v4-pro and deepseek-v4-flash-vision-exp — with 1M context and 384K max output, the off-peak/peak price columns (deepseek-v4-flash: $0.007 cache-hit input, $0.22 cache-miss input and $0.66 output off-peak, doubled at peak), and published concurrency limits of 2,500 / 500 / 2,500.

Note what did not change: the per-token price. Deep​Seek's claim that V4.1 Flash is "lower cost" is therefore an efficiency claim, not a rate cut — a point several first-day testers made the expensive way, watching a five-minute session drain noticeably more from their balance than the same wall-clock time on DeepSeek V4 Flash, because the model spends tokens far faster. Whether the per-task cost actually falls depends on whether it finishes the job with fewer tokens and fewer retries, which no one can judge from two days of speed tests.

What "new architecture, native multimodal" means — so far, unverified

Deep​Seek's official description of V4.1 Flash runs to four claims: a new model structure, native multimodal support, stronger capability, and faster generation. The architecture claim is the one that stands out, because it breaks a pattern. The previous refresh, DeepSeek V4 Flash 0731, was a post-training update that kept the same architecture and scale as its preview; Deep​Seek's wording here points to a structural change, and several outlets read it as a sign the model was re-pretrained rather than fine-tuned. Beyond that, everything is inference. Community speculation about changed attention mechanisms, expert networks, or an integrated vision module is exactly that — speculation. No parameter count, no context-window figure, no benchmark table, and no technical report has been published.

The "native multimodal" phrasing matters for a specific reason. DeepSeek-V4-Flash-Vision-Exp, launched August 21 and open-weights since August 31, bolted a vision encoder onto the DeepSeek V4 Flash text base — Deep​Seek's own materials described the pair as a text model plus an external vision path. V4.1 Flash is described as multimodal from the ground up, with image-and-text input handled by the base model itself. That is a real architectural difference in principle, but the modality spec has not been published: what testers have exercised is text-plus-images, and a reader should treat "multimodal" as confirmed only in that text-and-image sense until a model card appears. Whether a broader input set is part of the plan is, as of this writing, unknown rather than confirmed.

Speed is the headline — and it is a community measurement

The number doing the rounds on the first day is roughly 420 output tokens per second in long generations, with peaks reported above 500 tokens/s on individual runs. One widely shared test generated more than 71,000 tokens in under three minutes. None of this is official: these are developer measurements against the beta endpoint, under uncontrolled conditions, and Deep​Seek has published no speed benchmark of its own. For scale, Artificial Analysis measures the public DeepSeek V4 Flash 0731 endpoint at roughly 128 tokens/s, so the early reports point to a raw-throughput jump on the order of three times or more — with the usual caveat that throughput depends on input length, concurrency, and server conditions.

The end-to-end comparisons against DeepSeek-V4-Flash-Vision-Exp are where the gap looks widest, because they measure the same task back to back. Testers reported roughly 6× faster end-to-end on an SVG-code generation task, about 5.2× on a 49k-token long-context retrieval, around 5× on a large SQL generation-and-optimization task, 4.6× on an algorithmic problem, and 3.9× on an asynchronous-refactor task. Treat those as indicative, not precise: same-task end-to-end numbers fold in thinking time, first-token latency, and generation speed, and they come from individual testers rather than a controlled suite. The consistent direction across all of them is the interesting signal, not any single multiplier.

The question at the heart of the beta

The most strategic detail is hiding in the feedback form Deep​Seek attached to the test. Alongside questions about agent frameworks and real-world scenarios, it asks testers directly whether the DeepSeek V4.1 Flash intermediate version "can fully replace the online DeepSeek V4 Pro". That is a remarkable question for a company to put in front of users, because DeepSeek V4 Pro sits three price tiers above Flash: at current official rates the Pro model charges $0.66 per 1M input tokens (cache miss) and $1.98 per 1M output tokens off-peak, against $0.22 and $0.66 for DeepSeek V4 Flash — a clean 3× on every line. If a Flash-tier model genuinely approaches Pro-level capability at Flash-level prices, the question answers itself for a large slice of users, and Deep​Seek knows it.

Read together with the "new structure" phrasing, the questionnaire hints that V4.1 Flash is the first visible step of something bigger than a point refresh — a Flash line repositioned to compress the gap to the flagship, the way Deep​Seek has repeatedly used a cheaper, faster model to reset the market's expectations of what a price tier buys. None of that is confirmed by a benchmark; it is a strategic read of a two-sentence question. But it is the most interesting thing Deep​Seek has said about V4.1 Flash all day.

Where to try it, and what to run in production meanwhile

During the test window, the only place to reach V4.1 Flash is Deep​Seek's own API — there is no third-party hosting of a build that expires in two days, and no one should wire a production path to a model ID that is scheduled to stop answering. The sensible use of the next two days is evaluation: run your representative workloads, measure quality and real token economics, and treat every speed number you see as anecdotal until a formal release appears with a model card.

For traffic that has to keep working, the public Deep​Seek lineup is unchanged, and that is where a routing layer earns its keep. On OrcaRouter, DeepSeek V4 Flash, DeepSeek V4 Pro and DeepSeek-V4-Flash-Vision-Exp are all reachable through one API at Deep​Seek's list price passed through with zero markup — so the August price cut has been live on our side since the day it landed, not whenever a margin spreadsheet got around to it. When a formal V4.1 Flash eventually ships to providers, the same setup is the low-switch-cost way to adopt it: send a slice of traffic to the new model, keep DeepSeek V4 Flash or V4 Pro as an automatic failover, and let the router DSL decide which calls deserve the unproven model and which should stay on the workhorse. You do not have to bet a production path on a two-day beta to find out whether it is good.

A screenshot of the OrcaRouter model page for deepseek/deepseek-v4-flash (captured September 5, 2026) showing the Tools, JSON and Reasoning badges, a 1M-token context, 384K max output, text input, a p50 time-to-first-token of 434 ms, and the description 'DeepSeek V4 Flash efficient MoE — 284B total / 13B active params, 1M context, optimized for fast everyday workloads.'

Open questions

• When does the official release come? The signal that flagged this model says a release post is expected "very soon"; the beta's two-day lifespan suggests Deep​Seek does not intend to keep the world waiting long.

• Does the final build match this one? Independent testers who track Deep​Seek's testing cycles note the company appears to be running multiple candidate builds in parallel, and the fastest candidate in an early test is not always the one that ships — so today's speed numbers may not describe the GA model.

• What are the specs? Parameter count, architecture details, context length, and the full input-modality list are all unpublished, and all are the first things a real review needs.

• Does the price hold, and does "lower cost" survive contact with real workloads? The beta bills at DeepSeek V4 Flash rates; a launch could bring a new rate card, and per-task cost is unmeasured.

• Can it really pull Pro duty? The questionnaire asks the question, but only independent evals on agentic and reasoning suites — the benchmarks that separate the Flash tier from the Pro tier today — can answer it.

If you have a Deep​Seek key, the two-day clock is the story: call deepseek-v4.1-flash-expires-on-0910 before it vanishes, run your own workloads, and file the numbers under "community-measured, conditions vary". For everyone else, the thing to watch is the release post, and the question behind it — whether Deep​Seek's cheapest tier just started aiming at its most expensive one.

While the two-day beta sorts itself out, the stable Flash model you can actually run today is one API call away: DeepSeek V4 Flash on OrcaRouter — at Deep​Seek's list price, passed through with 0% markup.

And the flagship the beta questionnaire keeps asking testers to compare V4.1 Flash against: DeepSeek V4 Pro on OrcaRouter.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily