A generated hero card titled MiniMax M3.1 Flash Preview vs DeepSeek V4 Flash, subtitled 'The same 1M-token window, two very different sales pitches'. A left card reads 'MiniMax M3.1 Flash Preview: 1M context, price not published, Token Plan only, thinking cannot be disabled'; a right card reads 'DeepSeek V4 Flash: 1M context, $0.15/$0.60 off-peak, text only, superseded by V4.1 Flash'. A footer reads 'MiniMax figures per platform.minimax.io; DeepSeek figures per DeepSeek's published rate card'.
Engineering & Research

MiniMax M3.1 Flash Preview vs DeepSeek V4 Flash: Two Cheap 1M-Context Tickets, One Huge Asterisk

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

If you are shopping for a million-token context window at a low price per token, MiniMax-M3.1-Flash-Preview and DeepSeek V4 Flash are the two names that keep coming up — and they are not really the same kind of product. DeepSeek V4 Flash is a retired model whose identifier still answers, living inside a vendor rate card you can read to the cent. MiniMax-M3.1-Flash-Preview is a live preview with no published rate at all. The comparison below is therefore partly a spec comparison and partly a demonstration of how differently these two vendors are choosing to sell the same bracket.

Both sides of it are worth stating precisely, because the numbers that decide it are scattered across a vendor pricing page, a vendor documentation tree and our own catalogue, and one of the most frequently repeated claims about DeepSeek V4 Flash — its price — is one DeepSeek has since replaced.

What the two spec sheets actually match on

The headline specs are close enough that they are not the deciding factor, with one exception nobody advertises.

• Context window — 1,000,000 tokens for MiniMax-M3.1-Flash-Preview against 1,048,576 for DeepSeek V4 Flash: nominally identical, a 48,576-token difference
• Maximum output — not published for MiniMax-M3.1-Flash-Preview; 384,000 tokens for DeepSeek V4 Flash, which is unusual at this price and is the number that should decide a lot of purchasing
• Input modalities — text, image and video for MiniMax-M3.1-Flash-Preview; text only for DeepSeek V4 Flash
• Thinking control — an effort ladder from low to max, thinking permanently on, for MiniMax-M3.1-Flash-Preview; thinking or non-thinking on DeepSeek V4 Flash, with non-thinking as a real option
• Protocol support — Anthropic-compatible, OpenAI-compatible and OpenAI Responses endpoints for MiniMax-M3.1-Flash-Preview; the same three plus chat prefix completion on DeepSeek V4 Flash
• Weights — none for MiniMax-M3.1-Flash-Preview; DeepSeek's V4 Flash weights were published, though the model itself has been superseded
• Price — unpublished for MiniMax-M3.1-Flash-Preview; published and low for DeepSeek V4 Flash

Read that list twice, because the output ceiling is the quiet one. A million tokens of context with 384,000 tokens of output is a genuinely long single-pass generation window. A million tokens of context with an undisclosed output ceiling is a promise about reading, not about writing. If your workload is "read a repository, emit a migration plan", the second number is the one that determines whether the task fits in one call.

What each one is measured at

This is the least comfortable section of the comparison, and it is short.

DeepSeek V4 Flash has a benchmark record, and it is independent. Artificial Analysis scores it 34.3 on the Intelligence Index — 42nd of 145 models on that index — with 69.1 on the Coding Index and 90.8% on GPQA Diamond. The 95.0 on τ²-Bench that our own catalogue entry leads with is the same figure DeepSeek's launch material carried, so it is a vendor number in independent company rather than an audited one. Its long-context recall is 79.7% and its terminal-bench figure is 78.7% on the v2.1 harness. Those are real, comparable measurements of a known model.

MiniMax-M3.1-Flash-Preview has none. MiniMax published no evaluation table with the release, and no third party has one either — the model does not appear on Artificial Analysis's index at all. That is not a gap in our research; it is the current state of the public record, and it means every quality claim about the newer model is, at this moment, a vendor adjective. Anyone who hands you a MiniMax-M3.1-Flash-Preview benchmark table today is either quoting the older MiniMax-M3 or making it up.

A generated two-column scoreboard titled 'MiniMax M3.1 Flash Preview vs DeepSeek V4 Flash — the scoreboard'. The left column, MiniMax M3.1 Flash Preview, reads 'AA Intelligence: none published', 'Coding score: none published', 'Context window: 1M', 'Max output: not published', 'Weights: none', 'Price: not published'. The right column, DeepSeek V4 Flash, reads 'AA Intelligence: 34.3', 'Coding score: 69.1', 'Context window: 1,048,576', 'Max output: 384,000', 'Weights: published', 'Price: $0.15 / $0.60 off-peak'. A footer reads 'MiniMax figures per platform.minimax.io documentation, 27 September 2026; DeepSeek figures per DeepSeek's published rate card and Artificial Analysis. The MiniMax column is empty because nothing has been published, not because a result is pending.'

DeepSeek V4 Flash is not sold at the price you remember

Here is the trap. The figure widely quoted for DeepSeek V4 Flash — $0.14 per million input tokens and $0.28 per million output — was the card the model carried before 10 September 2026. On that date DeepSeek released DeepSeek-V4.1-Flash, retired both V4 Flash and the V4-Flash-Vision-Exp experiment, and cut prices. The deepseek-v4-flash identifier still works, but DeepSeek's documentation states that it is temporarily routed to V4.1 Flash and billed at the Flash price. In other words, you can still type the old model name; you are not calling the old model.

So the real card to compare against is the current one, per 1M tokens — and off-peak is the cheaper half of it:

• Cache-miss input — $0.15 off-peak, $0.30 at peak
• Cache-hit input — $0.003 off-peak, $0.006 at peak
• Output — $0.60 off-peak, $1.20 at peak
• Peak windows — 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, excluding Chinese public holidays; everything else, including full weekends, is off-peak

Against that, MiniMax-M3.1-Flash-Preview has no per-token rate of any kind. Its cost is whatever your Token Plan seat costs — $22, $55 or $132 per month for Plus, Max and Ultra — burned down through a shared quota pool at each model's pay-as-you-go list price, with the vendor reserving the right to change the composition of that pool. For a fixed monthly budget with predictable volume that can be excellent value. For a project that needs to attribute cost per call, it is opacity dressed as a subscription.

The reason this matters more than usual on the DeepSeek side is the peak multiplier. A team in Europe or Asia running its working day is inside a peak window for a meaningful share of its traffic and pays double for the privilege, which turns a headline $0.15 input rate into $0.30 for exactly the hours most products are busy. Budget off the peak column unless your traffic genuinely runs at night in UTC.

Where MiniMax M3.1 Flash Preview is the wrong call

Three cases, and the first two are easy to miss because they are documented rather than subtracted.

The output ceiling is unknown. If your task ends in a long generation — a full specification, a transcript rewrite, an annotated report — you are choosing an output budget you cannot see. DeepSeek V4 Flash publishes 384,000. That asymmetry favours the older model for anything that writes a lot.

Thinking cannot be turned off. Send thinking: {"type": "disabled"} to MiniMax-M3.1-Flash-Preview and you get an HTTP 400: the model requires adaptive thinking. On DeepSeek V4 Flash, non-thinking mode exists and is a supported switch. For high-throughput classification, routing, or short structured extraction, a model that always reasons first is paying latency and tokens you did not ask for — and the only lever you have is dropping effort to low, not removing the reasoning.

It is not callable on pay-as-you-go. The vendor's documentation is unambiguous that MiniMax-M3.1-Flash-Preview is available only through Token Plan and MiniMax Code right now, and the Subscription Key is not interchangeable with a pay-as-you-go API key. If you already hold a seat, it is a one-line change. If you do not, evaluating this model means buying a subscription.

Where the DeepSeek number is the wrong call

The mirror image is that DeepSeek V4 Flash is text-only and already superseded. It never accepted images — that work needed the separate experimental build, which has now been retired alongside it — and the model identifier now serves different weights than the name suggests. A pipeline pinned to deepseek-v4-flash for reproducibility is relying on a redirect DeepSeek describes as temporary.

MiniMax-M3.1-Flash-Preview takes text, image and video natively, and it is the vendor's current top-of-table text model. Video input at a Flash price point is not common in this bracket.

Both caveats point the same way for anyone running production traffic: the identifier on the invoice and the model that answered are not guaranteed to be the same thing indefinitely, and a routing layer is where that gets handled rather than discovered.

A generated infographic titled 'The two Flash brackets, side by side' listing six paired rows: 'Sales motion — subscription seat vs per-token rate card', 'Effective input rate — Token Plan quota vs $0.15 off-peak / $0.30 peak', 'Peak surcharge — not applicable vs doubles 01:00-04:00 and 06:00-10:00 UTC', 'Cost attribution — pooled quota vs metered per call', 'Failure mode — no published output ceiling vs identifier now serves a successor', and 'Best fit — fixed-budget teams vs metered pipelines'. A footer reads 'Both models carry a 1M-token context window; neither figure is a quality claim.'

How to decide this in one afternoon

Start from the ending of the task rather than the beginning.

If the job is long-form generation — anything that finishes by writing tens of thousands of tokens — DeepSeek V4 Flash is the only one of the two that publishes a ceiling big enough to promise it, and its rate card is small enough that the peak multiplier is survivable. Model the cost at peak, not off-peak, and treat the retired identifier as a migration you have not done yet rather than a stable contract.

If the job is reading and reasoning over multimodal input under a fixed monthly budget — screen recordings, scanned documents, video review — MiniMax-M3.1-Flash-Preview is the more capable shape, and inside a Token Plan seat it is the cheapest way to get video input at this context length. Accept that you are buying an unmeasured model, and cap the blast radius: keep the workload that matters on something with a published rate card, and let the preview handle the exploratory half.

If both descriptions fit your pipeline, that is a routing problem rather than a model-selection problem, and it is the case we exist for. DeepSeek V4 Flash on OrcaRouter sits at the provider's rate with 0% markup — its catalogue entry shows $0.22 per million input tokens and $0.66 per million output, which is the peak column of the current Flash card, passed straight through — alongside MiniMax-M3 and MiniMax M2.7 in the same price band on the same key. One endpoint reaches 200+ models with automatic failover behind it, so a workload can move between price brackets without a second integration. MiniMax-M3.1-Flash-Preview is not on our catalogue; reaching it means the vendor's Token Plan and its coding product.

The one-line version: DeepSeek V4 Flash is the known quantity with the published output ceiling, the readable rate card and a successor quietly answering to its name; MiniMax-M3.1-Flash-Preview is the newer, multimodal, unmeasured one that costs a subscription to try. Neither of those is a reason to pick the other — they are just the facts you were not going to get from a price-per-token comparison that quotes a card DeepSeek retired in September.

A headless Chromium screenshot of the OrcaRouter model page for deepseek/deepseek-v4-flash, showing the model title 'DeepSeek: DeepSeek V4 Flash', a pricing row reading 0.22 US dollars per million input tokens and 0.66 per million output tokens, a context window of 1,048,576 tokens, and a maximum output of 384,000 tokens, captured 28 September 2026.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily