
Dots vs MiniMax M3: Renting the Worker or Downloading the Reader
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 349 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 210 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 680 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 49 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 102 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 219 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
The cheapest model in this batch is also the one you are allowed to keep, and that combination is the whole comparison. MiniMax M3, published May 31, 2026, is described as MiniMax's flagship open-weight foundation model: text, image and video in with text out, a 1,048,576-token context window, up to 512,000 tokens of output in a single response, and a list price of $0.30 per million input tokens and $1.20 per million output tokens with cached reads at $0.06 — roughly a thirteenth of what the frontier charges for the same tokens. Dots is the company's always-on agent, announced at DevDay on September 29, 2026: it gets its own cloud computer and browser, works across more than 4,000 connected apps, runs on GPT-6 Astra, and comes included with Pro and Business Premium with no allowance figure published anywhere. One is a component you can hold. The other is a service you can only subscribe to.
The price is not the interesting number
Thirty cents a million input tokens is the headline, but the interesting figures are the two that shape what you can ask for. The first is 512,000 tokens of maximum output — six times what GPT-5.6 Sol will emit in one response, and the largest output ceiling of any model in this comparison set. The second is 83 on long-context recall, which is high enough that the context window is a usable working surface rather than a number on a specification sheet.
Put those together and a specific class of task becomes cheap: generate the long thing from the long source. A 400-page technical manual rewritten for a different audience. A migration guide produced from an entire repository. A research summary that is itself the length of a report. At $0.30 and $1.20 per million, a job that would be a four-figure line item on a frontier model is a rounding error here — and the fact that the output ceiling is measured in hundreds of thousands of tokens rather than tens of thousands is what makes it one request instead of twenty.
There is a caveat worth stating about measurement. Our catalogue's latency panel reports a 10.00-second median time to first token for M3 over the same seven-day window that shows 18.35 million tokens of traffic, and that figure sits exactly on the top of the panel's displayed range. Read it as an artefact of the window rather than a characterisation of the model. The traffic figure is the one that tells you people are actually using it.

Video is the input nobody plans for
Vision is now table stakes, so the modality that still differentiates is video. M3 accepts it natively, alongside text and images, and that changes the shape of a pipeline rather than just adding a capability to a list. A support archive of screen recordings, a set of dashcam clips, a semester of lecture captures — these were previously a preprocessing problem, where you sampled frames with one tool and described them with another. A model that reads video directly collapses two stages into one call.
It also makes the cost arithmetic less predictable, because video is expensive to look at in a way text is not, and the biggest number in a bill is usually the input nobody budgeted for. A cheap per-token rate is what makes that safe to experiment with: at $0.30 per million input tokens, an occasional five-hour ingest is affordable enough to be prototyped rather than argued about in a planning meeting.
What the open weights actually change
MiniMax describes M3 as an open-weight model, which means the weights are published rather than only served, and that is a different kind of guarantee from a subscription. If MiniMax stopped serving tomorrow, or repriced, or deprecated the identifier, the model would not disappear — it would keep running on hardware you had already bought, at whatever quantisation fitted. That does not make self-hosting cheaper in the general case; serving a large sparse mixture-of-experts yourself is a real engineering project. It makes the dependency survivable, which is a different claim and the one that matters when you are deciding what to build on.
A dot offers the opposite guarantee. OpenAI holds the computer, the browser, the connectors and the allowance, and your remedy if any of it changes is to stop using it. That is a fair arrangement — renting is how most teams should start — but it is not the same arrangement, and the difference only becomes visible on the day it matters.

Six differences that change a decision
• Guarantee — a service that can change terms, against weights you can keep (MiniMax M3 is described as open-weight; a dot is not downloadable in any sense)
• Cost per unit of work — unpublished for a dot, against $0.30 and $1.20 per million tokens with cached reads at $0.06, the lowest rate in this comparison set
• Single-request ceiling — undocumented for a dot, against 1,048,576 tokens of input and 512,000 tokens of output
• Modalities in — messages and connected-app data for a dot, against text, image and video for the model
• Modalities out — changes inside other software, against text and only text
• Independent standing — no benchmark exists for a dot; M3 carries an Artificial Analysis Intelligence Index of 29.2 placing it sixtieth of 145 models measured, which is mid-field and worth knowing before you assume the cheap rate buys frontier reasoning. It does not.
Where each one earns its line on the invoice
The honest read of M3's published profile is that it is a volume model, not a frontier one: 92.9 on GPQA Diamond is competitive, 59 on SWE Bench Pro is respectable, and the mid-field Intelligence Index is the tell. What it wins on is cost per unit of work, output ceiling, video and deployability, in that order. That is a good model to put underneath the expensive layer of a pipeline — the summarise, extract, transcribe, generate-long-text and fan-out-many-attempts steps — and a bad one to put at the top of it.
OrcaRouter serves it as minimax/minimax-m3 at MiniMax's provider list price with 0% markup, in the same catalogue as 200-plus models behind a single key, with automatic failover across upstream providers and a routing DSL for sending an individual request to the model that should handle it. That arrangement is what makes the two-layer split practical rather than theoretical: the same key that reaches a cheap model for volume reaches a frontier model for the calls that must be right, and nothing arrives inside a dot because dots are not on the catalogue at all — no endpoint, no identifier, no substitute intelligence.
What to do on Monday
If you have a corpus of audio or video and no cheap way to look at it, start there: M3 is the only model in this set that reads video natively at this price, and the experiment costs almost nothing. If you have a class of work that keeps getting dropped because nobody is chasing it, the dot is the product shaped for that, and the metered model will not pretend to be it.
The two only conflict if you assume one must win. The subscription chases; the downloaded reader reads. Own the second, rent the first, and keep the boundary between them explicit enough that you could replace either one without rewriting the other.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
