
Dots vs Gemini 3.1 Pro: An Agent With a Desktop Against a Model With Five Senses
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 383 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 209 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 680 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 50 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 102 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 219 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
The word "multimodal" hides the real difference between these two, so start with what each one does with the world. Gemini 3.1 Pro Preview is Google's flagship reasoning model, released in preview on February 19, 2026: it accepts text, images, audio, video and files as input and returns text, works across a 1,048,576-token context window with up to 65,536 tokens of output, and costs $2.00 per million input tokens and $12.00 per million output tokens below a 200,000-token prompt, repricing to $4.00 and $18.00 above it. Dots is the company's always-on agent, announced at DevDay on September 29, 2026: an agent with its own cloud computer and browser that works across more than 4,000 connected apps, runs on GPT-6 Astra, and comes included with Pro and Business Premium. Gemini 3.1 Pro watches the world and describes it; a dot acts inside it. Those are different verbs, and almost every practical question about this pair follows from which verb your problem needs.
Input is not the same as action
Gemini 3.1 Pro's media support is genuinely broad — a two-hour recording, a product photo, a spreadsheet export and a contract can all arrive in the same request — but every one of those inputs is something that already existed before the call. The model's output is text: an answer, an extraction, a plan, a draft. Nothing outside the conversation changes because the model ran.
A dot inverts that. Its inputs are mundane — your messages, your app data, a task you described — and its output is a change in the world: a file moved, a ticket updated, a message sent after your approval, a page opened in its own browser. OpenAI's launch materials describe it carrying several projects forward at once, learning from feedback, and accepting new tasks without a new thread. That is a description of a worker, not of a model, and it explains why OpenAI published no benchmark for it: you do not benchmark a colleague on a leaderboard, you watch what it gets done and check the Activity View.
• What it takes in — messages, task descriptions and connected-app data on one side; text, image, audio, video and file on the other
• What it produces — changes inside other software, subject to rules and approval gates, against text you then handle yourself
• Where it runs — OpenAI's cloud computer with its own browser, isolated from your machine unless you link them, against Google's serving infrastructure, reachable through an API from anywhere you like
• Context — undocumented for a dot, against 1,048,576 tokens in and 65,536 tokens out
• Evidence — vendor demos, no independent benchmark, against an Artificial Analysis Intelligence Index of 29.7 with 94.1 on GPQA Diamond and 82 on long-context recall, listed independently of Google
What Gemini 3.1 Pro is worth having
The reason to keep a model like this around is that perception is hard and most agents are bad at it. Gemini 3.1 Pro's published independent profile is unusual: 94.1 on GPQA Diamond is a near-frontier result, 82 on long-context recall is the second-best figure in the set of models we list, and 58.7 on SciCode puts it at the top of that particular board. Where it does not shine is agentic decision-making — a τ²-Bench score of 95.6 is respectable but its τ-banking slice, the one that measures multi-step tool use against a bank's systems, comes in at 21.4. All of those are third-party measurements listed against their source in our catalogue, not vendor figures, which is the reason they are worth more than a launch deck.
The other half of that per-request story is that the model is not free. It entered preview in February, months before the current frontier shipped, and it is priced above the cheap tier: $2.00 and $12.00 per million tokens is roughly $0.0006 of input per page of dense text, which is nothing until you run a video ingest over a library and then it is a line item. Reading a frame of video costs the same as reading a paragraph, and the model will happily bill you for both.

Two cost models that refuse to convert
A dot's cost is a subscription with an unpublished allowance: the first one is included with Pro or Business Premium, extended limits apply in the first month, conversations with your dot do not draw down ChatGPT usage limits, and there is no published price for a second dot, for extra speed or capacity, or for a finished task. CNBC's account of the keynote has Sarah Friar pricing the new $500 Pro 500 tier as a usage allowance plus the Ultrafast mode rather than as a per-dot rate, so the most expensive plan on OpenAI's sheet does not produce a cost per unit of agent work either.
Gemini 3.1 Pro converts everything into tokens, including the parts that are not text. Audio, video and images are billed as input, so the cost of a task depends on how much of the world you made the model look at — a useful property, because you can measure it, and a dangerous one, because the biggest number in the bill is frequently the one nobody planned for. Model routing helps here in a way that is specific rather than generic: with 200-plus models behind a single key at provider list price with no markup, the expensive multimodal read and the cheap follow-up work can sit on different models in the same pipeline, and a rate change from Google lands on your account the same day instead of at renewal.

Where the two work together
The natural pairing is a dot that coordinates and a multimodal model that perceives. Work arrives, a dot routes it: anything that needs to be looked at or listened to goes to Gemini 3.1 Pro for a structured description, and the description goes back into the workflow where a schedule or a rule can act on it. The dot handles the chasing; the model handles the reading. Splitting it that way also keeps the expensive modality calls auditable, which matters because they are the ones that surprise people.
On OrcaRouter the model is routed as google/gemini-3.1-pro-preview at provider list price passed through with 0% markup, alongside the rest of the catalogue, with automatic failover across upstream providers when one degrades and a routing DSL for composing several models into a single call. It is worth being precise about what that does not include: dots are not on our catalogue, there is no endpoint for one, and no model identifier exists to point a request at. If you want the perception layer under your own control — and a preview model from February is exactly the kind of dependency worth containing — the model belongs behind your endpoint and the agent stays wherever you rented it.
Which verb does your problem need
If the work involves looking at or listening to something and producing an answer, Gemini 3.1 Pro is the tool and a dot is the wrong purchase. If the work involves getting something finished across several applications without you driving, a dot is the tool and no amount of multimodal input substitutes for it. The teams that get this right run both and keep the boundary explicit: perception on a metered model you can measure, action behind a product you can cancel.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
