
Dots vs Grok 4.6: The Agent That Books Its Own Meetings Against the Model That Answers
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 349 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 210 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 680 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 49 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 102 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 219 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Try to write a phone number for a dot and you will fail, and that failure is the most useful thing about this pairing. Dots is the company's always-on agent, announced at DevDay on September 29, 2026: each dot gets its own cloud computer and browser, connects to more than 4,000 apps, and is reachable inside ChatGPT, Slack and Teams, running on GPT-6 Astra and included with Pro and Business Premium at no separate charge. Grok 4.6 is SpaceXAI's frontier model, published on August 12, 2026, and it is a completely different kind of object: a 500,000-token context window, text, image and file input with text output, configurable reasoning effort, native tool calling and structured outputs, priced at $2.00 per million input tokens and $6.00 per million output below a 200,000-token prompt and double that above it. One of these is a colleague with a schedule. The other is a function you call.
Both are genuinely strong at agentic work, which is what makes the confusion common. Grok 4.6 posts 88.4 on Terminal-Bench 2.1 and 50.7 on the τ-banking tool-use slice — the kind of numbers that get a model described as "an agent". It is not one. It is the reasoning engine you would build one out of.
What each one actually is
• Identity — a hosted product with no model identifier you can address (Dots) against a named, routable model with a stable string you can put in a config file (Grok 4.6)
• Where execution happens — on OpenAI's cloud computer, with its own browser, isolated from your machine unless you link them (Dots) against wherever you run the call: a container you control, a queue worker, a cron job (Grok 4.6)
• Reach — 4,000+ apps through OpenAI's connectors (Dots) against whatever tools you wire up yourself, because tool calling is a protocol and not a directory (Grok 4.6)
• Memory of a task — carried by the product across days, across apps, across new threads (Dots) against a 500,000-token window and whatever you persist yourself (Grok 4.6)
• Cost — included with a subscription, allowance unpublished (Dots) against $2.00 and $6.00 per million tokens, doubling past 200,000 input tokens, with cached reads at $0.50 (Grok 4.6)
• Evidence — vendor demonstration and capability statements, no independent benchmark (Dots) against an Artificial Analysis Coding Index of 76.8, fifth of 138 models measured, and 94.9 on GPQA Diamond (Grok 4.6)
The number that decides most questions: the allowance
OpenAI published one price for dots and it is not a price in any unit you can measure work in. The first dot is included with Pro or Business Premium; conversations with it do not draw down ChatGPT usage limits; extended limits apply in the first month. That is the whole published cost structure. No allowance figure, no second-dot price, no per-task rate, no enterprise price for the specialist dots reported to be in internal testing. CNBC's account of the keynote has finance chief Sarah Friar pricing the $500-per-month Pro 500 tier as a usage allowance plus the new Ultrafast mode rather than as a per-dot rate, which means the most expensive plan on OpenAI's sheet still does not yield a cost per unit of agent work.
Grok 4.6 goes the other way and publishes everything, including the part that hurts. The tier boundary at 200,000 input tokens is a cliff: a request at 199,000 tokens bills at half the rate of one at 201,000. That is not a hidden trap, it is a number in the rate card, and it changes how you design a long-context job. You chunk at the boundary or you accept the second tier knowingly.

A successor shipped while this pair was being compared
Grok 4.7 was published on September 21, 2026 and is now the flagship of the Grok line, with the same 500,000-token context, the same input surface, the same base pricing and a much larger 450,000-token maximum output. Grok 4.6 remains available and still listed as a current model — a version bump is not a retirement — but anyone choosing a Grok model today should be choosing between 4.6 and 4.7 rather than assuming 4.6 is the frontier. The reason to say so plainly is that the article you are reading is about a comparison that survives the transition: whatever the top of the Grok line is called next quarter, it is still a metered model with a rate card, and a dot is still a subscription with an unpublished allowance.
That is also the argument for routing your model calls through an abstraction rather than a hard-coded vendor. When the successor lands, changing the model is a one-line edit rather than a migration.

Building the agent instead of renting it
The interesting consequence of this pairing is that Grok 4.6 lets you build the thing Dots sells, badly and then better. A loop that reads a queue, calls the model with tools attached, writes results and retries on failure is not hard; what is hard is the parts a dot gives you free — a browser that behaves like a person's, connectors that were written by the app vendors, an approval flow and an activity log. Those are worth a subscription precisely because they are tedious.
Where the metered side wins is the moment the work stops being generic. The moment you need the same call to run against a different vendor next quarter, to be replayed from a log, to be priced to the cent, or to fail over to a cheaper sibling mid-flight when an upstream provider wobbles. On OrcaRouter, Grok 4.6 is served as grok/grok-4.6 at SpaceXAI's list price with 0% markup on top — the provider's own rate, passed through, so a vendor price cut lands on your bill the same day — in the same catalogue as 200-plus other models behind one key, with automatic failover and a routing DSL for composing several models into a single call. There is no equivalent for a dot: no endpoint, no identifier, nothing to fail over from.
The test to apply
Ask whether the work would still exist if nobody looked at it. If the answer is yes — it has to happen on a schedule, across several tools, without a person opening a chat window — a dot is the product shaped for it, and the metered model is a component inside whatever you build. If the answer is no, if someone is waiting on the output and the output has a format and a deadline, then a metered model is the right purchase and an agent is an expensive detour.
The mistake worth avoiding is buying a dot to get a better answer. A dot is not a smarter model; it is a longer leash on one. And the second mistake is paying for a subscription to get a capability that a rate card already sells you by the token, with the arithmetic visible in advance.

Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
