A generated hero card titled 'Dots vs Grok 4.6', subtitled 'The agent that books its own meetings against the model that answers'. A vertical rule splits the card: the left panel, labelled 'Dots', reads 'own computer, browser, 4,000+ apps'; the right, labelled 'Grok 4.6', reads '500K context - 88.4 Terminal-Bench - $2.00 / $6.00'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Dots vs Grok 4.6: The Agent That Books Its Own Meetings Against the Model That Answers

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Try to write a phone number for a dot and you will fail, and that failure is the most useful thing about this pairing. Dots is the company's always-on agent, announced at DevDay on September 29, 2026: each dot gets its own cloud computer and browser, connects to more than 4,000 apps, and is reachable inside ChatGPT, Slack and Teams, running on GPT-6 Astra and included with Pro and Business Premium at no separate charge. Grok 4.6 is SpaceXAI's frontier model, published on August 12, 2026, and it is a completely different kind of object: a 500,000-token context window, text, image and file input with text output, configurable reasoning effort, native tool calling and structured outputs, priced at $2.00 per million input tokens and $6.00 per million output below a 200,000-token prompt and double that above it. One of these is a colleague with a schedule. The other is a function you call.

Both are genuinely strong at agentic work, which is what makes the confusion common. Grok 4.6 posts 88.4 on Terminal-Bench 2.1 and 50.7 on the τ-banking tool-use slice — the kind of numbers that get a model described as "an agent". It is not one. It is the reasoning engine you would build one out of.

What each one actually is

• Identity — a hosted product with no model identifier you can address (Dots) against a named, routable model with a stable string you can put in a config file (Grok 4.6)

• Where execution happens — on OpenAI's cloud computer, with its own browser, isolated from your machine unless you link them (Dots) against wherever you run the call: a container you control, a queue worker, a cron job (Grok 4.6)

• Reach — 4,000+ apps through OpenAI's connectors (Dots) against whatever tools you wire up yourself, because tool calling is a protocol and not a directory (Grok 4.6)

• Memory of a task — carried by the product across days, across apps, across new threads (Dots) against a 500,000-token window and whatever you persist yourself (Grok 4.6)

• Cost — included with a subscription, allowance unpublished (Dots) against $2.00 and $6.00 per million tokens, doubling past 200,000 input tokens, with cached reads at $0.50 (Grok 4.6)

• Evidence — vendor demonstration and capability statements, no independent benchmark (Dots) against an Artificial Analysis Coding Index of 76.8, fifth of 138 models measured, and 94.9 on GPQA Diamond (Grok 4.6)

The number that decides most questions: the allowance

OpenAI published one price for dots and it is not a price in any unit you can measure work in. The first dot is included with Pro or Business Premium; conversations with it do not draw down ChatGPT usage limits; extended limits apply in the first month. That is the whole published cost structure. No allowance figure, no second-dot price, no per-task rate, no enterprise price for the specialist dots reported to be in internal testing. CNBC's account of the keynote has finance chief Sarah Friar pricing the $500-per-month Pro 500 tier as a usage allowance plus the new Ultrafast mode rather than as a per-dot rate, which means the most expensive plan on OpenAI's sheet still does not yield a cost per unit of agent work.

Grok 4.6 goes the other way and publishes everything, including the part that hurts. The tier boundary at 200,000 input tokens is a cliff: a request at 199,000 tokens bills at half the rate of one at 201,000. That is not a hidden trap, it is a number in the rate card, and it changes how you design a long-context job. You chunk at the boundary or you accept the second tier knowingly.

A screenshot of the OrcaRouter model page for Grok 4.6, showing the identifier 'grok/grok-4.6' with Vision, Tools, JSON and Reasoning badges, a publication date of 2026-08-12, a description of it as SpaceXAI's smartest model to date and the current flagship of the Grok line with a 500K-token context window, a side panel reading 500K tokens of context over text, image and file input, and a price strip reading $2.00 and $6.00 with tiered rates above a 200,000-token prompt.

A successor shipped while this pair was being compared

Grok 4.7 was published on September 21, 2026 and is now the flagship of the Grok line, with the same 500,000-token context, the same input surface, the same base pricing and a much larger 450,000-token maximum output. Grok 4.6 remains available and still listed as a current model — a version bump is not a retirement — but anyone choosing a Grok model today should be choosing between 4.6 and 4.7 rather than assuming 4.6 is the frontier. The reason to say so plainly is that the article you are reading is about a comparison that survives the transition: whatever the top of the Grok line is called next quarter, it is still a metered model with a rate card, and a dot is still a subscription with an unpublished allowance.

That is also the argument for routing your model calls through an abstraction rather than a hard-coded vendor. When the successor lands, changing the model is a one-line edit rather than a migration.

A generated scoreboard titled 'Dots vs Grok 4.6 - product against protocol', with two columns. The left column, 'Dots', lists: price included with Pro or Business Premium; allowance not published; computer its own cloud computer and browser; connectors 4,000+ apps; channels ChatGPT, Slack, Teams; independent benchmark none. The right column, 'Grok 4.6', lists: context 500,000 tokens; input text, image, file; price $2.00 in / $6.00 out per 1M; above 200K input $4.00 in / $12.00 out per 1M; cached reads $0.50 per 1M; AA Coding Index 76.8, 5th of 138. A footer reads 'Dots details vendor-stated from OpenAI's DevDay launch; Grok 4.6 figures from independent listings as carried in the OrcaRouter catalogue.'

Building the agent instead of renting it

The interesting consequence of this pairing is that Grok 4.6 lets you build the thing Dots sells, badly and then better. A loop that reads a queue, calls the model with tools attached, writes results and retries on failure is not hard; what is hard is the parts a dot gives you free — a browser that behaves like a person's, connectors that were written by the app vendors, an approval flow and an activity log. Those are worth a subscription precisely because they are tedious.

Where the metered side wins is the moment the work stops being generic. The moment you need the same call to run against a different vendor next quarter, to be replayed from a log, to be priced to the cent, or to fail over to a cheaper sibling mid-flight when an upstream provider wobbles. On OrcaRouter, Grok 4.6 is served as grok/grok-4.6 at SpaceXAI's list price with 0% markup on top — the provider's own rate, passed through, so a vendor price cut lands on your bill the same day — in the same catalogue as 200-plus other models behind one key, with automatic failover and a routing DSL for composing several models into a single call. There is no equivalent for a dot: no endpoint, no identifier, nothing to fail over from.

The test to apply

Ask whether the work would still exist if nobody looked at it. If the answer is yes — it has to happen on a schedule, across several tools, without a person opening a chat window — a dot is the product shaped for it, and the metered model is a component inside whatever you build. If the answer is no, if someone is waiting on the output and the output has a format and a deadline, then a metered model is the right purchase and an agent is an expensive detour.

The mistake worth avoiding is buying a dot to get a better answer. A dot is not a smarter model; it is a longer leash on one. And the second mistake is paying for a subscription to get a capability that a rate card already sells you by the token, with the arithmetic visible in advance.

A screenshot of the OrcaRouter model catalogue page headed 'Models', showing the line '206 models · 16 providers · one API key, one bill' above filter controls for input modalities, context length, input price, status and supported parameters, with model cards visible in a grid beneath.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily