A generated hero card titled 'Dots vs Kimi K3', subtitled 'One reads the whole library, the other goes and finds it'. A vertical rule splits the card: the left panel, labelled 'Dots', reads 'acts across 4,000+ apps - allowance unpublished'; the right, labelled 'Kimi K3', reads '1M context - long-context recall 88.7 - $3.00 / $15.00'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Dots vs Kimi K3: One Reads the Whole Library, the Other Goes and Finds It

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

There is a specific kind of job that only one of these two can do, and it is worth naming before anything else. Kimi K3, published by Moonshot AI on July 15, 2026, is a 2.8-trillion-parameter mixture-of-experts model with a 1,048,576-token context window that accepts text and images and returns text, priced at $3.00 per million input tokens and $15.00 per million output tokens with cached reads at $0.30. Dots is the company's always-on agent, announced at DevDay on September 29, 2026, with its own cloud computer and browser, connectors to more than 4,000 apps, and a place inside ChatGPT, Slack and Teams, running on GPT-6 Astra and included with Pro and Business Premium. K3 can hold an entire archive in one request and answer questions about it. A dot can go and fetch the next document you have not thought of yet. Reading and retrieving are different jobs, and picking the wrong one is the expensive mistake.

The number Kimi K3 owns outright

Long-context recall is the measurement that separates marketing from capability, because a big context window is trivial to claim and hard to use. K3 scores 88.7 on it — the highest figure of any model we list, above GPT-5.6 Sol's 84, MiniMax M3's 83, and Gemini 3.1 Pro's 82. It also posts 93.5 on GPQA Diamond and 59.5 on SciCode, and carries an Artificial Analysis Coding Index of 76.2 placing it ninth of 138 models measured and an Intelligence Index of 43.6 at nineteenth of 145. Those are all third-party figures, carried in our catalogue with their sources attached rather than repeated from a launch post.

What that buys you is a class of task that simply does not decompose: reading a whole codebase before answering a question about it, diffing a year of contracts against a new template, or holding a long agentic session without the middle of it falling out of memory. K3 is built for exactly that — Moonshot positions it for programming agents and long-horizon knowledge work — and it is unusual in refusing sampling parameters entirely. There is no temperature, no top_p, no seed; reasoning depth is a single dial called reasoning_effort. That is a deliberate trade: less knobs, more consistency across a long session.

A screenshot of the OrcaRouter model page for Kimi K3, showing the identifier 'kimi/kimi-k3' with Vision, Tools, JSON and Reasoning badges, a publication date of 2026-07-15, a description of it as Moonshot AI's flagship 2.8-trillion-parameter MoE with a 1M-token context window and native visual understanding, a side panel reading 1M tokens of context over text and image input, and a price strip reading $3.00 and $15.00 with a p50 time to first token of 8.82 seconds and a seven-day traffic figure of 831.3M tokens.

What a dot is for, stated without the marketing

A dot is a worker with a calendar. Its inputs are mundane — your messages, your app data, a task you described in a sentence — and its output is a change somewhere else: a file moved, a ticket updated, a message drafted and sent after your approval, a page opened in its own browser. OpenAI's launch materials describe one carrying several projects forward at once and learning from feedback, and the honest reading is that you are buying the connective tissue, not the intelligence.

The connective tissue is the expensive part to build. A browser that behaves like a person's, connectors maintained by the app vendors, an activity view, approval gates, isolation from your own machine — none of that is hard in principle and all of it is tedious in practice. That is what the subscription pays for. What it does not pay for is measurability.

• Cost — included with Pro or Business Premium, allowance unpublished (Dots) against $3.00 per million input and $15.00 per million output, cached reads $0.30 (Kimi K3)

• Ceiling on one request — undocumented for a dot, against 1,048,576 tokens of context and a mobile-friendly 88.7 on long-context recall (Kimi K3)

• Input and output — messages and connected-app data in, changes inside other software out (Dots) against text and image in, text out (Kimi K3)

• Sampling control — none published (Dots) against no temperature, no top_p and no seed, with reasoning depth set by reasoning_effort alone (Kimi K3)

• Speed you should expect — undocumented for a dot, against an 8.82-second median time to first token and about 39 output tokens per second in our own seven-day measurement (Kimi K3)

• Independent evidence — vendor demos and capability statements, no benchmark (Dots) against AA Coding Index 76.2 at ninth of 138, GPQA Diamond 93.5, long-context recall 88.7 (Kimi K3)

The thing nobody mentions about buying an agent

A subscription with an unpublished allowance has a property that a rate card does not: you cannot tell the difference between a task that was cheap and a task that was expensive. Every job costs the same — nothing marginal — right up until the moment it does not, and then the answer arrives as a limit rather than a bill. For a team whose work is exploratory, that is a feature. For a team that has to attribute cost to a project, it is a hole in the accounting.

There is a second-order effect, too. When the marginal cost of a call is invisible, you stop asking whether the call was necessary, and the cheapest optimisation in any pipeline — not making the request — becomes unavailable to you. A metered model forces that question every time somebody reads the rate card.

A generated scoreboard titled 'Dots vs Kimi K3 - what each one measures', with two columns. The left column, 'Dots', lists: price included with Pro or Business Premium; allowance not published; cost per task not published; model GPT-6 Astra (vendor-stated); computer its own cloud computer and browser; independent benchmark none. The right column, 'Kimi K3', lists: context 1,048,576 tokens; input text and image; price $3.00 in / $15.00 out per 1M; cached reads $0.30 per 1M; architecture 2.8T total MoE; long-context recall 88.7. A footer reads 'Dots details vendor-stated from OpenAI's DevDay launch; Kimi K3 figures from independent listings as carried in the OrcaRouter catalogue.'

Composing them without pretending they are alternatives

The shape that works is a dot that notices and a long-context model that reads. Work arrives — a document in a shared drive, a message in a thread — and the dot decides whether it matters. If it does, the document goes to Kimi K3 with everything else it might need attached, in one request, and the answer comes back fully grounded in the source rather than in a summary of the source. The dot handles the chasing and the boring logistics; the model handles reading, at a volume no loop of small calls would have assembled correctly.

Kimi K3 is routed on OrcaRouter as kimi/kimi-k3 at Moonshot's list price with no markup added, sitting in the same catalogue as 200-plus other models behind one key, with automatic failover across upstream providers when one degrades and a routing DSL for composing several models into a single call. That is the metered half of the pipeline and nothing more: dots are not in the catalogue, there is no endpoint for one, and no identifier exists to point a request at. If the reading layer is the part you need to control — and for anything you intend to run for a year, it usually is — that is the layer to own.

Which purchase wastes money

Buying a dot to read a large corpus is a waste, because the dot will fetch and summarise rather than hold the whole thing, and summarisation is where the detail you needed goes to die. Buying Kimi K3 to chase a project across five applications is a waste for the mirror-image reason: a model that answers when called does not call you.

If the job is retrieval and logistics, rent the worker. If the job is comprehension at scale, buy the meter and give it everything at once. The two overlap far less than the word "agentic" suggests, and the overlap is not where either one is good.

A screenshot of the OrcaRouter model catalogue page headed 'Models', showing the line '206 models · 16 providers · one API key, one bill' above filter controls for input modalities, context length, input price, status and supported parameters, with a grid of model cards visible beneath.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily