A generated hero card for GPT-6.1 Sol vs DeepSeek V4 Pro headed 'Three meters, three different answers', with three metric cards reading 'Blended price 4x', 'Cost per index task 1.07x' and 'Intelligence Index 16 points', a line reading 'One is proprietary and sees images. One is MIT-licensed open weights and text-only.', a footer reading 'Prices per OpenAI and DeepSeek; index and cost-per-task figures per Artificial Analysis, which compares open-weights and proprietary models in different peer groups.', and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

GPT-6.1 Sol vs DeepSeek V4 Pro: Three Meters, Three Different Answers

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Ask how far apart GPT-6.1 Sol and DeepSeek V4 Pro are and the honest answer is that it depends which meter you read, and the three meters disagree by more than most model comparisons disagree about everything. On price per million tokens, blended at a 3:1 input-to-output ratio, they are $4.00 against $0.99 — a factor of four. On what it actually cost to run the same evaluation suite, they are $0.72 against $0.67 per task — seven percent. On the Artificial Analysis Intelligence Index they are 51.8 against 36 — sixteen points, which sounds decisive until you look at who each number was compared against, because that evaluator's own methodology places the two models in different leagues. GPT-6.1 Sol was released September 29, 2026 at $2.00 in and $10.00 out with a 1,050,000-token window. DeepSeek V4 Pro is an MIT-licensed mixture-of-experts flagship, 1.6 trillion parameters with 49 billion active, released August 13, 2026 at $0.66 and $1.98 with a 1,048,576-token window. One is a text model you can download. The other sees images. Sorting out which of the three meters you should actually steer by is the whole job.

The index row is not the comparison it looks like

Start with the number that gets quoted most, because it is the one most often quoted wrong.

Artificial Analysis publishes its methodology alongside its leaderboard, and two sentences in it matter here. Open-weights models, it says, are compared only with other open-weights models of the same size class — tiny, small, medium, large, with the boundary at 150 billion parameters. Proprietary models are compared across a price band instead, at a blended 3:1 input-to-output ratio. Those are two different peer groups. DeepSeek V4 Pro 0813 is open weights — the evaluator's own FAQ says so, and points at the published weights — and at 1.6 trillion total parameters it lands in the open-weights large class. GPT-6.1 Sol is proprietary and is scored against models in its own price band.

A 51.8 and a 36 sitting in adjacent rows of the same table are therefore not a head-to-head. They are a score against one peer set and a score against another, which is precisely why the per-task cost figures are comparable and the raw index numbers are not. If you want to know whether DeepSeek V4 Pro is good for a downloadable model, 36 is the number that answers it. If you want to know how it does against a hosted frontier model, the index column will not tell you, and this is a case where the honest move is to say so rather than subtract 36 from 51.8 and call it a gap.

What can be compared cleanly: the price cards, the context and output limits, the modality lists, and the cost of running the same evaluation suite — because that last figure is an accounting of identical work, not a score.

What the raw meters say

A screenshot of the OrcaRouter model page for deepseek/deepseek-v4-pro, credited to DeepSeek and dated 2026-04-24, showing the model identifier, text-only input with tools, JSON and reasoning, a 1M-token context window with a 384K maximum output, and a rate row reading $0.66 input and $1.98 output per million tokens alongside a 1.92-second median latency and a 1030.7M seven-day traffic figure.

• Input price — GPT-6.1 Sol $2.00 vs DeepSeek V4 Pro $0.66 per million tokens

• Output price — GPT-6.1 Sol $10.00 vs DeepSeek V4 Pro $1.98, stepping to $1.32 and $3.96 inside two four-hour UTC windows on weekdays

• Cache read — GPT-6.1 Sol $0.10 vs DeepSeek V4 Pro $0.022 per million tokens; both cache, and the cheap side is 4.5× cheaper again

• Cost per index task — GPT-6.1 Sol $0.72 vs DeepSeek V4 Pro $0.67

• Output tokens emitted on the index suite — GPT-6.1 Sol 67M vs DeepSeek V4 Pro 160M

• Maximum output — DeepSeek V4 Pro 384,000 tokens vs GPT-6.1 Sol 128,000 tokens

• Modalities — GPT-6.1 Sol takes text, images and files; DeepSeek V4 Pro takes text only

The fourth and fifth lines are the interesting pair. DeepSeek V4 Pro emits 2.4 times as many tokens on the same suite and still finishes 7% cheaper per task, because its output rate is a fifth of GPT-6.1 Sol's. That is what a price-driven model looks like when you measure output rather than rate: the verbosity is real, and it does not erase the advantage. The timing window is the asterisk. Two four-hour blocks on weekdays — 40 of the 168 hours in a week — double the listed rate to $1.32 and $3.96, and that surcharge applies to the same tokens that would otherwise be the cheapest on either card.

What the cheaper model does not do

Three things, and each one is a hard limit rather than a trade-off.

It does not see. DeepSeek V4 Pro is text-in, text-out. No screenshots, no scanned PDFs, no chart reading, no diagram-in-a-ticket. Computer-use and document-heavy workflows — the exact workloads GPT-6.1 Sol is positioned for — are off the table at any price. If your pipeline already routes images to a vision model, this is a non-issue; if it does not, it is the deciding constraint.

It does not have a release cadence you can plan around. The name "deepseek-v4-pro" has resolved to the 0813 build since August, and getting there was not quiet: the vendor announced the build's retirement for September 14, withdrew the announcement two days before the date, and kept serving the API with billing unchanged. Our own catalogue still lists it as the flagship with the same context, output ceiling and concurrency cap. A vendor that can withdraw a deprecation in two days can also decline to; the operational takeaway is not that the model is unstable but that the name is a pointer, and pinning behaviour to a build identifier rather than a route name is the cheap insurance.

It does not carry the surrounding knowledge. GPT-6.1 Sol is eight days old and already has one independent configuration published; DeepSeek V4 Pro has a longer evaluation history and a much larger practitioner base, including a published serving stack and third-party throughput work. For a self-hosted deployment that body of material is worth more than the index row, because it is what you consult when the model behaves strangely on your hardware rather than on a benchmark.

When the four-times-cheaper meter is the wrong meter

If the cheap model were simply worse on everything, this would be a short article. The awkward fact is that the two are 7% apart per task on identical work, and 2.4× apart in how much they write to get there. That means the blended-token meter — the one that says 4× — is measuring a difference you mostly do not pay, because it prices a token mix you may never send. And the index meter, the one that says 16 points, is measuring across two peer groups and cannot be subtracted.

So the practical split is not "frontier model for hard work, cheap model for easy work." It is a modality split and a volume split. Anything with an image in it goes to GPT-6.1 Sol, and so does anything where the answer is short and the reasoning is the expensive part — its five effort levels, low through max with medium as the default, let you buy reasoning without buying prose, and its cached input at $0.10 makes a long repeated prompt cheap. Anything text-only, high-volume and tolerant of a longer answer — classification, extraction, summarisation, bulk generation against a 384,000-token output budget — goes to DeepSeek V4 Pro, ideally outside the two UTC windows.

Both on one key, and the split is a config line

Both models are on OrcaRouter at their providers' own published rates, nothing added: GPT-6.1 Sol at $2.00 / $10.00, stepping to $4.00 / $15.00 above 272,000 prompt tokens, and DeepSeek V4 Pro at $0.66 / $1.98 with the two timed windows carried through as listed. One key, one bill, one OpenAI-compatible endpoint for both.

That is the reason the split above is worth writing down rather than arguing about. A modality split is expressible as a routing rule, so the image-bearing requests go to the vision model and the bulk text goes to the cheap one without an application change; if a route degrades, failover moves the call instead of failing it. And because both are on the same endpoint, the price comparison that actually matters — yours, on your traffic, at your token mix — can be produced from your own invoices rather than from a benchmark suite's average.

A screenshot of OpenAI's developer model page for GPT-6.1 Sol, headed 'GPT-6.1 Sol' with the tagline 'Near-Astra performance for complex work at a lower cost', showing a 1,050,000-token context window, 128,000 max output tokens, an Apr 30, 2026 knowledge cutoff, a price row reading $2 input and $10 output per million tokens, and the note that reasoning.effort supports low, medium (default), high, xhigh and max while none and minimal are not supported.A generated scoreboard titled 'GPT-6.1 Sol vs DeepSeek V4 Pro — the scoreboard' showing two columns on six shared rows: input price $2.00 against $0.66 per million tokens; output price $10.00 against $1.98; cache read $0.10 against $0.022; cost per index task $0.72 against $0.67; maximum output 128,000 against 384,000 tokens; modalities text, image and file against text only. A footer reads 'Prices per OpenAI and DeepSeek as listed on OrcaRouter; cost-per-task figures per Artificial Analysis, whose index compares open-weights and proprietary models in different peer groups.'

The verdict, in one line, is that the four-times-cheaper reading is the one to distrust and the seven-percent reading is the one to check. Four times is what the blended meter says about a token mix you may not send. Seven percent is what the same evaluation cost on both, and it comes with a 2.4× verbosity difference baked in. The sixteen points are real, they are also across two peer groups, and the constraint that actually decides most deployments of this pair is not a number at all: one of these models can read a screenshot and the other cannot.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily