A generated hero card for GPT-6.1 Sol headed 'The independent score arrived', with badges reading 'Released: September 29, 2026', 'Intelligence Index: 52 at max effort' and 'Cost per index task: $0.72', a footer reading 'Independent figures per Artificial Analysis, index v4.3.2, read September 30, 2026', and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

GPT-6.1 Sol's First Independent Score Is In: 0.84 Points Behind Astra, at Under a Quarter of the Task Cost

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Every write-up of GPT-6.1 Sol published in the days after Ope​nAI's September 29, 2026 DevDay release — this blog's included — carried the same caveat: the capability numbers were the vendor's and no third party had scored the model. That caveat is now out of date. Artificial Analysis has a GPT-6.1 Sol row on its Intelligence Index, run at maximum reasoning effort, and it is the first figure in this model's short life that Ope​nAI did not produce. It reads 51.8 against 52.7 for GPT-6 Astra and 47.5 for GPT-6 Sol — a gap of less than one point to the $10/$50 flagship, and a 4.3-point step up from the model GPT-6.1 Sol replaces. The number that makes the row interesting is the one beside it: the board prices the evaluation run behind each score, and GPT-6.1 Sol's index run cost $0.72 per task where Astra's cost $3.26. Four and a half times the money for 0.84 points is the whole comparison in one line, and it is now a measured line rather than a vendor's framing of one.

The row itself, and what it was run on

A screenshot of the Artificial Analysis model page for GPT-6.1 Sol at maximum reasoning effort, headed 'GPT-6.1 Sol (max) Intelligence, Performance & Price Analysis', with a class-rank strip above the summary naming the model's Intelligence, Speed, Cost and Verbosity positions out of 221 models. The summary paragraph reports an Intelligence Index score of 52 against a median of 26 across 221 models in the class, $2.00 per million input tokens and $10.00 per million output tokens, $0.72 per task to evaluate on the Intelligence Index, and 67 tokens per second against an average of 74, with 67 million output tokens generated against a board median of 82 million; a panel to its right lists a 1M-token context window together with text-and-image input. The page names Intelligence Index v4.3.2 and its ten constituent evaluations.

Artificial Analysis publishes its Intelligence Index as a composite of ten evaluations, and the version running on September 30, 2026 was v4.3.2 — AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. All three OpenAI models in this piece sit on that same version, so the comparison below is like-for-like in a way that a vendor's own launch table never is. The board also records the release date it has for the model — 2026-09-29, the DevDay date — and evaluates it in the max-effort configuration, which is the setting the page names in its own heading.

• Intelligence Index — 51.8 for GPT-6.1 Sol (max), 52.7 for GPT-6 Astra (max), 47.5 for GPT-6 Sol (max).
• Cost per index task — $0.72 for GPT-6.1 Sol, $3.26 for Astra, $1.05 for GPT-6 Sol. This is the board's own figure for what one full index evaluation cost to run, not a rate card.
• Output tokens spent on the run — 67.2 million for GPT-6.1 Sol, 60.0 million for Astra, 76.8 million for GPT-6 Sol. AA's own commentary calls 67 million "fairly concise" against a board median of 82 million, which is a second, quieter win over the model it replaces.
• Output speed — 66.8 tokens per second for GPT-6.1 Sol, 57.0 for Astra, 76.2 for GPT-6 Sol.
• Prices as the board lists them — $2.00 input and $10.00 output per million tokens for GPT-6.1 Sol, with cached input at $0.10 and cache writes at $2.50. Those four numbers match the vendor's pricing page exactly, which is the least dramatic and most useful thing about the row: an independent list of the rates we already quoted.

One framing note before the numbers get used as ammunition. AA's index is ten evaluations, and the evaluations OpenAI led its own announcement with — DeepSWE v1.1, OSWorld 2.0, Terminal-Bench Science 0.1 — are not among them. AA runs its own AutomationBench variant, which is not the same harness as the AutomationBench version the vendor cited. So the independent record does not confirm or refute the four headline claims from the launch post; it measures a largely different set of tasks, and it is worth reading as a second opinion on the shape of the model rather than as a verdict on those specific sentences.

Astra still wins, by less than a point, for four and a half times the money

A generated scoreboard titled 'GPT-6.1 Sol - the independent scoreboard', listing six rows: Intelligence Index 51.8 at maximum effort, cost per index task $0.72 against GPT-6 Astra's $3.26, output tokens on the run 67.2 million, output speed 66.8 tokens per second, price $2.00 input and $10.00 output per million tokens with cached input at $0.10, and release date September 29, 2026. A footer reads that all figures are per Artificial Analysis, Intelligence Index v4.3.2, read September 30, 2026.

Both models expose an effort ladder, and the board has now published a score and a run cost for every rung of both. Lined up, the ladders say something the single headline comparison cannot: GPT-6.1 Sol's best configuration is not competing with Astra's best. It is competing with Astra's second-best.

• GPT-6.1 Sol, max — 51.8, $0.72 per index task. GPT-6 Astra, max — 52.7, $3.26.
• GPT-6.1 Sol, xhigh — 51.0, $0.39. GPT-6 Astra, xhigh — 52.4, $2.31.
• GPT-6.1 Sol, high — 50.2, $0.32. GPT-6 Astra, high — 50.9, $1.73.
• GPT-6.1 Sol, medium — 47.8, $0.21. GPT-6 Astra, medium — 49.6, $1.54.
• GPT-6.1 Sol, low — 42.1, $0.13. GPT-6 Astra, low — 45.8, $0.82.

Astra leads at every rung, which is the honest summary of the matchup and also the least interesting part of it. The interesting part is the exchange rate. If you want the cheapest Astra configuration that beats GPT-6.1 Sol's best, it is Astra at xhigh: 52.4 against 51.8, for $2.31 against $0.72. That is 0.56 of an index point for $1.59 more per evaluation run — roughly 3.2 times the cost for less than one percent of the score. Move the other way and take GPT-6.1 Sol down to xhigh, and you give up 0.8 points while the bill falls to $0.39, which is 5.9 times cheaper than Astra's xhigh and one-ninth of Astra's maximum-effort run.

There is a second reading of the same ladder that argues against buying Astra at max effort. Astra's four lower rungs are all cheaper than its max, and its xhigh is 0.29 points below its max — little more than half a percent of score for a 29% cut in run cost. Any procurement conversation about this pair that starts from "Astra at max" and ends at "Astra costs 4.5 times more" is leaving Astra's own cheaper rungs out of the comparison.

Six evaluations go to Astra, two go to GPT-6.1 Sol, two tie

The composite hides a split. Scored eval by eval at maximum effort, GPT-6.1 Sol is not a uniformly slightly-worse Astra — it is better at two things and worse at six.

• GDPval-AA — Elo 1575.1 for GPT-6.1 Sol against 1541.9 for Astra, a 33-point lead on professional-work tasks. This is the largest single independent win for the cheaper model.
• AA-LCR v1.1, long-context reasoning — 0.830 against 0.807. GPT-6.1 Sol leads.
• AA-Briefcase v1.1 — Elo 1564.2 against 1568.9. Astra by 4.7.
• AutomationBench-AA — 0.649 against 0.685. Astra by 3.6 points.
• Terminal-Bench 4.0 — 0.561 against 0.591. Astra by 3.0 points.
• SciCode — 0.542 against 0.565. Astra by 2.3 points.
• Humanity's Last Exam — 0.529 against 0.547. Astra by 1.8 points.
• AA-Omniscience — 41.5 against 43.4. Astra by 1.9 points.
• GDP.pdf and CritPt — dead ties at 0.310 and 0.317 respectively, on both models.

Two of those rows are worth more than their position in the list. The GDPval-AA margin is the only place where the cheaper model's advantage is large enough to survive a change of harness, and it lands on exactly the workload the vendor's positioning line claims — professional, document-heavy work. And the two dead ties are on the evaluations with the narrowest spreads, which is a reminder that a 0.84-point composite gap is being assembled from rows that individually range from a 33-point win to a 3.6-point loss.

Against the model it replaces, the bump is real but not uniform

The comparison that matters to anyone already running GPT-6 Sol on a production path is the one 6.1 Sol makes against its own predecessor, and the independent record is sharper about that than the vendor's post was. Eight of ten evaluations improve. Two go backwards.

• SciCode — GPT-6.1 Sol scores 0.542 where GPT-6 Sol scored 0.576. The newer model is worse at scientific code by 3.5 points.
• AA-LCR v1.1 — 0.830 for 6.1 Sol against 0.837 for GPT-6 Sol. A hair's breadth, but in the wrong direction, on the long-context evaluation.
• AA-Omniscience — 41.5 against 27.1. This is the biggest move in the row: a 14.4-point jump on the knowledge evaluation, and accuracy on that set rising from 0.545 to 0.621.
• Hallucination rate on the same set — 0.543 for GPT-6.1 Sol, down from 0.601 for GPT-6 Sol and still above GPT-6 Astra's 0.513. AA-Omniscience splits its score into "got it right" and "confidently wrong", and the split is where the 6.1 refresh earns its keep: it is a meaningfully less dishonest model than Sol, and it is still the most dishonest of the three.
• Terminal-Bench 4.0 — 0.561 against 0.439, a 12.2-point gain. AA-Briefcase — 1482.8 to 1564.2 Elo. GDP.pdf — 0.248 to 0.310.

The reading is that this was a knowledge-and-tooling refresh more than a reasoning refresh. The evaluations that moved most are the ones that reward knowing things and not inventing them; the evaluation that fell is a narrow, technical one. Anyone who moved to 6.1 Sol for the cache saving gets the knowledge gain as a bonus and should not assume it extends to every harness they run.

Where the board ranks it, and what sits above

A generated two-column scoreboard titled 'GPT-6.1 Sol vs GPT-6 Sol — the independent scoreboard', with both columns carrying the same six labels. Left column 'GPT-6.1 Sol (max)': Intelligence Index 51.8, cost per index task $0.72, SciCode 0.542, long-context reasoning 0.830, AA-Omniscience 41.5, hallucination rate 0.543. Right column 'GPT-6 Sol (max)': Intelligence Index 47.5, cost per index task $1.05, SciCode 0.576, long-context reasoning 0.837, AA-Omniscience 27.1, hallucination rate 0.601. A footer reads 'Figures per Artificial Analysis, Intelligence Index v4.3.2, read September 30, 2026.'

In the leaderboard's default Intelligence Index ordering, GPT-6 Astra at maximum effort sits 7th and GPT-6.1 Sol at maximum effort sits 10th, with GPT-6 Sol at maximum effort 20th. The nine rows above GPT-6.1 Sol are two Astra configurations, three Claude Opus 5.5 configurations, one Claude Sonnet 5.5 configuration and two Claude Fable 5.1 configurations — so the model GPT-6.1 Sol is nearest to is the flagship of its own family, and the model it has actually displaced in that ordering is its own predecessor, thirteen places down.

Drop the effort setting and the ordering compresses in a way that matters for cost planning: GPT-6.1 Sol at xhigh ranks 13th, at high 15th, at medium 19th and at low 35th, while Astra sits 14th at high, 16th at medium and 26th at low. In other words, GPT-6.1 Sol at xhigh outranks Astra at high on the board, and GPT-6.1 Sol at medium outranks Astra at low. Effort is a bigger lever here than model choice, at least among these two.

What to do with the number, honestly

The vendor claim this row tested was that GPT-6.1 Sol nearly matches the flagship at roughly one-fifth the price. The independent version of that sentence is: it comes within 0.84 index points of the flagship, and the evaluation run behind its score cost 22% of the flagship's. That is close enough to the claim that nobody should feel misled by it, and it is a stronger statement than "unverified" in every direction that matters to a buyer.

What the row does not do is price your workload. An index run is a specific mix of tasks, and the number that decides a real bill is the rate card against your own token profile — which is why the cache line is still the line to model: GPT-6.1 Sol's $0.10 cached input against Astra's $1.00 is a ten-times difference that no index score touches. On our side, the same pass-through applies: OrcaRouter serves GPT-6 Astra, GPT-6 Sol and GPT-6 Luna at OpenAI's own list prices with no markup added, so the day a vendor rate moves, the rate on our side moves with it rather than waiting on a second contract. GPT-6.1 Sol is not on our routes — the public catalogue returns "model not found" for openai/gpt-6.1-sol today, and there is no honest way to describe a model we cannot serve. What one API key does buy you right now is the pair the benchmark actually argues about — Astra for the hardest work, Sol for the volume — with provider list prices passed through unmarked and automatic failover between providers when one of them errors.

Two things would sharpen this further. A second independent harness on the same model would tell us whether the GDPval-AA lead is a property of the model or of that evaluation, and it is the single most useful missing data point in the row. And an independent measurement of the 272,000-token long-context tier would matter more than another composite point, because that is where this family's pricing stops being cheap.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily