
GPT-6.1 Sol Pricing: The Rate Card, the One Meter That Moved, and the 272K Cliff
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 161 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 79 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 53 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 301 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
This page is the rate card. GPT-6.1 Sol costs $2.00 per million input tokens, $0.10 per million cached input tokens, $2.50 per million cache-write tokens and $10.00 per million output tokens on the vendor's standard tier, and GPT-6 Sol — the rung this model replaces, and the one our catalogue serves — reads $2.00, $0.20, $2.50 and $10.00 for the same four meters. One number differs between the two cards, the threshold above which both reprice sits in the same place, and the discounts below the list price are multipliers rather than a second price list. Everything on this page is taken from the vendor's own developer pricing page and its two model pages for GPT-6.1 Sol and GPT-6 Sol, all read on 7 October 2026, and from our own model cards where a figure is labelled as ours. No launch narrative is repeated here; if you want what shipped, that is a different page.
The rate card, with GPT-6 Sol's four meters directly beneath
OpenAI publishes these as prices per million tokens on the standard processing tier, for a request at or under 272,000 input tokens. The two rows that matter are adjacent by design, because the whole argument of the release is visible in the pair: read down each meter in turn and exactly one of them changes.
• GPT-6.1 Sol input — $2.00 per million tokens, unchanged from GPT-6 Sol.
• GPT-6.1 Sol cached input — $0.10 per million tokens, halved from GPT-6 Sol's $0.20.
• GPT-6.1 Sol cache write — $2.50 per million tokens, unchanged from GPT-6 Sol.
• GPT-6.1 Sol output — $10.00 per million tokens, unchanged from GPT-6 Sol.
• GPT-6 Sol input — $2.00 per million tokens.
• GPT-6 Sol cached input — $0.20 per million tokens.
• GPT-6 Sol cache write — $2.50 per million tokens.
• GPT-6 Sol output — $10.00 per million tokens.
Three of those eight lines are the same number written twice, which is the point: this is not a repricing, it is one meter moving on one model. Cache writes are 1.25× the uncached input rate on both cards, so $2.50 is $2.00 × 1.25 on both — the write premium is a property of the family, not of the version. Cached reads are the only line where the two models separate: on GPT-6.1 Sol a cached read is 5% of the input rate, on GPT-6 Sol it is 10%, and OpenAI states both as fractions of the input rate rather than as independent prices.
One note on how this page states figures. OpenAI publishes the rate card as a flat list price per meter with no reasoning-effort qualifier attached to any of the eight lines above, so none of them is quoted at a particular effort level. Effort changes how many tokens a task spends, not what a token costs, which is why the card is effort-independent while your invoice is not. Nothing on this page is a matched-effort head-to-head against a rival model, and no such comparison is implied anywhere below — where a figure would depend on effort, this page names the effort it was produced at.

The one line that moved: cached input, $0.20 to $0.10
A cached-input meter prices the part of a request that did not change. When a long stable prefix — a system prompt, a set of tool schemas, a retrieved corpus, the accumulated transcript of a loop — is deliberately stored so the next request can reuse it, the reused tokens bill at the cached rate instead of the input rate, and the tokens that get stored in the first place bill at the cache-write rate. It is a discount on repetition, and it exists because the alternative for the provider is recomputing the same prefix from scratch on every call.
Halving that meter is not the same kind of news as halving the input meter, and the difference is worth being precise about. The input rate is what every uncached token in the family costs, so halving it would be a repricing of the whole catalogue line. The cached rate is a discount depth — how far below the input rate a reused token sits — so moving it is a change to the reward for a specific behaviour rather than a change to the baseline. GPT-6.1 Sol's input rate is $2.00, exactly what GPT-6 Sol charges, so a request with no cached prefix bills identically on the two models. Halving the cached meter takes the discount from 90% off input to 95% off input, and that reaches exactly the workloads whose prefixes are long enough and reused often enough to register on the line at all. Read the moved meter as "the same task costs less if you cache" rather than "the model got cheaper", and the card stops being surprising.
The arithmetic follows from the two rows and nothing else. One million cached input tokens costs $0.10 on GPT-6.1 Sol and $0.20 on GPT-6 Sol, so the release is worth ten cents per cached million — a figure that is trivial on a single call and material on a long run, because a cached prefix is read once per turn. Cache writes are untouched at $2.50, so the break-even point for caching at all is unchanged by the release: the payback on a write premium comes from reusing the prefix, not from which of the two cached rates applies.
Above 272,000 input tokens the whole request reprices
Both cards carry a long-context tier, and OpenAI's pricing page states the rule the same way for both: prompts with more than 272K input tokens are priced at 2× input and cache rates and 1.5× output for the full request. The whole request reprices — not just the tokens past the line. That sentence is the one to carry away from this section, because it is where a rate card stops being a lookup table: at 272,001 input tokens the entire request, including the first token, is billed at the long-context columns.
The repriced card for GPT-6.1 Sol, above 272,000 input tokens:
• Input — $4.00 per million tokens, 2× the short-context rate.
• Cached input — $0.20 per million tokens, 2× the short-context rate.
• Cache write — $5.00 per million tokens, 2× the short-context rate.
• Output — $15.00 per million tokens, 1.5× the short-context rate.
And the same tier on GPT-6 Sol, so the two rows can be read against each other:
• Input — $4.00 per million tokens.
• Cached input — $0.40 per million tokens.
• Cache write — $5.00 per million tokens.
• Output — $15.00 per million tokens.
Again the whole request reprices, not the excess, so a 273,000-token prompt is billed at $4.00 input for all 273,000 tokens rather than at $2.00 for the first 272,000 and $4.00 for the last thousand. That makes crossing the threshold a step rather than a slope, and it makes the cached-input change look different above the line than below it. The $0.20 that GPT-6.1 Sol charges for a cached read above 272K is exactly the rate GPT-6 Sol charges below it, so a long-context cached workload on the newer model pays the pre-release figure. The saving survives in relative terms above the line — $0.20 against GPT-6 Sol's $0.40 is still a halving — but the headline ten cents does not apply to the request that crosses the threshold, because that request is not billed on the short-context card at all.

The discount tiers, as multipliers
Everything below the standard list price is a multiplier applied to the card above rather than a separate price list. Stating them as multipliers is what makes them portable: the same three factors apply to all four meters on both models, so a reader who has the standard card has the whole tier structure.
• Batch — 50% of standard. OpenAI's Batch pricing is exactly half across all four meters on both models: $1.00 input, $0.05 cached input, $1.25 cache write and $5.00 output per million on GPT-6.1 Sol, and $1.00, $0.10, $1.25 and $5.00 on GPT-6 Sol.
• Flex — 50% of standard, and identical to Batch on both models line for line, so the choice between them is a scheduling choice rather than a price choice.
• Fast mode — 2× standard across all four meters on both cards, with no exception listed on either pricing table. Fast mode is $4.00 input, $0.20 cached input, $5.00 cache write and $20.00 output per million on GPT-6.1 Sol, against $4.00, $0.40, $5.00 and $20.00 on GPT-6 Sol — the two rows differ only on the cached line, and only by the same 2× factor each model's own standard cached rate already carries.
• Regional processing — a 10% uplift on top of whichever tier applies, for models released on or after 5 March 2026, on any residency other than the default.
Fast mode has one availability restriction that the price alone does not tell you. OpenAI's own model page for GPT-6.1 Sol states that Fast mode is unavailable with EU data residency, and its data-residency guidance lists GPT-6.1 Sol among the models that support EU residency with Standard, Flex and Batch processing. So inside the EU the tier ladder is two rungs shorter, not one: the half-price tiers are there and the doubled tier is not. A reader comparing a European deployment against a US one should compare it against Standard, because that is the fastest tier available to it — quoting a Fast-mode figure at an EU workload describes a configuration the vendor does not offer.
Nothing on this card is promotional, and nothing carries an expiry
The GPT-6.1 Sol rows above are the list rate. OpenAI's pricing page attaches no promotional qualifier and no end date to them, and none of the four meters on either the short- or long-context card is marked as a limited-time figure. That is worth stating explicitly, because the family has a counter-example sitting on the same page: GPT-5.6 Sol's promotional pricing carries a footnote committing it at least through 21 November 2026, which means any GPT-5.6 Sol price quoted today is a price with a known review point attached. Nothing on GPT-6.1 Sol has that shape.
The practical consequence is that this page does not need a date-check caveat on the numbers themselves — it needs one on the page you are reading, because a list rate can change without an expiry date having fired. It also means the GPT-6.1 Sol card is the right row to build a cost model on if you are choosing between the two Sol rungs and want the comparison to hold still: GPT-6 Sol's $0.20 cached meter is also an ordinary list rate with no promotion attached, so the ten-cent gap between them is a structural difference in the cards rather than a temporary discount one model is running against the other.
How to read anyone else's GPT-6.1 Sol price claim
Two questions settle almost every disagreement about what a Sol-class model costs, and both of them are answerable from the card rather than from the prose around it.
• Which meter is being quoted? Input, cached input, cache write and output are four different numbers on the same request, and a single headline price is almost always the input meter with the other three left implicit. A claim that reads like a total cost is usually an input rate multiplied by an assumed token mix, and the assumed mix is where the disagreement lives.
• Which context tier was assumed? The short-context and long-context cards differ by 2× on input and cache and 1.5× on output, and the tier is selected by the whole request rather than by the excess, so a figure quoted without a tier is a figure that changes the moment the request crosses 272,000 input tokens.
• Which processing tier was assumed? Standard, Batch, Flex and Fast differ by factors of 2× and 50%, and a price quoted without saying which is a price with a factor of four of ambiguity built into it before any meter question is asked.
• Which effort level was assumed, if a token count is involved? Rates are effort-independent, but token counts are not, so any per-task cost figure inherits the reasoning effort it was measured at — and a comparison between two models at different effort settings is not a comparison of their prices.
Answer those four and a price claim either survives or it does not. What it should not do is get compared against this page's four meters without the tier and the meter being named first, because the two cards are adjacent precisely so that the single moving line is visible and nothing else is mistaken for it.

Where the GPT-6 Sol rung is served
OrcaRouter serves the GPT-6 rung at OpenAI's own list rate with nothing added. The model page at https://www.orcarouter.ai/models/openai/gpt-6-sol carries GPT-6 Sol at $2.00 input and $10.00 output per million tokens, matching the card above line for line, at 0% markup, and the catalogue holds more than 200 models on the same terms. Our card for that route carries the same $0.20 cached-input and $2.50 cache-write figures OpenAI publishes, so the numbers agree meter for meter; treat every cached figure on this page as a restatement of OpenAI's published rate rather than as an independent quote from us. GPT-6.1 Sol is listed in our catalogue as well — the live card for openai/gpt-6.1-sol carries a release date of 29 September 2026, both pricing tiers mirroring OpenAI's, and a seven-day token count in the hundreds of millions — so a reader who wants the newer rung should read that page's own numbers rather than treat this page as the last word on them. Prices move; the page is the quote.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
