A generated title card reading 'GPT-6.1 Sol Context Window' with the subtitle '1,050,000 tokens, the 922,000 line, and the 272,000 cliff'. Three rounded panels sit below it: 1,050,000 labelled 'context window, on the vendor page', 922,000 labelled 'maximum input, in the docs form', and 272,000 labelled 'the input line that reprices the whole request'. The OrcaRouter logo appears in the bottom-right corner.
Guides & Insights

GPT-6.1 Sol Context Window: 1,050,000 Tokens, the 922,000 Line, and the 272,000 Cliff

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The vendor's model page for GPT-6.1 Sol, read on 7 October 2026, states a 1,050,000-token context window and a 128,000-token maximum output. The same page, in the machine-readable form you get by appending .md to its URL, carries a third figure the rendered page never prints: a maximum of 922,000 input tokens. GPT-6 Sol, the model 6.1 was released to succeed on 2026-09-29, publishes the identical pair of ceiling figures on its own page, and the same 922,000 line in its own markdown form. Our own model card for openai/gpt-6-sol reports the window as 1,050,000 tokens and the output cap as 128,000, displays the first of those as "1M" in its spec strip, and prints "1.1M" for the same model in a comparison table further down the same page.

So the pages do disagree, and it is worth being precise about how. the vendor's rendered spec strip gives a window and an output ceiling and no input ceiling. the vendor's markdown documentation for the same model gives all three. Anyone sizing a request off the rendered page is working with one fewer constraint than the vendor published, and the missing one is the number that decides whether a request fits.

Three figures, three sources, and one subtraction nobody writes down

Here is each number with the document it came from, all read on 7 October 2026.

• 1,050,000 context window — the vendor's model page for gpt-6.1-sol, in the rendered spec strip and in its markdown form, and the same figure on the page for gpt-6-sol. It is also what our catalogue returns for openai/gpt-6-sol and openai/gpt-6-luna, where the field is typed as 1,050,000 rather than rounded.

• 128,000 max output tokens — the same page, same two forms, for both generations. Our card's field says 128,000; its display rounds to "128K".

• 922,000 maximum input tokens — the markdown form of the vendor's gpt-6-sol model page and of its gpt-6.1-sol page. It is not in the rendered strip of either page, and it is not in our catalogue field for the model, which stops at the window and the output cap.

The three figures are arithmetically consistent with one another: 922,000 plus 128,000 is exactly 1,050,000. the vendor's own reasoning guide describes the mechanism that makes the identity meaningful without ever doing the sum on the model page — reasoning tokens, it says, "still occupy space in the model's context window", and if generated tokens "reach the context window limit or the max_output_tokens value you've set", the response comes back marked incomplete. A window shared between what goes in and what comes out is a window in which the input ceiling is the window minus the output reservation.

That reading is corroborated, not proven, and it is worth separating the two. What is documented is a window of 1,050,000, an output cap of 128,000, and a maximum input of 922,000. What is inferred is which of those is the constraint that fires first. The inference holds for every request that reserves its full output allowance and fails for any request that does not — set max_output_tokens to 4,000 and 1,046,000 input tokens is not obviously refused. Until the vendor writes the subtraction into the page the numbers live on, treat the pairing as the shape of the budget rather than as a hard admission rule, and validate against the counting endpoint rather than against a blog post.

Context is not a price: the 272,000-token step

A bigger window is a capacity claim. It is not a cost claim, and on this family the two come apart at a documented line. the vendor's pricing page states the rule in one sentence: prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request. The same page's own definition of its two columns is "Short context: ≤272K input tokens. Long context: >272K input tokens."

Read the word full carefully. The tier does not tax the tokens past the line — it reprices everything, including the first token. And it is not a 6.1 change: the identical rule, with the identical threshold, applies to GPT-6 Sol, which is why the cliff has to be attributed to the tier and not to the release.

Work one long-context job through both sides, on GPT-6.1 Sol's published rates, with the prefix already resident in cache so the cache-write charge does not muddy the comparison:

• 240,000 input tokens (200,000 cached, 40,000 fresh), 6,000 output — fresh input 40,000 at $2.00 per million is $0.080; cached input 200,000 at $0.10 is $0.020; output 6,000 at $10.00 is $0.060. Total, $0.160.

• 300,000 input tokens (260,000 cached, 40,000 fresh), 6,000 output — the request is now above the line, so every rate moves: fresh input 40,000 at $4.00 is $0.160; cached input 260,000 at $0.20 is $0.052; output 6,000 at $15.00 is $0.090. Total, $0.302.

Twenty-five percent more input tokens buys an 89 percent larger bill. Run the same pair on GPT-6 Sol and the shape holds with a steeper slope — $0.180 below the line against $0.354 above it, because the older card's cached rate of $0.20 sits at the same place 6.1's long-context cached rate does. The crossing is a step, not a slope, and the cheapest way to see that is to cross it by a hair. A fully uncached 271,000-token request with 1,000 output tokens costs $0.552; at 273,000 it costs $1.107. Seven-tenths of a percent more input tokens, 2.01 times the money. Trim 2,000 tokens back off that request and the bill falls from $1.107 to $0.552 — under one percent of the input for half the cost.

Nothing in this section is about GPT-6.1 Sol's window being large. It is about the window being large enough to reach a boundary that costs more than the window's size buys.

A generated two-column scoreboard titled 'GPT-6.1 Sol vs GPT-6 Sol - the scoreboard'. Both columns carry the same six rows. GPT-6.1 Sol reads: context window 1,050,000 tokens; max input 922,000 tokens (docs form); max output 128,000 tokens; input price $2.00 / $4.00 per M; cached input $0.10 / $0.20 per M; output price $10.00 / $15.00 per M. GPT-6 Sol reads the same on every row except cached input, which is $0.20 / $0.20 per M. A footer line credits OpenAI's model and pricing docs read Oct 7 2026 and explains that the second figure in each price row is the long-context rate above 272,000 input tokens. The OrcaRouter logo appears in the bottom-right corner.

What actually fills 1,050,000 tokens

A context budget is made of six things, and they are not equally cacheable. Approximate shares of an illustrative 240,000-token agentic request; the proportions are our own, the cacheability rules are the vendor's, from its prompt-caching guide read the same day.

• Provider-injected system content and request formatting — rendered ahead of your messages, billed as input, and explicitly excluded from the minimum cacheable length. Not yours to control, and not yours to trim.

• Your developer and system instructions, around 6,000 tokens. Cacheable. This is the front of the prefix, so a change here invalidates everything behind it.

• Tool definitions and schemas, around 14,000 tokens for the hosted tool surface plus your functions. Cacheable, and the most brittle part of the prefix: the guide lists tool names, descriptions, schemas, ordering and tool-specific instructions as things that move the prefix boundary.

• The retrieved corpus, around 180,000 tokens. Cacheable, and by far the largest line. It is worth caching only if it is byte-stable between calls — a corpus reassembled per request is a full-price corpus.

• The accumulated transcript — earlier turns and tool results — around 30,000 tokens and rising. Cacheable up to the newest change; the tool result that arrived on this turn is fresh input at the full rate.

• Reasoning tokens — generated, never cached, billed as output. They consume window and they are invisible in the response body.

Two operational notes fall straight out of that list. First, cache entries are held on individual machines: the guide says a request can reuse a prefix "only if it reaches a machine holding a matching entry that has not expired", and that overflow routing begins above roughly 15 requests per minute. A prefix that is stable in your code can still miss in production. Second, the minimum cacheable prefix is 1,024 visible input tokens, and hidden system tokens do not count toward it — so a small system prompt is not a cacheable prefix no matter how large the request around it is.

The output ceiling is a separate budget, not a second helping

128,000 max output tokens does not mean 128,000 tokens of answer. The reasoning guide is explicit that max_output_tokens caps the total the model generates, "including reasoning tokens, visible output tokens, and non-visible formatting tokens", and that reasoning tokens are billed as output while occupying space in the window.

That makes truncation a design decision rather than an edge case, because of how it fails. When generation reaches the limit, the response returns with a status of incomplete and a reason of max_output_tokens — and the guide warns this "might occur before any visible output tokens are produced, meaning you could incur costs for input and reasoning tokens without receiving a visible response." A budget that spends the whole window on input and leaves the output reservation to chance is a budget that can bill a full long-context request and return nothing a caller can parse. the vendor's own starting recommendation is to reserve at least 25,000 tokens for reasoning and outputs while you are still measuring what a prompt actually needs.

GPT-6.1 Sol sharpens this, and it is one of the few genuinely 6.1-specific lines in the release. Its reasoning effort ladder runs low, medium, high, xhigh and max, and the none and minimal settings are not supported. GPT-6 Sol accepts all six. There is therefore no setting on 6.1 that turns the reasoning spend off, the default is medium, and the output side of the budget is never free.

The cached-input halving, read where the window is widest

The only rate on the GPT-6.1 Sol card that moved against GPT-6 Sol is cached input: $0.20 down to $0.10 per million tokens, which the model page expresses as 5% of the uncached input rate, and which the vendor's caching guide names explicitly as the 0.05x case against the 0.1x that most GPT-5.6-and-later models read at. Input, cache writes and output are identical on both cards, and the long-context multipliers are identical too.

For exactly the workload this page is about, that is the right meter to have moved, and the reason is the composition of a long request rather than its size. On the 300,000-token job above, 260,000 of the input tokens are a cached prefix — 87 percent of everything the request sends. The cached line is therefore the largest single input meter in the bill, which is the general property of long-context work: the longer the window you use, the more of it is a prefix you have already sent. Halving that meter is worth $0.052 on that request.

And the cliff takes back more than the halving gives, on the same request. Priced at the short-context rates it would have paid below the line, the same 300,000-token job on GPT-6.1 Sol would cost $0.166 instead of $0.302 — a crossing cost of $0.136, or about 2.6 times what the release's one changed meter is worth. Above the line the cached rate reads $0.20, which is not a new number on this family: it is double the headline on the 6.1 card and exactly what GPT-6 Sol charged for a cached read below the line before this release. A long-context cached workload collects the headline change and then hands it back at the boundary, and the boundary — not the model — is the cause.

A screenshot of the machine-readable markdown form of OpenAI's GPT-6.1 Sol model documentation, captured October 7 2026, showing the Model details block with the three figures on consecutive lines — 1,050,000 context window, Maximum input tokens: 922,000, and 128,000 max output tokens — above the Text tokens pricing table listing Input $2, Cached input $0.1, Cache writes $2.5 and Output $10 per 1M tokens, the note that cached input tokens are priced at 5% of the uncached input rate, and the sentence 'Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.'

How to size a context budget

As a procedure, in the order the constraints bind:

• Count the request, do not estimate it. POST the exact payload — tools, images, files and all — to the input token count endpoint on the Responses API. The guide is blunt about why: the count includes formatting tokens for message roles and boundaries that never appear in text you can tokenize locally, and local estimates like characters over four are inaccurate for images, files and schemas.

• Reserve the output side first. Choose max_output_tokens remembering that it covers reasoning, visible output and formatting together, and start from the vendor's 25,000-token buffer rather than from zero. Your input budget is the window minus that reservation, and the nine-hundred-twenty-two figure is the vendor's version of the same subtraction.

• Price the request on both sides of 272,000 before you send it. The step is large enough that a request designed to land just under the line and a request designed to land just over it are different products.

• Order the prefix by stability. Instructions, then tool schemas, then the corpus, then the transcript. Anything that changes between calls belongs at the end, where it costs a prefix match rather than the whole cache.

• Clear 1,024 visible input tokens before expecting a cache. Below that minimum nothing caches, and hidden provider tokens do not count toward it.

• Check reuse is plausible. A prefix must be reused inside the 30-minute cache lifetime and must land on the machine holding the entry; both are described in the guide and neither is a property of your code.

• Re-measure after any model or setting change. Moving to GPT-6.1 Sol removes the reasoning-off position, which changes reasoning token counts and therefore the output side of the budget — and a change to reasoning effort, tools, structured output schema or context management can also move a prefix boundary and cost you the cached rate entirely.

• Decide what happens when the job will not shrink. Compaction is the documented escape: a Responses request can set context_management with a compact threshold, and the server replaces earlier conversation content with an opaque compaction item that carries key state forward in fewer tokens. It is a budget decision rather than a free trim, because the guide notes compaction "can prevent reuse from the first changed token onward" — a compaction pass invalidates the prefix behind it.

A screenshot of the OrcaRouter model page at www.orcarouter.ai/models/openai/gpt-6-sol, captured October 7 2026, showing the OrcaRouter nav bar, the breadcrumb Home -> Models -> OpenAI, the model identifier openai/gpt-6-sol attributed to OpenAI with the date 2026-09-22, the Vision, Tools, JSON and Reasoning capability badges, the spec tiles reading Max output 128K, input text + image + file, output text and a p50 TTFT of 1.44 s, the description stating a 1.05M-token context, and the /v1/chat/completions rate row of $2.00 in and $10.00 out per 1M tokens.

The arithmetic on this page starts from the generation GPT-6.1 Sol replaces, and that rung is the one that is callable today: our card for openai/gpt-6-sol reports a 1,050,000-token context window with 128,000 tokens of maximum output at OpenAI's list rates with 0% markup — the vendor's price is the price on the page, and a vendor move lands there the same day rather than at a renewal. Our card carries no maximum-input field of its own, so the 922,000-token figure for that model has to come from the vendor's own documentation, which is what this page has used throughout. What the card is useful for is sizing: the window and the output ceiling it does publish are the two numbers the budget procedure above subtracts from each other, and the rung below is the one you can actually run that procedure against while 6.1 is still new. It is at https://www.orcarouter.ai/models/openai/gpt-6-sol.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily