A generated hero card for a GPT-6.1 Sol versus GLM-5.3 comparison, headed 'Seven index points, 2.8 times the money', with badges reading 'GPT-6.1 Sol: $0.72 per index task', 'GLM-5.3: $2.01 per index task' and 'GLM-5.3 weights: downloadable since Aug 25, 2026', a footer reading 'Independent figures per Artificial Analysis v4.3.2, max effort; AI-generated card', and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

GPT-6.1 Sol vs GLM-5.3: The Weights Shipped, and the Bill Didn't Move

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Open​AI released GPT-6.1 Sol on 29 September 2026 at its DevDay keynote, and Z.ai released GLM-5.3 six weeks earlier, on 18 August 2026, then held the checkpoint back for a week. That hold is over: zai-org/GLM-5.3 has been on Hugging Face since 25 August, a BF16/FP8 checkpoint at roughly 753 billion parameters of which about 40 billion are active per token, and it has since passed 1.4 million downloads. So the question most comparisons have been asking since August — is this an open model or a rental? — has an answer, and the answer did not change the arithmetic. On the independent board, GPT-6.1 Sol scores 51.8 and costs $0.72 per completed index task; GLM-5.3 scores 44.8 and costs $2.01. You can now own the cheaper-by-the-token model and still pay nearly three times more per finished job.

What each model actually is

GPT-6.1 Sol is Open​AI's mid-tier GPT-6 model, one rung below the GPT-6 Astra flagship and above the budget GPT-6 Luna. It carries a 1,050,000-token context window with a 128,000-token output ceiling and an April 30, 2026 knowledge cutoff, and it accepts text, images and files as input. Reasoning runs from low to max with a default of medium; the none and minimal values that GPT-6 Sol accepted now return an error, which Open​AI's own migration guidance addresses by telling developers to substitute low. It costs $2.00 per million input tokens and $10.00 per million output, with cached input at $0.10.

GLM-5.3 is Z.ai's current flagship for software engineering and long-horizon agentic work, built on the same mixture-of-experts base as GLM-5.2 with the gains coming from post-training rather than from a bigger model. It is text-in, text-out — no vision, at any price — with a 1M-token window, a 128,000-token output cap and thinking switched on permanently. There is no non-reasoning mode, which is a hard constraint for a latency-sensitive call. Z.ai lists it at $1.40 in and $4.40 out per million tokens with cached input at $0.26; our own catalogue carries it at $1.26 and $3.96 with a $0.234 cache read, so the vendor's card is the more conservative number and the one to argue from.

The rate card, and the line that inverts it

One line per dimension, both sides on each:

• Input — GPT-6.1 Sol $2.00 per 1M vs GLM-5.3 $1.40 per 1M; GLM is 30% cheaper
• Output — GPT-6.1 Sol $10.00 per 1M vs GLM-5.3 $4.40 per 1M; GLM is 56% cheaper
• Cached input — GPT-6.1 Sol $0.10 per 1M vs GLM-5.3 $0.26 per 1M; the ordering flips, Sol is 2.6× cheaper
• Long-context clause — GPT-6.1 Sol reprices the whole request above 272,000 input tokens at 2× input and cache and 1.5× output vs GLM-5.3 no equivalent clause on the published card
• Context and output — 1,050,000 in / 128,000 out vs 1M in / 128,000 out; a 5% difference, immaterial
• Modality — text, image and file in, text out vs text only
• Reasoning floor — GPT-6.1 Sol low, medium (default), high, xhigh, max vs GLM-5.3 always on, no floor to drop to
• Weights — closed, API only vs open, downloadable, FP8
• Independent score, Artificial Analysis v4.3.2 at max effort — 51.8 vs 44.8
• Independent cost per index task — $0.72 vs $2.01
• Output tokens for the index run — 67 million vs 210 million

img src="2.png" alt="A generated two-column comparison scoreboard titled 'GPT-6.1 Sol vs GLM-5.3 - the scoreboard'. Left column 'GPT-6.1 Sol': rows reading 'Intelligence Index: 51.8', 'Cost per index task: $0.72', 'Input price: $2.00 per 1M', 'Output price: $10.00 per 1M', 'Output tokens on index run: 67M', 'Weights: closed'. Right column 'GLM-5.3': rows reading 'Intelligence Index: 44.8', 'Cost per index task: $2.01', 'Input price: $1.40 per 1M', 'Output price: $4.40 per 1M', 'Output tokens on index run: 210M', 'Weights: open on Hugging Face'. A footer reads 'Independent figures per Artificial Analysis v4.3.2 at max effort; vendor list prices.' The OrcaRouter logo sits in the bottom-right corner." /> That last pair is where this matchup is decided, and it is not a subtle effect.

A generated two-column comparison scoreboard titled 'GPT-6.1 Sol vs GLM-5.3 - the scoreboard'. Left column 'GPT-6.1 Sol': rows reading 'Intelligence Index: 51.8', 'Cost per index task: $0.72', 'Input price: $2.00 per 1M', 'Output price: $10.00 per 1M', 'Output tokens on index run: 67M', 'Weights: closed'. Right column 'GLM-5.3': rows reading 'Intelligence Index: 44.8', 'Cost per index task: $2.01', 'Input price: $1.40 per 1M', 'Output price: $4.40 per 1M', 'Output tokens on index run: 210M', 'Weights: open on Hugging Face'. A footer reads 'Independent figures per Artificial Analysis v4.3.2 at max effort; vendor list prices.' The OrcaRouter logo sits in the bottom-right corner.

The 210-million-token problem

Artificial Analysis runs the same ten-evaluation suite against every model on the board and publishes the token counts behind each score. GLM-5.3 generated roughly 210 million output tokens completing the run. GPT-6.1 Sol generated 67 million. Read against each model's own comparison class on the board, GLM-5.3's 210 million sits above a class median of about 140 million while Sol's 67 million sits well below the roughly 81 million median of the flagship class it is scored in. Since output is the expensive side of every rate card, GLM-5.3's headline 56% discount on the output line does not survive contact with the measurement: it emits 3.1× the tokens at 0.44× the price, and 3.1 × 0.44 is 1.37 — GLM-5.3 is the more expensive model on this workload despite a materially cheaper rate card.

The board's own cost-per-task figure confirms the direction and the rough size. $2.01 against $0.72 is 2.8×, wider than the token ratio alone implies, because GLM-5.3 also pays more on the input and cache lines per task. It is worth being precise about what this does and does not prove: the index run is ten fixed evaluations, and verbosity on them is not the same as verbosity on your traffic. But for a buyer, tokens are what gets invoiced, and the shape of the difference is the same on any task set where the model has to produce a long answer.

The partial offset is the cached-input line. GLM-5.3's cache read is $0.26 against a $1.40 base, and GPT-6.1 Sol's is $0.10 against a $2.00 base — Sol wins by a factor of 2.6 on a rate that dominates any workload with a stable prefix. If your traffic is a large fixed corpus with a one-paragraph answer, GLM-5.3 is straightforwardly the cheaper option and you should take it. If your traffic looks like the traffic both of these models were built for — agent loops, repo-scale edits, long generated outputs — the cached-input advantage shrinks to noise next to a 3× token ratio.

Where the vendor numbers land

Z.ai's launch figures are strong and, as always, unreproduced. Terminal-Bench 2.1 at 88.2, DeepSWE v1.1 at 66.9, CyberGym at 84.5, ExploitBench at 54.4, AutomationBench at 48.2, GDPval-AA v2 at 1769 Elo. The one figure worth reading twice is Terminal-Bench 3.0 at 28.3 — the successor benchmark, and the correction to the 88.2 on the older one. A model can be excellent on the benchmark its post-training targeted and ordinary on the next generation of the same benchmark.

Open​AI's launch numbers for GPT-6.1 Sol are framed entirely around cost per completed task rather than raw score: DeepSWE v1.1 beating GPT-6 Sol's best by 6.4 percentage points at a lower reasoning effort, AutomationBench 1.0.6 at 2.2 points above Claude Opus 5.5 at medium effort, OSWorld 2.0 up seven points on GPT-6 Sol at maximum effort, Terminal-Bench Science 0.1 more than doubling GPT-6 Sol at less than half the cost per task. Note that these are Open​AI's harness at Open​AI's settings on Open​AI's selected benchmark set, and that none of them has been reproduced independently.

The overlap between the two vendors' published sets is thin, and the one direct comparison available is on the same benchmark: Z.ai reports 66.9 on DeepSWE v1.1, Open​AI reports 68.8 for GPT-6 Sol — the model 6.1 replaces — and a 6.4-point improvement for 6.1 Sol over its own predecessor. Two points apart on the vendor-reported comparison, and roughly six apart on Open​AI's own delta. That is close to a controlled comparison, and it says what the independent board says: near parity on capability, a wide gap on what finishing the job costs.

One API, or two contracts

Both models are on OrcaRouter, which is the unusual part of this pairing. GPT-6.1 Sol landed in the catalogue at Open​AI's own list price on release day — $2.00 and $10.00 per million tokens, cached input $0.10, with the 272,000-token step to $4.00 and $15.00 passed through exactly as the vendor lists it — and GLM-5.3 sits alongside it at $1.26 and $3.96 with a $0.234 cache read. Nothing is added on top of either: the platform passes provider list price through unchanged, which is why a vendor price cut shows up on our side the same day rather than after a reseller renegotiates.

A screenshot of the OrcaRouter model page for z-ai/glm-5.3, showing the listing dated 2026-08-18, the model described as Z.ai (Zhipu AI)'s latest flagship for complex software engineering and long-horizon agentic tasks, a 1M-token context with 128K maximum output, text in and text out, and the pass-through list price of $1.26 input and $3.96 output per million tokens, with the page's own pricing tabs and a 4.00 s median time to first token visible.

That matters more for this specific pair than for most. The decision between them is not "which is better" — it is a threshold on your own output length, and it will move the first time Z.ai sharpens its rate card or Open​AI changes the 6.1 tier's tier boundary. Holding both behind one key, with automatic failover between them and a routing rule you change in configuration rather than in a deployment, is what lets you test that threshold on real traffic instead of arguing about it. If Sol's cache line wins your workload, route on it; if your answer lengths stay short and GLM-5.3's flat card with no context cliff wins the segment, route on that instead — the same key, no second contract, no second bill.

A screenshot of the OrcaRouter model page for openai/gpt-6.1-sol, showing the listing dated 2026-09-29, a 1M-token context with 128K maximum output, text, image and file input, a reasoning effort ladder running from low to max, and the pass-through list price of $2.00 input and $10.00 output per million tokens with cached input at $0.10.

Which one, if you have to pick today

• Long generated outputs, agentic loops, output-heavy work — GPT-6.1 Sol, and the margin is close to threefold once the token counts are applied. This is where the rate card lies and the measurement does not.
• Stable prefix, short answer — GLM-5.3. Its cached-input and input rates are genuinely better and the verbosity never gets a chance to bite.
• Any request above 272,000 input tokens — GLM-5.3, because GPT-6.1 Sol reprices the entire request at that boundary and GLM-5.3's card has no equivalent step.
• Anything with an image or a document as visual input — GPT-6.1 Sol. GLM-5.3 is text-only and this is not close.
• Inference inside your own boundary, or a hedge against a vendor's terms changing — GLM-5.3, and now the download exists rather than being promised. Budget for the hardware: the checkpoint is a large FP8 artifact and self-hosting is a capital decision, not an API one.
• Latency-sensitive calls — neither without care. GLM-5.3 has no way to switch reasoning off; GPT-6.1 Sol has a floor at low, and its own time-to-first-token on the independent board measured 332 seconds against a board median of 3.7.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily