
Pareto 26.10 Preview vs Pareto: Same Model String, Two Different Models
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 221 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 103 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 214 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Unbiased shipped Pareto 26.10 Preview on October 1, 2026 and left the previous release, Pareto, listed beside it. The two are not variants of a family you choose between with a flag. They are two releases behind one brand, and the vendor's own model string — pareto — now resolves to the newer, explicitly unstable one. So the comparison question is not really "which is better." It is: which one does your request actually reach, and what happens to the answer when that changes underneath you?
The only place the two are described in the same breath is the model's public technical listing. Unbiased's own model card, pricing page and FAQ describe one stacked product and never mention the split. That asymmetry — a detailed public diff on one side, silence on the other — is the honest starting point for a head-to-head, because it tells you where every number below came from.
Six dimensions, both sides
Everything that separates the releases sits in these six lines. The left side is Pareto 26.10 Preview; the right is the Pareto release from September 17, the one the preview's catalogue text calls the stable choice.
• Context window — 1,048,576 tokens vs 262,144 tokens. Four times the room, on paper, with no smaller tier offered on either.
• Maximum output — 131,072 tokens vs 131,072 tokens. Unchanged. The larger window did not come with a larger answer.
• Input price, as routed — $0.80 per million vs $2.50 per million. The preview is roughly a third of the price on that listing.
• Output price, as routed — $3.20 per million vs $7.50 per million. Under half.
• Cached input, as routed — $0.03 per million vs $0.25 per million. The longest agent trajectories gain the most here.
• Published benchmark sheet — four scores, explicitly "preliminary" and "may change before final publication" vs no published benchmark table at all.
Two of those rows are the ones that decide real work. The context window is the headline the vendor will not state on its own site — it lives on the listing, not the model card — and the cached-input price is the one that changes an agent's bill rather than its ceiling. Everything else is either a marketing delta or, in the case of the output score, an admission: publishing a preview's numbers with a "preliminary" label is more honest than the alternative, and it is also a warning.

The price you see depends on where you look
Run the two rate cards side by side and they do not reconcile, which is the part of this comparison nobody has explained.
Unbiased's own platform, on the day 26.10 Preview became the current release, still lists $2.50 per million input, $0.25 cached and $7.50 output. Those are the September numbers, unchanged, on the pricing page and the model card alike. The lower set — $0.80, $0.03, $3.20 — appears on the routed catalogue listing for the same preview. So a customer calling pareto straight through the vendor's API pays the old card for the new model, while the routed listing advertises roughly a third of it. Both are live. The vendor has published no promotional end date and no statement reconciling them.
That is a channel question, and on a preview release it is also a timing question: whichever number you model your costs on, the other one is one release note away from becoming correct. A router that passes a provider's list price through at 0% markup removes exactly this friction — a vendor price change is live the same day rather than at the next contract review — and that is the one structural thing worth borrowing from OrcaRouter here. We do not carry either Unbiased release today, so the direct answer to "where do I call it" is Unbiased's own API, on whichever of the two price sets the listing you buy through happens to carry. The reason to say it out loud is that the difference is a factor of three, and a preview is not the place to discover you priced the wrong one.
What the preview buys you, and what it costs you
The preview's own catalogue description sets the terms: it "may change without notice," and anyone who needs predictable behaviour is pointed back at the September release. That is not boilerplate — it is the product definition. Compare that with the practical shape of a long agent session and the trade becomes concrete.
An agent that keeps a conversation going accumulates context, and long conversations are where prompt caching earns its keep. A changed model behind an unchanged endpoint is the classic way to lose it: the cache keys to the model that produced the tokens, so a silent version swap mid-trajectory can invalidate what you had cached and re-bill work you thought was already paid for. Unbiased's own documentation argues this point in the other direction — it positions its blend as cache-stable where a switching router is not — so this is a risk the vendor understands well and is simply accepting in exchange for shipping a preview. It is worth naming because the failure is invisible in a dashboard: your tokens still bill, your answers still return, and the number is just larger than last week.
The preview's upside is real and mirrors the downside. Four times the context, cached input at an eighth of the stable release's rate on the routed listing, and a benchmark sheet at all — the September release never had published scores. If your workload is long-context and cost-sensitive, the preview is the one you want, and the way to take it without betting the stack is to pin it deliberately: a named endpoint for the preview, a named endpoint for the stable release, and a fallback between them configured once. The routing DSL exists for exactly this shape of problem — compose the two as separate legs, send a slice of real traffic down each, and let the request log tell you what each actually cost. A preview should be a decision you can reverse in one config line, not a property of your endpoint that you discover by invoice.

A performance picture that is not stable enough to trust
Every figure above comes from the public listing, and the listing's own performance panel moved between reads on the day of launch. One read of the 26.10 Preview page reported a median latency of 5.28 seconds and 35 tokens per second; a side-by-side view published the same day reported 3.54 seconds and 21 tokens per second for the preview and 1.05 seconds and 45 tokens per second for the September release. Those are not small discrepancies. A brand-new model served by a single provider has a thin measurement base, and a rolling window over hours will move.
The honest position is that neither release has a throughput or latency figure stable enough to plan against, and for the preview that is baked into the product rather than a snapshot problem. The September release has at least had weeks of traffic — its listing shows multi-day availability in the high nineties with a full provider history behind it. The preview has hours. If your reason for choosing between them is speed rather than price or context, the evidence to choose on does not exist yet, and the responsible move is to measure on your own prompts rather than read a number off either page.

Which one to call, and how to stop caring
Call the preview if your work is long-context, cost-sensitive, and evaluable — agent sessions, batch processing, anything where four times the window and a third of the input price pay for the risk of a mid-week change. Treat it as a trial with a config line underneath it, and keep your own numbers, because the vendor's will move.
Call the stable Pareto if your workload is production, latency-sensitive, or covered by a compliance review that does not want a model described as possibly changing without notice. You give up the 1M window, the cheaper cache, and the published scores; you keep an endpoint with a track record and a name that means the same thing next week. For a lot of teams, that is the whole requirement.
The mistake is to treat this as a permanent fork. Unbiased built the split to be temporary — the preview exists to be evaluated and then absorbed — and the decision that matters is not which of the two you pick, but whether switching between them is a five-minute change or a quarterly project. On one endpoint with one model string, it is currently neither: it is an announcement you have to catch. Configure the choice as a routing decision instead and the version you call becomes a knob, which is the only posture a preview release has ever really rewarded.
