A generated title card headed "Unbiased's Pareto 26.10 Preview" with the subtitle "A 1M-context swap behind the same model string", above three cards reading "Released: Oct 1, 2026", "Context: 1,048,576 tokens" and "Routed price: $0.80 / $3.20 per 1M", with a footer reading "Vendor-run, preliminary benchmarks; no independent scores published as of Oct 1, 2026."
Guides & Insights

Unbiased's Pareto 26.10 Preview: A 1M-Context Swap Behind the Same Model String

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The integration does not change. The model behind it does. On October 1, 2026, Unbiased moved its model string — pareto — onto Pareto 26.10 Preview, a new release with a four-times-larger context window, a lower price on the routed catalogue, and a benchmark sheet that, by the vendor's own admission, is preliminary and unpublished in final form. If you call Pareto today by its name, you are calling 26.10 Preview, and the vendor's changelog says so in the plainest possible terms: "Pareto 26.10 Preview is now the current Pareto release… The model string stays pareto."

That single line is the whole story for anyone with Pareto in a stack. This is not a second product launched alongside the first. It is the first product, replaced in place, under a name that did not change — which means the version you get is decided by the release calendar, not by anything in your request.

What actually changed on October 1

Four things moved, and they do not all point the same way.

• Context window — 262,144 tokens on the September Pareto release, now 1,048,576. Unbiased's own documentation never states the figure; the number comes from the model's public technical listing, where it is the per-request limit with no smaller tier offered.
• Price, as routed — the pair commonly quoted for Pareto has been $2.50 per million input tokens and $7.50 per million output, with cached input at $0.25. The 26.10 Preview listing carries $0.80, $3.20 and $0.03 respectively — roughly a third of the input rate and a little over 40% of the output rate.
• Benchmark scores — published for the first time for this release, and preliminary by the vendor's own label.
• The stable option — 26.10 Preview is a preview. Unbiased's catalogue text says it "may change without notice" and points anyone who needs predictable behaviour at the previous release instead.

Everything else held. Maximum output is still 131,072 tokens. Input is still text and image; output is still text only. There is still no Hugging Face identifier on the listing, so the weights remain closed. And there is still exactly one serving provider — Unbiased itself — with no second host inside the listing.

A generated single-column card titled "Pareto 26.10 Preview - what changed" with six rows reading "Context: 262K to 1M", "Max output: 131,072 (unchanged)", "Input price: $2.50 to $0.80 per 1M", "Output price: $7.50 to $3.20 per 1M", "Benchmarks: four, preliminary" and "Providers: one, the vendor", with a footer reading "Preview figures per the public listing; the vendor's own site keeps the $2.50 / $7.50 card."

The price is the part worth reading twice

On Unbiased's own platform, the rate card did not move. The model card and pricing page both still read $2.50 per million input, $0.25 cached, $7.50 output — the same three numbers the September release carried — and because the model string stayed pareto, a customer on the vendor's API is paying those rates for 26.10 Preview. The lower figures are the ones on the routed catalogue listing, and the gap between the two is exactly the kind of thing that decides a migration.

Two readings are possible and the vendor has not said which is true. Either 26.10 Preview is genuinely being served cheaper through that route as an introductory push, or the listing represents a different commercial arrangement than the direct platform. Unbiased publishes no promotional end date, and a price set on the day of launch is not a commitment.

This is the one part of the story a router changes the shape of. OrcaRouter passes the provider's list price straight through with 0% markup, so a vendor price cut is live on our side the same day it lands rather than on a contract cycle — and automatic failover means a preview release that behaves differently next week can be swapped out without a code change. We do not carry Unbiased's Pareto 26.10 Preview, so the honest version is this: if you want this specific model, call it through Unbiased's own API. The point is only that on a preview release, the price and the version are both variables, and neither should be pinned in code.

The benchmark sheet, with the labels it came with

Unbiased published four scores for 26.10 Preview, each paired with a measured cost per task, and each carrying a caveat the vendor wrote itself. Read them as vendor-run and unreproduced, because that is what they are:

• DeepSWE v1.1 (agentic coding) — 69.9%, at $0.24 per task
• Terminal-Bench 4.0 (agentic terminal use) — 50.8%, at $0.48 per task
• Humanity's Last Exam, text-only (expert reasoning) — 49.9%, at $0.008 per task
• GPQA-Diamond (science reasoning) — 92.4%, at $0.004 per task

A generated single-column scoreboard titled "Pareto 26.10 Preview - the benchmark sheet" listing six rows: "DeepSWE v1.1: 69.9% at $0.24 per task", "Terminal-Bench 4.0: 50.8% at $0.48 per task", "HLE text-only: 49.9% at $0.008 per task", "GPQA-Diamond: 92.4% at $0.004 per task", "Context: 1,048,576 tokens" and "Independent scores: none yet", with a footer reading "Vendor-run, preliminary results; not independently validated."

The vendor's framing is that 26.10 Preview is "matching or setting the Pareto frontier on all four benchmarks." The sheet it publishes alongside that claim is more specific, and it is to Unbiased's credit that it is published at all: on Terminal-Bench 4.0, GPT 6 Astra is listed at 58 and Fable 5.1 at 56, both above Pareto's 50.8, and the vendor's own "where other models score higher" section says so directly. On DeepSWE v1.1 the release sits in a cluster — GLM-5.3 and Kimi K3 are each listed at 69.0 against Pareto's 69.9 — which is a tie in practice and inside whatever error bar a single run carries.

The comparison numbers carry a second caveat that matters more than the first: they are transcribed from a September 21 comparison and, in the vendor's words, are "not independently validated," and they were produced in a separate evaluation of individual models — so they should not be paired with Pareto's cost-per-task figures as though both came off the same bench. The vendor says so in the evaluation details. A reader who skips that line will over-read the chart.

What none of this is: independent. Artificial Analysis, the board most teams check before trusting a new name on a price sheet, has no page for Pareto at all. Every number above was produced by the company selling the model. The one thing that softens it is that the eval harness is public — the vendor publishes a reproducible head-to-head suite that anyone can point at an endpoint — which turns "trust our score" into "rerun our score." That is a real offer and it is not the same as a verified number. Nobody has published the rerun yet.

What "preview" means in the contract

The word is doing more work here than the release notes do. Unbiased's own documentation describes current access as evaluation access under its terms, and states that commercial or production use requires a separate written agreement. New accounts are reviewed by hand before an API key is issued. The API is OpenAI- and Anthropic-compatible — base URL, key, model string, no rewrite — but the documented surface is chat completions and messages only, with known gaps the vendor lists openly: GET /v1/models is not implemented and unknown model identifiers are not rejected, so callers must send pareto explicitly rather than discovering what is available.

Put the preview label and the private-beta contract together and the operating advice writes itself. This is a model to evaluate against your own workload behind a routing layer, not one to put in front of revenue traffic and forget about. The mechanism for that is unglamorous: a fallback chain configured once, so a release that changes under you mid-week becomes a config change rather than an incident. Unbiased spent September fixing exactly this class of problem on its own side — a changelog entry from September 16 records that roughly 13% of long non-streaming requests were hitting a 120-second timeout, and the gateway ceiling was raised to 300 seconds. That is a vendor being candid about a rough edge. It is also a reminder that a preview's operational profile is still being written.

A screenshot of the Unbiased Pareto 26.10 Preview model card page, headed "Pareto 26.10 Preview model card" and subtitled "Pareto is our blended AI model, available through the Unbiased AI platform. Send one request and receive one answer.", showing the DeepSWE v1.1 and Terminal-Bench cost-per-score chart with the Pareto 26.10 Preview point at 69.9% and $0.24 alongside other labs' published points, over an evaluation-details block stating the Pareto 26.10 Preview results are from October 1, 2026 runs and may change before final publication.

What to watch before you commit

Three specific things would settle the open questions, and each is checkable.

The first is an independent score. Until a third party runs 26.10 Preview on a published harness, the 92.4 on GPQA-Diamond and the 50.8 on Terminal-Bench are vendor claims about vendor work — good enough to decide whether to test, not good enough to decide whether to migrate.

The second is what the preview does to the price. If the routed listing stays at $0.80 and $3.20 while the vendor's own platform holds $2.50 and $7.50, the difference is a channel decision worth understanding before you build a cost model on either number.

The third is what "may change without notice" turns out to mean in practice. A preview that is revised twice in a month is a different instrument from one that is frozen and renamed at the end of the quarter, and the changelog is where that will show up first.

None of this argues against trying it. A 1M-token multimodal model that reports frontier-adjacent scores at $0.24 a task is worth an afternoon of real work on your own prompts, and that evaluation costs less than the argument about it. It argues against assuming the model you benchmark this week is the model you get next week — which, for a release the vendor itself calls a preview, is the only assumption that needs to be right.

A preview release is exactly the case automatic failover was built for: automatic failover turns a model that changes under you mid-week into a config change rather than an outage.