
Boreal vs Qwen3.8 Max: The Row Each Vendor Hopes You Skip
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Every model launch publishes a card of numbers, and every card has a row the vendor would rather you did not linger on. For Qwen3.8 Max, released on 3 August 2026, the flattering rows are easy to find: it took first place on the CodeArena front-end leaderboard at 1,691 points, it reached an Artificial Analysis Agentic Index of 58 to tie Claude Opus 5 at high effort — the closest a Chinese lab had come to the top of that board — and it landed a GDPval-AA score of 1,739 Elo, ahead of GPT-5.6 Sol. The row underneath them is harder to read. On the same evaluator's suite, measured hallucination rose from 23% on the previous generation to 40%, and the model's long-context retrieval score fell two points. Boreal, Creatify's advertising video model, launched 15 September 2026 at one cent per second, has a row of the same kind on its own card — and it is more specific, because the vendor published it.
Neither row is disqualifying. Both are more informative than the headline scores, which is why it is worth putting them side by side rather than comparing the banners.
What Qwen3.8 Max is genuinely best at
The wins are real and they are not small.
Qwen3.8 Max is a 2.4-trillion-parameter mixture-of-experts model with roughly 95 billion parameters active per token, and it is the first Max-class Qwen that Alibaba has said will be released as open weights — announced for the week of 10 August 2026 without a licence or a firm date, a detail worth noting because the announcement is not the release. It carries a 1M-token context window with a 131,072-token output ceiling, and it accepts text, images and video as input. That last point deserves emphasis in this particular comparison: of the four frontier text models this series sets against Boreal, Qwen3.8 Max is one of only two that can be handed a finished video clip and asked a question about it.
The pricing is aggressive in a way that reshaped the segment. At $2.00 per million input and $6.00 per million output, with cache reads at $0.25, it undercuts the previous generation by about 20% while doubling the maximum output length. On our board that lands at $2.00 and $6.00, a 1M-token context, a p50 time to first token of 10.00 seconds — the ceiling our instrumentation reports, so read it as "slow" rather than precise — and 183.0 million tokens of traffic over the last seven days.
And the row underneath
Here is the part that does not make the launch post's headline.
Artificial Analysis's independent evaluation recorded two regressions between Qwen3.7-Max and Qwen3.8 Max. Its long-context retrieval measure fell two points. Its omniscience measure fell ten, and the hallucination rate the evaluator measures rose from 23% to 40%. At the same time, the model's cost per task on the intelligence index rose from $0.53 to $1.14 — it took roughly 64 turns per task where its predecessor took 14.
Read those together and a coherent picture appears. Qwen3.8 Max is a model that got better at the things benchmarks with a single right answer can measure — code, agentic task completion, structured problem solving — and got worse at the thing that is hardest to score: knowing when it does not know. A model that deliberates more and hallucinates more is not contradicting itself. Longer reasoning traces create more opportunities to commit to a confident wrong answer, and a benchmark that rewards task completion does not penalise that until the task has a hidden invariant to violate.
For a buyer, the practical consequence is narrow and specific. If your use of the model is code generation or an agentic workflow where a wrong answer fails loudly — a test fails, a build breaks — the regressions are largely invisible and the gains are the whole story. If your use is retrieval, summarisation, or anything where a fluent wrong answer reaches a human without a check, the 23-to-40 move is the number that should decide your evaluation, not the CodeArena placement.
Boreal's own worst row, published by the vendor
The contrast that makes this pairing useful is that Creatify did something most vendors do not: it put its weakest result in the same post as its strongest.
Boreal was tested blind against its own untouched base model, with prompt, resolution, frame count and seed held fixed, across 40 production cases. Creatify reports an 81% preference for Boreal — 25 to 6 with 9 ties, sign test p<0.001 — and then breaks that average down. Product ads: 83%. Creator and UGC scenes: 85%. Single-person talking clips: 60%.
That last figure is the row. Talking heads are among the most common formats in direct-response advertising, and it is the format where the model's advantage over its own base is thinnest — barely better than a coin flip against a model nobody was going to ship. Rendering is quoted at 1:1 with realtime at 720p-class, and fine-detail retention improved 11.3% at p=0.004, so the aggregate story is genuinely positive. The breakdown simply tells you which half of the category this model was tuned for, and it is not the half with a person talking to camera.
Note also what kind of evidence each of these rows is. Qwen3.8 Max's regression was found by an independent evaluator running its own harness. Boreal's weak row was disclosed by the vendor, in a test the vendor designed and ran. The second is more trustworthy as a statement of fact and less informative as a comparison, because there is no independent number anywhere in Boreal's file.

The spec contrast
• Output — Boreal: rendered video, 720p-class, text-to-video and image-to-video. Qwen3.8 Max: text only, up to 131,072 tokens.
• Input — Boreal: a text prompt or a still image. Qwen3.8 Max: text, images and video.
• Price — Boreal: $0.01 per second of finished video. Qwen3.8 Max: $2.00 per 1M input / $6.00 per 1M output, cache reads at $0.25.
• Context — Boreal: not published. Qwen3.8 Max: 1M tokens in, 131K out.
• Latency — Boreal: quoted at 1:1 with realtime. Qwen3.8 Max: p50 time to first token of 10.00 seconds on our board.
• Weights — Boreal: proprietary, owned and served by Creatify. Qwen3.8 Max: proprietary at launch; Alibaba has announced an open-weights release for the first Max-class Qwen, with no licence or date confirmed.
• Evidence — Boreal: vendor-run 40-case blind test, no independent evaluation. Qwen3.8 Max: independent benchmark coverage including two measured regressions, alongside vendor rows of 86.1 on OSWorld-Verified and 86.6 on Terminal-Bench 2.1.
What a regression is worth to someone buying
Regressions on leaderboards are usually treated as trivia. They are closer to the most decision-relevant thing a benchmark publishes, because they are the only rows that tell you what a vendor traded away.
The useful move is not to read either card, but to run both models on your own traffic — and the practical obstacle is that this normally means two contracts, two SDKs and two billing relationships, which is enough friction that most teams skip the test and buy on the headline score instead.
That friction is the part a routing layer removes. Qwen3.8 Max is in our catalogue alongside the previous generation, so comparing them on your own evaluation set is a one-line change to the model string rather than a procurement cycle, and the same key reaches the rest of the catalogue when the comparison says you want a different model for a different stage of the pipeline. It sits at provider list rate with 0% markup, which matters here for a specific reason: Qwen3.8 Max arrived about 20% cheaper than the generation it replaced, and a vendor repricing of that kind is the sort of thing that should land on your bill when it happens rather than at the next renewal.
Boreal is not in our catalogue. OrcaRouter does not serve video generation models, and nothing here should be read as claiming otherwise. What we serve is the layer above it — the copy, the variants, the compliance pass, the analysis of what performed — and two of the models in this series, Qwen3.8 Max among them, can be handed the finished clip to review.

How to read either card
Qwen3.8 Max is the stronger buy if your work is code, agents or structured reasoning at a price that undercut its own predecessor, and the two independent regressions are the reason to test it on your retrieval and summarisation traffic before you route production there. The 40% hallucination figure is not a reason to avoid the model; it is a reason to know which of your workloads can tolerate it.
Boreal is the only one of the two that produces the artifact this comparison is nominally about, and it is the cheapest credible option on published rates for short-form advertising video. Its constraint is format-specific and stated by its own vendor: 60% preference on single-person talking clips. Test that format before the product ads, because that is where the headroom is smallest.
Two things would move this. If Alibaba ships the open weights it announced for Qwen3.8 Max, the pricing and the durability of the model both change, and the licence it arrives under will matter more than the date. And for Boreal, the first independent evaluation — of any kind, by anyone — would replace a 40-case vendor test that currently carries the entire evidence load for a model selling into a format where brands are unusually sensitive to how their product looks.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
