A generated title card reading 'AREX-2 vs Qwen3.8-Max - A Substrate and What Runs On It', with two panels. The left panel, headed Qwen3.8-Max, lists: live since 3 August 2026; 1M-token context, multimodal; $2 in / $6 out per 1M; callable today. The right panel, headed AREX-2, lists: repo created 29 Sep 2026; no weights, no model card; no rate card anywhere; not callable anywhere. A footer reads 'The first AREX generation was post-trained on a Qwen backbone.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

AREX-2 vs Qwen 3.8: One Is a Substrate the Other Is Probably Built On

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The matchup between AREX-2 and Qwen3.8-Max looks like a straight fight between two Chinese labs' flagship systems, and it is not one — not because AREX-2 is unproven, though it is, but because the two sit at different layers of the same stack. Qwen3.8-Max has been live since August 3, 2026, at $2 per million input tokens and $6 per million output with a 1M-token context window; its open-weight siblings, Qwen3.8-27B and Qwen3.8-Flash, shipped on August 13 and August 26 with published benchmarks and published prices. AREX-2 is a repository BAAI created on Hugging Face on September 29, 2026 at 17:56 UTC, containing a .gitattributes and nothing else. And the reason the two are not rivals is in the underlying fact that neither vendor advertises loudly: every AREX model BAAI has shipped so far is built on a Qwe​n backbone.

That is not a speculation about AREX-2. It is documented for the generation that exists. AREX-Base is a 122-billion-parameter mixture-of-experts with 10 billion active parameters built on Qwen3.5-122B-A10B, and AREX-Turbo is a dense 4B built on Qwen3.5-4B; both were released on July 23, 2026 under Apache 2.0. So the honest shape of this comparison is not agent-versus-flagship. It is: what does the Alibaba substrate buy you today, what does BAAI's agent layer add on top, and how much of the second question can be answered before BAAI uploads a file?

What is live on the Qwen side, and what it costs

The Qwen3.8 generation is unusually easy to reason about because it is fully specified and priced in public, at three distinct tiers.

• Qwen3.8-Max — released August 3, 2026. A 1M-token context window, native text, image and video input, thinking mode, function calling and structured outputs, and three wire formats (OpenAI-compatible, Anthropic-compatible and native DashScope) across five regions. $2.00 per million input tokens and $6.00 per million output, with cache reads at $0.25.

• Qwen3.8-27B — released August 13, 2026. Dense, 27B parameters, 256K context (262,144 positions), text, image and video in, Apache 2.0 weights. Published on Qwe​n's own card at 90.3 on LiveCodeBench v6, 89.2 on GPQA Diamond, 84.3 on OSWorld-Verified and 79.0 on QwenSWEBench — vendor-reported, not independently reproduced. Independent measurement puts it at an Artificial Analysis Intelligence Index of 33.7 and 68.1 on the AA Coding index.

• Qwen3.8-Flash — released August 26, 2026. Same 1M-token window and the same multimodal input, at $0.15 per million input tokens and $0.47 output. The cheap tier of the family, and the one that makes high-volume agentic traffic affordable.

There is also an untagged Qwen3.8 entry: the 2.4-trillion-parameter flagship whose weights Alibaba has said are coming, and which is not callable anywhere yet. If you are searching for "Qwe​n 3.8" and expecting one model, that is why the answer is a family of four, not one.

A screenshot of the OrcaRouter model page for Qwen3.8-27B, showing the Featured badge, a 256K-token context window, text, image and video input with text output, badges for Vision, Tools, JSON and Reasoning, the provider line 'by Qwen - 2026-08-13', a pricing row of $0.33 input and $2.40 output per million tokens with measured latency and 120.2M tokens of traffic, and the beginning of the model description covering the 27B dense multimodal model's Apache-2.0 licence and 256K context.

Against all of that, AREX-2's specification is one line long: a repository, a timestamp, and a name that implies a second generation of something BAAI already built once.

The detail that reframes the whole comparison

AREX is not a competitor to Qwe​n; it is a consumer of it. Both first-generation AREX models are post-trained on Qwen3.5 backbones and then wrapped in the framework that makes them agents rather than chat models: an inner loop that searches, browses and integrates evidence into a candidate answer carrying a confidence figure, and an outer loop that checks that answer against the original constraints and either accepts it, refines it, or discards the trajectory and starts again. The interesting engineering in AREX is not the network. It is the loop, the context-update mechanism that keeps verified findings, open candidates and unresolved constraints straight across a long horizon, and the tool interface that exposes search and browsing to the model as first-class actions.

That is why the family's own published numbers — 82.5 on BrowseComp, 85.4 on GAIA, 71.0 on xbench-2510, 89.9 on DeepSearch QA and 82.0 on WideSearch-en for the 122B Base, with 70.7 / 81.6 / 57.0 / 78.5 / 68.5 for the 4B Turbo — are not comparable to Qwen3.8-27B's coding and agentic scores, even though the two share an architectural lineage. They measure different things. QwenSWEBench asks whether a model can resolve a repository issue. BrowseComp asks whether an agent can find a fact that is deliberately hard to find, across many searches, without losing the question. A raw Qwen3.5 checkpoint scores badly on the second and well on the first, which is precisely the gap BAAI's post-training and harness exist to close.

A generated pricing card headed 'The Qwen3.8 family is priced. AREX-2 is not callable.' It lists Qwen3.8-Max at $2.00 in / $6.00 out with a 1M context released 3 Aug 2026, Qwen3.8-27B at $0.33 in / $2.40 out with 256K context and open weights released 13 Aug 2026, Qwen3.8-Flash at $0.15 in / $0.47 out with a 1M context released 26 Aug 2026, Qwen3.5-122B-A10B at $0.115 in / $0.917 out as the backbone AREX-Base post-trains up to 128K, and a highlighted AREX-2 row reading 'no rate card' because there are no weights, card or licence tag. A footer reads 'Qwen3.8-27B scores 33.7 on the AA Intelligence Index; AREX-2 does not exist to score.' The OrcaRouter logo is composited in the bottom-right corner.

What a second generation would most plausibly be

If AREX-2 continues the pattern, the most likely reading — inference, not fact — is a second-generation research agent post-trained on a newer Qwe​n backbone than the 3.5 generation the current AREX models use. Qwen3.8-27B is dense, multimodal and Apache 2.0, which makes it a natural candidate for a quality-tier agent; Qwen3.8-Flash is cheap per token, which matters enormously when your agent makes dozens of calls per query and re-reads a growing context on every step.

None of that is confirmed. The name carries no size, no backbone and no date, and a "2" in a product line is a naming convention rather than a specification. What is confirmed is far narrower and should be stated that way: BAAI created the repository on 29 September 2026, put nothing in it, and has announced nothing. No weights, no card, no parameter count, no licence tag, no benchmarks, no provider serving it. If you see AREX-2 numbers quoted anywhere this week, they are not measurements.

What you can actually buy this afternoon

The practical asymmetry between these two is not quality, it is that one side has a rate card and the other has a directory entry.

If your work is agentic and you want the substrate the AREX family was built on, the exact backbone is callable: Qwen3.5-122B-A10B is on OrcaRouter's catalogue at $0.115 per million input tokens and $0.917 per million output up to 128K tokens, rising to $0.287 / $2.294 beyond that, with vision, tool use and reasoning enabled. That is the model AREX-Base post-trains, available without waiting for anything.

If you want the current generation, both open Qwen3.8 tiers are on the catalogue: Qwen3.8-27B at $0.33 per million input and $2.40 output, running on our own infrastructure because the weights are open and there is no vendor per-token cost to pass through, and Qwen3.8-Flash at $0.15 / $0.47 for the volume tier. One API covers both, provider list price passed through with nothing added per token, so when Alibaba moves a price the number here moves the same day.

And when AREX-2 does ship, the reason it belongs behind that same layer is structural rather than promotional. An agent loop that makes twenty calls per query multiplies every provider's failure rate by twenty; a routing rule with automatic failover that decides at runtime is the difference between a timeout being a retry and a timeout being a confidently wrong answer. That argument applies to any research agent, and it will apply to AREX-2 whenever there is an AREX-2 to route.

Who should act, and on what

If you need a research agent today, AREX-Base is the AREX model that exists, it shipped in July under Apache 2.0, and its agent benchmarks are the vendor's own but at least they are attached to a downloadable model. If you need frontier multimodal reasoning or high-volume agentic traffic today, the Qwen3.8 tiers are live, priced and independently scored in part, and nothing about AREX-2's existence changes that decision.

The only scenario in which AREX-2 matters to a decision you are making this week is the one where you are choosing a backbone for a fine-tune. And there the answer is already visible in the catalogue: the 3.5 backbones AREX used are still cheap and available, the 3.8 generation is better and also available, and a second-generation AREX — whenever it arrives — will almost certainly be evidence about which Qwe​n base BAAI thinks is worth building on, not a replacement for it.

Three things would settle the rest: files in the tree, a card with a parameter count and a base-model field, and a licence tag. The first of those is the one that matters, because until it lands, this comparison is between a priced family and a placeholder.