A generated title card reading "Ember-1 vs Qwen3.8-Max" with the subhead "A token-thrifty derivative against the flagship", above two cards: Ember-1 — "40% fewer tokens claimed" and "no published price"; Qwen3.8-Max — "2 dollars in, 6 dollars out" and "1M-token context".
Guides & Insights

Ember-1 vs Qwen3.8-Max: The Token-Thrifty Derivative Against the Flagship

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The two models in this matchup both have a real number on the same benchmark, and it still does not settle anything. Ember-1 is Fireworks Research's specialized derivative of Moonshot AI's Kimi K3, published 23 September 2026, reporting 82.0% on Terminal Bench 2.1 with roughly half the tokens its base model spends. Qwen3.8-Max is Aliba​ba's flagship — a sparse mixture-of-experts model of roughly 2.4 trillion parameters with about 95 billion active per token — generally available since 3 August 2026, reporting 86.6% on the same benchmark. Both those figures are vendor-reported, both labs ran their own harness, and a 4.6-point gap between two unpublished evaluation setups is not a verdict. What is comparable is the money, and that is where this matchup gets interesting: one model's entire thesis is that you should not pay for tokens you do not need, and the other is a flagship priced at $2 per million in and $6 per million out.

What each model actually is

Ember-1 is not a new model in any architectural sense. It is Kimi K3 retrained to reason more concisely, built by Fireworks Research over more than 50 training experiments and over 200 evaluations on its own serverless training stack. The lab's claim is that K3's reasoning is longer than tasks require and that the excess can be removed without changing answers — reasoning length fell by 35–50% without accuracy loss across seven benchmarks and two customer production A/B tests. It is available as a research preview on the vendor's own serverless platform, with no published weights and no published price.

Qwen3.8-Max is a different kind of object entirely. It is Aliba​ba's frontier tier, natively multimodal for text, image and video input, carrying a one-million-token context window, and refreshed twice since launch: a post-trained update on 2 September 2026 aimed at coding and professional office work, and a faster serving tier announced on 22 September 2026. Aliba​ba positions it against GPT-5.5, Claude Opus 4.7 and Gemini 3.1 Pro, and its published evaluation sheet reflects that ambition — GPQA Diamond 92.6, OSWorld-Verified 86.1, SWE-bench Pro 67.7, PaperBench 93.0, IFBench 82.8 and Humanity's Last Exam at 43.6, all vendor-reported.

A two-column scoreboard titled "Ember-1 vs Qwen3.8-Max — the scoreboard" comparing six dimensions. Ember-1: input/output price not published, computed from Kimi K3 rates of $3 and $15; context not published, base model carries 1M; text input; Terminal Bench 2.1 82.0% vendor-reported; research preview with a two-week window; not routed on OrcaRouter. Qwen3.8-Max: $2.00 in and $6.00 out per 1M tokens; 1,000,000-token context; text, image and video input; Terminal Bench 2.1 86.6% vendor-reported; generally available and priced; routed on OrcaRouter. The footer notes both Terminal Bench figures are vendor-reported and were produced on different harnesses.

The rate cards, side by side

One of these has a price and the other does not, which is the first practical asymmetry.

• Input — Ember-1: none published; the benchmark sheet computes dollars from Kimi K3's rate card. Qwen3.8-Max: $2.00 per million tokens on OrcaRouter.

• Output — Ember-1: none published. Qwen3.8-Max: $6.00 per million tokens.

• Cached input — Ember-1: not applicable. Qwen3.8-Max: $0.25 per million tokens for cache reads, which is an eighth of the input rate and matters enormously in agent loops where the same context is resent every turn.

• Context window — Ember-1: not published for the preview; its base model carries 1,048,576 tokens. Qwen3.8-Max: 1,000,000 tokens.

• Multimodality — Ember-1: text. Qwen3.8-Max: text, image and video in, text out.

• Weights — Ember-1: none. Qwen3.8-Max: an open-weights release exists but is text-only and lacks the full context window and vision input of the served endpoint.

• Access — Ember-1: research preview, two-week serverless window, permanence tied to demand. Qwen3.8-Max: generally available and priced.

Note one thing about the Qwen3.8-Max rates before modelling anything: Aliba​ba has revised this model's packaging repeatedly since launch, and the international rate card has not always matched the August headline. Treat $2.00 and $6.00 as the current route price on our platform — which passes provider list price through with no markup, so a vendor change shows up the same day — and check the card before you commit a budget to it.

Why the benchmark comparison does not work

Ember-1's 82.0% and Qwen3.8-Max's 86.6% are both Terminal Bench 2.1, which makes them look comparable and does not make them comparable. Ember-1's sheet reports its number against Kimi K3 variants evaluated by Fireworks Research, on 89 samples, with token and cost deltas attached. Qwen3.8-Max's number comes from Aliba​ba's own evaluation sheet. Different labs, different harnesses, different agent scaffolds, and in a benchmark whose scores move several points on scaffold choice alone, a 4.6-point gap is inside the noise floor of the methodology.

What you can read from Ember-1's sheet is the token column, which is the only thing the model was trained to change. Against its own base, Ember-1 spends 51.9% fewer tokens on Terminal Bench 2.1, 32.5% fewer on SWE-Interact, 23.7% fewer on DeepSWE 1.1 and just 5.9% fewer on τ-2 Bench Airline. The spread is the honest headline. A model that cuts deliberation is worth a great deal on benchmarks where the base model over-thinks and almost nothing where it does not, and you cannot know which kind of workload you have without measuring your own.

Cost per completed task, worked through

Here is the arithmetic that actually decides this, using Kimi K3's published rates as the stand-in for Ember-1's cost, since Ember-1 has none. Kimi K3 is $3.00 per million input tokens, $0.30 cached and $15.00 output — exactly the rate card Fireworks Research used in its own comparison. Qwen3.8-Max is $2.00 in and $6.00 out, with cache reads at $0.25.

On the input side Qwen3.8-Max is already cheaper, by a third. On the output side it is 2.5 times cheaper per token. That means Ember-1's token reduction is not competing against a static bill — it is competing against a flagship that is cheaper per token before any efficiency work happens. Fireworks Research's own production A/B test showed output tokens falling from 49.3K to 29.9K, a 39% total reduction; apply that to Qwen3.8-Max's output rate and you get the same 39% saving, which is a smaller absolute win than the same percentage applied to K3's $15 rate.

The efficiency case for Ember-1 is therefore strongest when measured against its own base model, and it is exactly the case the release makes. Against Qwen3.8-Max the comparison shifts to capability per dollar, where the flagship's larger context, multimodal input and priced availability are the counter-arguments — and where nobody has published a head-to-head that would settle it.

A screenshot of OrcaRouter's model page for Qwen3.8 Max, showing the qwen/qwen3.8-max listing priced at $2.00 per 1M input tokens and $6.00 per 1M output tokens, a p50 time-to-first-token of 1.98s, 33.3M tokens of traffic over seven days, a 1M-token context window accepting text, image and video, and a Python snippet calling api.orcarouter.ai/v1.

What our own routing data shows

Two of the three models involved here are live routes on OrcaRouter, which gives a view no benchmark sheet contains. Qwen3.8-Max is served at 1,000,000 tokens of context with a median first-token latency around 2.0 seconds and roughly 52.7 output tokens per second over the last seven days, with an error rate of about 4.2%. Kimi K3 — Ember-1's base model, and the reference point for every saving Ember-1 claims — runs at 1,048,576 tokens of context, a median latency near 8.0 seconds, about 43.6 output tokens per second and an error rate of roughly 0.27%, on materially higher traffic. Ember-1 itself is not on OrcaRouter.

Those numbers reframe the decision. Kimi K3 is substantially slower to first token and roughly twice as expensive per output token as Qwen3.8-Max, while carrying a much lower error rate under far heavier load. A token-thrifty derivative of K3 narrows the cost gap but does not close the latency gap, because shortening reasoning shortens the generation phase, not the queue. If your workload is latency-sensitive rather than bill-sensitive, the flagship is the better route today and the efficiency story is beside the point.

Which one to route

Route Qwen3.8-Max when you need the context window, the multimodal input, a published price and a service you can put a budget against. It is the safer of the two by construction: generally available, priced, and refreshed twice in two months by a vendor with a track record of keeping the endpoint current.

Try Ember-1 when your bill is dominated by reasoning tokens in long agent loops and you run against a hosted model rather than your own hardware. The right test is the two weeks the preview gives you, run in shadow against production traffic, measuring tokens per completed task rather than tokens per request. What you should not do is swap it in as a like-for-like replacement for a flagship on the strength of a 40% figure that ranges from 6% to 52% depending on the workload.

If the real question is whether to keep running Kimi K3 at all, that one you can answer today without any preview: both Kimi K3 and Qwen3.8-Max are on OrcaRouter behind one key, priced at provider list with no markup, with automatic failover between providers. Comparing your workload's cost on both is a routing change, not a procurement project — and it gives you the baseline that makes any efficiency claim about Ember-1 measurable in your own numbers rather than the lab's.

A screenshot of OrcaRouter's model page for Kimi K3, showing the MoonshotAI Kimi K3 listing priced at $3.00 per 1M input tokens and $15.00 per 1M output tokens, a p50 time-to-first-token of 8.00s, 749.2M tokens of traffic over seven days, a 1M-token context window and a Python snippet calling api.orcarouter.ai/v1.

The honest summary is that these two models are optimized against different constraints. Qwen3.8-Max is built to be the most capable thing available and priced accordingly; Ember-1 is built to make an already-expensive model cheaper to run. Both are legitimate, and the reason this matchup resists a clean verdict is that the derivative's savings are defined relative to a rate card the flagship materially undercuts.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily