Generated hero title card headlined 'GPT-6 Luna vs Qwen3.8-Max' with the subtitle 'The cheapest model here is not the one that watches video', a footer reading 'Prices per OpenAI and Alibaba Cloud; index figures per Artificial Analysis.', and the OrcaRouter logo composited bottom-right.
Guides & Insights

GPT-6 Luna vs Qwen3.8-Max: The Cheapest Model in This Comparison Is Also the Only One That Watches Video

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The obvious framing for this pairing is the wrong one. Qwen3.8-Max lists at $2.00 per million input tokens and $6.00 per million output with cache reads at $0.25. GPT-6 Luna lists at $0.10 and $0.50 with cache reads at a cent. Twenty times on input, twelve times on output, fifty times on cached input — a price gap so wide that the comparison looks settled before it starts.

It is not settled, because the two models are not competing for the same request. Qwen3.8-Max accepts text, image and video, and its 131,072-token output ceiling is slightly above GPT-6 Luna's 128,000. GPT-6 Luna accepts text and image only. Alibaba's model scores 58 on the independent Intelligence Index — first in its class on the agentic board — and GPT-6 Luna scores 37 at maximum effort, 29 at its default. One is a flagship with an open-weight lineage and a video input modality. The other is a volume tier with a speed figure and a one-cent cache rate. The price gap is not a discount on the same thing; it is a description of two different things.

What each of these two is

Qwen3.8-Max is Alibaba's flagship of the 3.8 generation, a Max-class checkpoint whose underlying model — Qwen3.8-2.4T-A95B — was released publicly under a custom licence on August 12, 2026, making it the first Max-class Qwen to have open weights at all. The hosted flagship itself is still listed by the independent board as proprietary. It carries a 1,000,000-token context window, a 131,072-token output ceiling, text, image and video in, and reasoning effort selectable at low, medium and extra-high. A pinned Qwen3.8-Max-0902 revision is available at the same price, which is a small detail with real operational value.

GPT-6 Luna is OpenAI's efficiency tier from September 22, 2026, sitting beneath GPT-6 Sol at $2.00/$10.00 and the flagship GPT-6 Astra at $10.00/$50.00. Text and image in, text out, a 1,050,000-token window with 922,000 maximum input, a 128,000-token output ceiling, and effort from none through low, medium, high, xhigh and max with medium as the default. Closed, single-identifier, available to anyone with an API key.

One accepts video and can be downloaded in its base form. One accepts images and is fast. Both are priced to be used at volume. The overlap between those two statements is smaller than the price ratio suggests.

The cards, line by line

One line per dimension, both sides on each:

• Input per 1M — GPT-6 Luna $0.10 vs Qwen3.8-Max $2.00

• Output per 1M — GPT-6 Luna $0.50 vs Qwen3.8-Max $6.00

• Cached input per 1M — GPT-6 Luna $0.01 vs Qwen3.8-Max $0.25 for implicit cache reads, with an explicit cache read at $0.17 and an explicit cache write at $2.50

• Long-context clause — GPT-6 Luna doubles input and cache and lifts output 1.5× above 272,000 input tokens, applied to the whole request vs Qwen3.8-Max a flat tier across the full window, which is the one place the Qwen card is structurally better rather than marginally cheaper

• Context window — GPT-6 Luna 1,050,000 tokens vs Qwen3.8-Max 1,000,000 tokens

• Maximum output — GPT-6 Luna 128,000 tokens vs Qwen3.8-Max 131,072 tokens

• Input modalities — GPT-6 Luna text and image vs Qwen3.8-Max text, image and video

• Reasoning effort — GPT-6 Luna none, low, medium (default), high, xhigh, max vs Qwen3.8-Max low, medium and extra-high

• Intelligence Index, independent — GPT-6 Luna 37 at max effort vs Qwen3.8-Max 58

• Agentic Index, independent — Qwen3.8-Max 58, first in its class; no comparable GPT-6 Luna figure published

• Default-effort score — GPT-6 Luna 29 at its default medium vs Qwen3.8-Max measured at its maximum setting

• Cost per Index task, independent — GPT-6 Luna about $0.07 vs Qwen3.8-Max $5.41

• GDPval, independent — Qwen3.8-Max 1,739 Elo at 64 turns per task, which works out to roughly $1.14 per completed task on that evaluation; no comparable GPT-6 Luna run is published

• Hallucination rate on AA-Omniscience — Qwen3.8-Max 40%, a regression from 23% on the previous generation; GPT-6 Luna 77%, down from 93%

• Dated snapshot — GPT-6 Luna a single rolling identifier vs Qwen3.8-Max-0902 available as a pinned September 2, 2026 revision at the same price

• Weights — GPT-6 Luna closed, API only vs Qwen3.8-Max the base Qwen3.8-2.4T-A95B checkpoint public under a custom licence since August 12, 2026, while the hosted flagship is listed as proprietary

• On OrcaRouter — GPT-6 Luna not in our catalogue vs Qwen3.8-Max routable at Alibaba's list price

Generated two-column scoreboard titled 'GPT-6 Luna vs Qwen3.8-Max - the scoreboard'. Left column GPT-6 Luna rows: Price in/out $0.10 / $0.50, Cached input $0.01, AA Index 37 at max effort, Context 1,050,000, Input text + image, Cost per index task $0.07. Right column Qwen3.8-Max rows: Price in/out $2.00 / $6.00, Cached input $0.25, AA Index 58, Context 1,000,000, Input text + image + video, Cost per index task $5.41. Footer: 'Prices per OpenAI and Alibaba Cloud; index figures per Artificial Analysis.'

Video is not a feature line, it is a different workload

The modality row is the one that decides more real comparisons than the price row, and it is usually read as a checkbox.

Qwen3.8-Max accepts video. GPT-6 Luna does not. If your pipeline needs to reason about a screen recording, a product demo, a lecture capture, a security camera segment or a UI walkthrough, the comparison ends at that line and the twentyfold price difference becomes irrelevant — you are not choosing between an expensive option and a cheap one, you are choosing between an option and no option. Extracting frames and passing them as images is a workaround, not an equivalent, and anyone who has built one knows the difference in both cost and quality.

If your pipeline is text and images, the modality line is silent and the price line speaks. That is the honest split, and it is worth being blunt about which side of it you are on before reading anything else on this page.

The agentic gap is real and it is the expensive half of the price difference

Qwen3.8-Max scores 58 on the independent Intelligence Index. GPT-6 Luna scores 37 at maximum effort and 29 at its default medium. Twenty-one points at matched settings, twenty-nine at defaults — on a composite where a single point is inside the noise of a re-run, a gap of that size is not noise.

Qwen3.8-Max's position is concentrated in agentic work. It scores 58 on the independent Agentic Index, first in its class, and 1,739 Elo on GDPval at 64 turns per task. That last figure is the informative one, because it is a measurement of a model holding a task together across sixty-four sequential decisions, and it comes with a price tag attached: roughly $1.14 per completed task on that evaluation. Compare that with GPT-6 Luna's cost per index task of about $0.07, and the honest statement is not that Qwen is expensive. It is that Qwen is being asked to do a much harder thing, and doing it.

The corollary is that GPT-6 Luna's twenty-one-point deficit is not distributed evenly across your workload either. On a short, well-specified, program-consumed task — extract this field, classify this ticket, route this request — both models will produce the same usable answer nearly all of the time, and the index gap will be invisible in your eval. On a task requiring sixty-four sequential actions with a checkpoint at the end, the gap is the difference between finishing and quietly stopping halfway, and that is the task Qwen3.8-Max was built for and GPT-6 Luna was not.

One number cuts the other way and deserves to be stated rather than buried. Qwen3.8-Max's hallucination rate on the omniscience board is 40%, up from 23% on the previous generation — a substantial regression, and the kind that shows up as confident wrong answers in production rather than as a benchmark line. GPT-6 Luna's is 77%, down from 93%. Both figures are high in absolute terms, and the cheaper model is still the worse of the two on this measure. Neither number is a reason to prefer one model; both are reasons to keep a human or a verifier in the loop on anything where a fabricated fact is expensive.

The long-context clause is the one line where Qwen wins on structure

GPT-6 Luna carries a clause that Qwen3.8-Max does not: requests above 272,000 input tokens are billed at twice the input and cache rates and 1.5× the output rate, applied to the entire request rather than the portion over the line. A 300,000-token call on GPT-6 Luna bills at $0.20 per million input for all 300,000 tokens because it went over by 28,000. Split the same document into two calls and the cost halves with no change to the prompt.

Qwen3.8-Max is flat across its full 1,000,000-token window. For genuinely enormous single requests, that makes the twelvefold output-price gap smaller than it looks and can invert it: above 272,000 input tokens GPT-6 Luna's effective input rate doubles to $0.20, and its output rate rises to $0.75, against Qwen's flat $2.00 and $6.00. The gap narrows from twenty times to ten on input and from twelve to eight on output. Still wide — but the workloads most likely to have a 300,000-token working context are exactly the agentic ones both vendors are selling into, and the teams running them are the teams that will notice.

The other structural line is the pinned revision. Qwen3.8-Max-0902 is available at the same price as the rolling identifier, which means a team that needs reproducible behaviour across a model update can have it without paying a premium. GPT-6 Luna offers no equivalent. That is a small line on a rate card and a large line in a post-mortem.

Running one of them, or both

Qwen3.8-Max is on OrcaRouter at Alibaba's list price, under the pass-through pricing that means a vendor rate change is live on our side the same day. Two things about our routing layer earn their place for this particular model: automatic failover, because a hosted flagship with a published open-weight base has more upstreams than a closed single-vendor model and more ways for one of them to degrade; and the pinned-revision handling, which lets a route hold Qwen3.8-Max-0902 explicitly rather than inheriting whatever the rolling identifier resolves to next week.

GPT-6 Luna is not in our catalogue as of this writing and is reached through OpenAI's own API, so a team that wants both runs one key for Qwen3.8-Max alongside whatever it already uses for OpenAI. This is also the pairing where the routing DSL is worth more than a static split, because the boundary between the two models is a modality check and a length check rather than a preference: requests carrying video go to Qwen3.8-Max because nothing else can serve them, requests above 272,000 input tokens go to Qwen3.8-Max because the other model's cliff makes them more expensive there, and everything else goes to the model that is a twentieth of the price. That is a rule, and it belongs in configuration rather than in an application branch.

Screenshot of OpenAI's developer documentation page for GPT-6 Luna, showing the model selector, the description 'Our most efficient model for focused, high-volume tasks', reasoning set to High, speed Fast, price $0.1 · $0.5, text and image input with text output, a 1,050,000-token context window, 128,000 maximum output tokens, a May 18, 2026 knowledge cutoff, and pricing cards of $0.10 input, $0.01 cached input, $0.125 cache writes and $0.50 output per 1M tokens.

Which one, and when

Take Qwen3.8-Max if any of the following is true. Your pipeline needs video in. Your requests routinely exceed 272,000 input tokens, where its flat tier beats GPT-6 Luna's repricing. Your output exceeds 128,000 tokens. Your work is long-horizon agentic execution with a verifiable endpoint, where its 58 on the agentic board and 1,739 Elo on GDPval at 64 turns per task are the evidence that matters. Or you need a pinned revision for reproducibility and are willing to pay for it — which, at the same price as the rolling identifier, means you are not.

Take GPT-6 Luna if none of those apply and your workload is high-volume, program-consumed, text-and-image, and short. Classification, extraction, routing, bulk summarisation, the inner loop of an agent that needs a fast answer rather than a careful one. At $0.10/$0.50 with a one-cent cached read and a cost per completed task around a fifteenth of Qwen3.8-Max's, the arithmetic is not close. Set the effort parameter explicitly — the default is medium and the headline 37 is a maximum-effort figure — and keep long calls on the right side of the 272,000-token line.

The mistake this comparison invites is reading the twentyfold input price as a verdict. It is a verdict on one workload shape, and it is silent on the three lines that decide the other one: whether the request carries video, whether it clears 272,000 tokens, and whether the answer has to survive sixty-four sequential steps.

Screenshot of the OrcaRouter model page for Qwen3.8 Max (0902) showing the breadcrumb Home > Models > Qwen > Qwen3.8 Max (0902), a NEW badge, the model id qwen/qwen3.8-max-0902, Vision, Tools, JSON and Reasoning chips, a Qwen attribution dated 2026-09-02, the description of a September 2, 2026 snapshot accepting text, image and video input with a 1M-token context, input pricing of $2.00 and output of $6.00 per 1M tokens, a p50 time to first token of 3.50s, a p95 of 9.38s, and traffic of 196.6M tokens over seven days.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily