
Grok 4.7 vs Qwen 3.8 Max: Identical Price, One Point Apart, Opposite Shapes
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Two closed flagship models, both at $2 per million input tokens and $6 per million output, both scoring within a single point of each other on the same independent benchmark, released nineteen days apart. Grok 4.7 arrived on September 21, 2026 from SpaceXAI; Qwen 3.8 Max was released by the vendor on August 3, 2026 and refreshed on September 2, 2026. On the current Artificial Analysis Intelligence Index, version v4.3.2, Grok 4.7 scores 46 and Qwen 3.8 Max scores 45. On price and on measured capability, these two models are the same purchase. On everything else — context, modality, serving speed, openness, and what each vendor chose to optimize — they could not be more different, and that is what makes this comparison worth running rather than skipping.
First, the timeline, because Qwen 3.8 Max has two dates
This trips up a lot of coverage, so it is worth setting straight at the top. Qwen 3.8 Max launched on August 3, 2026 as Alibaba's largest and most capable model, built from the Qwen 3.5 vision-language architecture on a sparse mixture-of-experts design reported at 2.4 trillion total parameters with 95 billion active per token. On September 2, 2026 Alibaba shipped a refreshed build — the 0902 revision, which is the version carrying current model IDs and the version Artificial Analysis dates its page to. If you have seen both an August and a September date for Qwen 3.8 Max, that is why: one is the launch, the other is the revision most people are actually calling today.
Grok 4.7 has a cleaner history and a messier one behind it. SpaceXAI released it on September 21, 2026 after a launch that slipped repeatedly — Musk had pointed at windows in late August, early September and mid-September before the model actually shipped. The release itself was narrow: a larger base model than Grok 4.6, a longer reinforcement-learning run aimed at multi-hour tasks, and no change to the $2/$6 rate card that Grok 4.6 had carried since August 12.
What each vendor optimized for
Read the two launch announcements as a pair and the design bets are almost perfectly opposed.
SpaceXAI optimized for agentic execution. Grok 4.7's vendor-reported numbers are concentrated in terminal and repository work: Terminal-Bench 4.0 at 38.0% against Grok 4.6's 20.3%, CursorBench 4.0 at 46.3%, DeepSWE v1.1 at 71.0% at high reasoning effort, EEBench at 64.0%. The company's own description of the release names self-verification and long-context handling inside the Grok Bot environment as the targets. Grok 4.7 serves a 500,000-token window, takes text and images in, returns text, and exposes reasoning effort from low through xhigh. It is a model built to finish jobs.
Alibaba optimized for reach. Qwen 3.8 Max serves a context window of roughly 984,000 tokens — effectively a million — and its published strengths are breadth rather than depth in one lane: a Terminal Bench 2.1 score of 86.6, SWE-bench Pro at 67.7, PaperBench at 93.0, GPQA Diamond at 92.6, MMMU-Pro at 82.3, OSWorld-Verified at 86.1. It accepts text, image and video input and returns text, supports reasoning effort levels with xhigh as the default, and its weaker published spot is Humanity's Last Exam at 43.6. It is a model built to handle anything you hand it.
Every figure in that paragraph is vendor-reported. Neither set has been independently reproduced, and the two sets are not directly comparable anyway — Terminal-Bench 2.1 and Terminal-Bench 4.0 are different benchmarks, so Qwen's 86.6 and Grok's 38.0 must never be placed side by side.
• Independent composite — Grok 4.7 46 on AA Intelligence Index v4.3.2 vs Qwen 3.8 Max 45 on the same version. A statistical tie.
• Price per million tokens — $2.00 in / $6.00 out on both. Qwen's cache discount is listed at 88%, giving a blended rate of $1.18 per million; Grok 4.7's cache terms are not published on the same basis.
• Context window — Grok 4.7 500,000 tokens vs Qwen 3.8 Max approximately 984,000. Qwen wins by nearly double, and this is the single largest spec difference between them.
• Input modalities — Grok 4.7 text and image vs Qwen 3.8 Max text, image and video.
• Weights — both proprietary. Neither model can be downloaded, which removes the usual open-weights argument from this comparison entirely.
• Parameters — Grok 4.7 undisclosed on any SpaceXAI spec sheet (the ~2.1T figure is a Musk claim) vs Qwen 3.8 Max reported at 2.4T total with 95B active, which Alibaba has not confirmed on a specification sheet either.
• Serving speed — Qwen 3.8 Max at 38.8 output tokens per second with 3.08 seconds to first token, per Artificial Analysis; Grok 4.7 positioned by SpaceXAI as roughly twice as fast as comparable models, vendor-reported, with no independent throughput figure published yet.

The context difference is the real decision
When price and capability are identical, the decision falls to whatever else differs, and here it is context. Doubling the window from 500,000 to roughly 984,000 tokens is not a spec-sheet nicety — it is the difference between chunking a repository and holding it.
Qwen 3.8 Max's longer window comes with a documented caveat that matters for buyers: the open-weights version Alibaba announced for the week of August 10, 2026 was text-only and did not include the full million-token context. That does not affect the hosted endpoint this comparison is about, but it is worth knowing if you were planning around the open release.
There is a second consideration on the Grok side. The Grok line has historically applied a pricing cliff at 200,000 input tokens, where a request crossing the threshold reprices the entire request at double the rate — $4 in and $12 out. With a 500,000-token window that threshold sits well inside the model's working range. Before routing long-context work to Grok 4.7, confirm the current rate card for the deployed model, because the interaction between a large window and a repricing cliff is where long-context bills go wrong.
What a month of work costs on each
Because the headline rates are identical, the cost comparison has to be built from token behaviour rather than from the rate card, and there the two models diverge in a way that is measurable.
Artificial Analysis measured Grok 4.7 generating 240 million output tokens while running its Intelligence Index, against a median of 92 million across models on the board — it is a verbose reasoner. Qwen 3.8 Max generated 190 million output tokens on the same index, with a total evaluation cost of $4,934.79 and a measured 38.8 tokens per second. Both are well above the median verbosity, and Grok 4.7 is the more verbose of the two.
So at an identical $6 per million output, the model that emits more tokens per task costs more per task. The published verbosity figures suggest Grok 4.7 will consume more output tokens for equivalent index work, which means the tie on price is really a small win for Qwen 3.8 Max on cost-per-task — offset by Grok 4.7's claimed speed advantage, which shortens wall-clock time and therefore the human cost of waiting on an agent run. Neither of those offsets is visible on a rate card, and both are worth measuring on your own traffic rather than assuming.
The one asymmetry that is not in dispute is caching. Qwen 3.8 Max publishes an 88% cache discount and a $1.18 blended rate at a 7:2:1 cache-hit/input/output mix. For workloads with a large stable prefix — a system prompt, a codebase header, a document set — that discount is where the real savings live, and it is worth checking how Grok 4.7's cache terms compare before you commit volume.

Choosing, when the numbers refuse to choose for you
A 46 and a 45 on the same index, at the same price, is as close to a genuine tie as this blog is likely to publish. That means the decision should be made on the axes where the two models actually differ, and there are four of them.
Pick Grok 4.7 if your work is agentic coding and terminal execution, if wall-clock speed on long runs matters more than context headroom, and if you want a per-request reasoning dial to keep easy traffic cheap. Pick Qwen 3.8 Max if you routinely handle inputs larger than half a million tokens, if video input is on your roadmap, if you want the 88% cache discount for prefix-heavy workloads, or if you want a model whose vendor publishes a broad evaluation matrix rather than one concentrated in a single lane.
If the honest answer is "some of both," that is the normal answer, and it is the case this pair is built for. Qwen 3.8 Max is on OrcaRouter at Alibaba's list price, and OrcaRouter passes provider list pricing through with 0% markup, so the $2/$6 above is what you actually pay and a vendor rate change is live here the same day rather than at renewal. One API key covers Qwen 3.8 Max plus 200-plus other models, which means the context decision becomes a per-request routing rule instead of a platform commitment: send the oversized inputs to the 984k window and keep everything else on whichever endpoint is fastest and cheapest for that task, with automatic failover if a provider degrades. Where a task genuinely needs both shapes — a long-context read followed by an agentic execution pass — the routing DSL composes models into a single call rather than forcing you to pick a side.
One clarification worth stating, since neither model is open: nothing in this comparison involves downloadable weights. Both are hosted endpoints, and self-hosting is not an option for either.
The one number to keep
Forty-six against forty-five is not a reason to pick a model, and it should not be. What decides this comparison is the context window — 500,000 tokens against roughly 984,000 — and the shape each vendor chose: Grok 4.7 optimized for finishing hard agentic jobs fast, Qwen 3.8 Max optimized for handling whatever you hand it. Same price, same measured capability, different jobs. That is a routing decision, and it is the kind of decision that should take an afternoon to test rather than a quarter to procure.

Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
