
Grok 4.7 vs Claude Opus 5: The Price Gap Is Real, the Parity Isn't
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
You can call Grok 4.7 for $2.00 per million input tokens and $6.00 per million output tokens. You can call Claude Opus 5 for $5.00 and $25.00. That is a 4.2x difference on output — the kind of gap that usually settles a comparison before the benchmarks are even opened. It does not settle this one. On the version of the Artificial Analysis Intelligence Index that both models are currently scored against — v4.3.2, the revision published on September 19, 2026 — Claude Opus 5 sits at 51 and Grok 4.7, released by the vendor on September 21, 2026, sits at 46. Five points is not a rout, but the shape of the five points is: Opus 5 wins eight of the ten evaluations in that index. The cheaper model's case is price, context is a wash, and the one place Grok 4.7's price advantage was supposed to matter most is exactly where it goes away.
Two flagships, two very different price sheets
Claude Opus 5 has been in market since July 24, 2026 and is Anthropic's Opus-tier flagship, sitting below Claude Fable 5.1 in the company's own lineup. It carries a 1M-token context window at flat pricing — Anthropic states plainly that a 900k-token request bills at the same per-token rate as a 9k-token one, with no long-context premium at all — and a 128K maximum output, extendable to 300K on the Message Batches API. Thinking is adaptive and on by default, with an effort ladder running low, medium, high, xhigh and max, defaulting to high. Weights are closed.
Grok 4.7 is a day old. It holds a 500k-token context, takes text and images in and returns text, and runs a reasoning-effort dial of its own. Its list price is $2.00 input and $6.00 output per million tokens below 200k prompt tokens, stepping to $4.00 and $12.00 at or above that threshold; cached reads are $0.50 and $1.00 in the same two bands. xAI also lists a faster variant running at roughly twice the output speed for roughly twice the price, though that configuration shipped without a published speed measurement attached. Weights are closed here too, which removes the "run it yourself" escape hatch from both sides of this comparison.
The independent board is not close, and it is not flattering
Both of these models are scored on Artificial Analysis's v4.3.2 index, which incorporates ten evaluations covering agents, coding, general knowledge and scientific reasoning. Putting the two columns side by side makes the five-point headline look generous to the cheaper model — a single five-point gap in a composite can hide a lot, and here it hides a near-sweep.
• Composite intelligence — Grok 4.7 (xhigh): 46 on Artificial Analysis Intelligence Index v4.3.2. Claude Opus 5 (adaptive reasoning, max effort): 51 on the same revision.
• Hard reasoning — Grok 4.7: Humanity's Last Exam 43, CritPt 18. Claude Opus 5: HLE 55, CritPt 29. A twelve-point and an eleven-point gap respectively, both in Opus 5's favour.
• Agentic terminals — Grok 4.7: Terminal-Bench 4.0 26. Claude Opus 5: 49. This is the widest single gap on the board, and it is on the benchmark closest to real autonomous work.
• Workplace automation — Grok 4.7: AutomationBench-AA 66%. Claude Opus 5: 57%. Grok 4.7's one clear win, and a real one — nine points on the evaluation that scores whether an agent finishes a workflow without tripping a guardrail.
• Knowledge and honesty — Grok 4.7: AA-Omniscience 32. Claude Opus 5: 37. Both models hallucinate; Opus 5 does it less on this measure.
• Long context — Grok 4.7: AA-LCR v1.1 77%, window 500k tokens. Claude Opus 5: 79%, window 1M tokens.
• Cost to run the index — Grok 4.7: 240M output tokens across the evaluation, flagged by Artificial Analysis as unusually verbose. Claude Opus 5: 140M. Grok 4.7 spends 1.7x the tokens to reach a lower score.

Where Grok 4.7's price advantage goes to die
Here is the arithmetic that the rate sheet hides. Take a realistic agentic request: 30k input tokens, 8k output. Grok 4.7 bills $0.06 in and $0.048 out — about eleven cents. Claude Opus 5 bills $0.15 in and $0.20 out — about thirty-five cents. That is a genuine 3x, and at this size Grok 4.7 is the obvious pick on cost alone.
Now take the request Grok 4.7's own marketing is built around: a long-context agentic session. Push the prompt past 200k tokens and Grok 4.7's rates double to $4.00 in and $12.00 out. Claude Opus 5 does not move — its 1M window bills at $5.00 and $25.00 flat. At 250k input and 8k output, Grok 4.7 costs $1.00 in and $0.096 out, about $1.10. Opus 5 costs $1.25 in and $0.20 out, about $1.45. The gap has collapsed from 3x to roughly 1.3x. Push further — 400k, 500k, the top of Grok 4.7's window — and Opus 5's flat rate keeps closing while its window keeps going to a million tokens. Grok 4.7 is the cheap model at short prompts and a near-peer on price at the long ones, which inverts the intuition the pricing page is designed to create.
Two caveats keep this honest. First, the long-context tier is a documented list-price structure from xAI's own model documentation, not a measurement — real bills depend on caching, and Grok 4.7's $0.50 cached-read rate is genuinely cheap. Second, Artificial Analysis has not published a speed measurement for Grok 4.7 yet, so the "it is faster" half of the cheap-and-fast argument is currently unverified. What is measured is Opus 5 at 59 tokens per second with a 49-second time to first token — slow, expensive, and ahead on almost every quality axis.
What you would actually route
This is a comparison where the honest answer differs by workload, and a router is the cheapest way to act on that. Claude Opus 5 is available on OrcaRouter, listed under its own model page with the provider's price passed through at 0% markup — so Opus 5's $5.00 and $25.00 on our catalog is Anthropic's list price, not a marked-up copy, and if Anthropic moves it the number changes on the same key the same day. Grok 4.7 is not in the OrcaRouter catalog as of this writing; to call it today you go through xAI's own API and the third-party platforms that carry it. That asymmetry is worth knowing before you architect around it.

The routing pattern that falls out of the numbers is straightforward. Send long-context, reasoning-heavy and terminal-agent work to Opus 5 — it wins Terminal-Bench 4.0 by 23 points and carries twice the window at flat pricing, which is exactly the profile of a hard, long job. Send high-volume automation and workflow tasks to Grok 4.7: it wins AutomationBench-AA by nine points and costs a third as much at short prompts. Because both are reached through one endpoint when they are in the catalog, that split is a routing rule rather than a second vendor contract, and automatic failover means a rate limit on one path becomes a routing event instead of an outage. Where a workload genuinely needs both — a cheap first pass and an expensive verification pass — the routing DSL composes them into a single call rather than two integrations.
The verdict
If you are choosing on evidence, Claude Opus 5 wins this matchup and it is not close: eight of ten evaluations, the two widest gaps, twice the context at flat pricing, and a token budget well under two-thirds the size for a higher score. The five-point composite undersells how thoroughly it leads. Grok 4.7's case is narrower than its price sheet suggests but it is not empty — it is genuinely cheaper at the prompt sizes most applications actually use, it is the better automation model, and it is a day old with a vendor benchmark suite still being independently reproduced. Buy it for cost-sensitive automation, watch its speed numbers when Artificial Analysis publishes them, and treat the long-context price tier as the thing that decides whether the cheap model stays cheap.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
