
Grok 4.7 Scores 46 on the Artificial Analysis Intelligence Index and Puts SpaceXAI in the Top Four
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 217 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 115 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 969 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 49 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 100 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 214 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Grok 4.7 Scores 46 on the Artificial Analysis Intelligence Index and Puts SpaceXAI in the Top Four
Grok 4.7 is out. SpaceXAI released it on Monday, September 21, 2026 — at the same $2 per million input tokens and $6 per million output tokens as Grok 4.6, with the same 500,000-token context window, a later knowledge cutoff and a new top reasoning level. It is live in Cursor and Grok Build, through the vendor's own API, and through third-party coding harnesses, model routers and cloud platforms. Grok 4.5, from July 8, stays below it as the cheaper option, and the models it still has to beat are unchanged: Claude Fable 5.1, Claude Opus 5 and GPT-5.6 Sol all sit above it on Artificial Analysis' v4.3.2 Intelligence Index, where Grok 4.7's first independent score — 46 — lands two points above Grok 4.6 and seven below Claude Fable 5.1. On the same evaluator's AA-Briefcase board for long-horizon knowledge work it lands fourth, behind three Claude entries and ahead of Claude Opus 5 at xhigh effort, at roughly half Claude Opus 5's cost per task — the first independent result in this release that reads as a price-performance win rather than a gap. Eleven days on, the picture has widened at the edges rather than in the middle: since October 1 it has been rolling out inside Grok's own web and mobile apps, where the mode picker now names it on its two hardest entries, and it is listed on OrcaRouter at SpaceXAI's own $2/$6.
That is the launch. The more useful numbers arrived the same day, from the one evaluator in this story who is not the vendor: Artificial Analysis published its first independent pass on Grok 4.7, and it is a three-part result — a 46 on the Intelligence Index that is only two points clear of Grok 4.6, a 56 on the Coding Agent Index that is nine points clear of it, and a 1,657 Elo on AA-Briefcase that is 111 points clear. Those three numbers point in different directions, and the spread between them is the most interesting thing about this release. What follows is the shipped specification, the vendor's own benchmark table with the label each row needs, and then the independent scores read properly — including what they say about the readiness caveat SpaceXAI never addressed.
What SpaceXAI shipped on September 21
Start with the specification, because it is unusually close to its predecessor's. Grok 4.7 is priced identically to Grok 4.6 — $2 per million input tokens and $6 per million output tokens below 200,000 input tokens, rising to $4 in and $12 out once a request crosses that threshold, the same structure Grok 4.6 introduced. The context window is 500,000 tokens. The knowledge cutoff is May 2026, three months later than Grok 4.6's February 1, 2026. Reasoning effort runs across four levels — low, medium, high and xhigh — and the published figures are quoted at xhigh. A fast variant runs at twice the output speed for twice the price.
Underneath, xAI describes three changes rather than a new architecture. The base model is larger than Grok 4.6's. The reinforcement-learning run was longer and weighted toward tasks that take many hours to complete. And the model was trained to verify its own work and to hold long context better, including native understanding of the Grok Bot harness. The company's own one-line summary is "twice as fast, at half the price of comparable models" — a positioning claim about rivals rather than a measurement, and worth reading as one.
Availability was broad on day one and has widened twice since. At launch: Cursor and Grok Build, the vendor's own API, third-party coding harnesses, model routers and cloud platforms; SpaceXAI's own Grok Build page now says outright to "start building with Grok 4.7". On September 28 Amazon Web Services added it to Amazon Bedrock with the same 500,000-token context window and the same four reasoning-effort levels, served through cross-Region inference profiles, so the model is now reachable through a hyperscaler's own catalog and not only the vendor's API. Then on October 1 it began rolling out inside Grok's consumer web and mobile apps, and two days later that is still where it sits. That last one is unusually quiet: SpaceXAI announced the equivalent step for Grok 4.5 with a blog post titled "Bringing Grok 4.5 to iOS, Android, Web, and X", and has published nothing for Grok 4.7, so the change is visible in the product rather than declared. It is server-side rather than shipped, which the app listings confirm: both Grok's iOS and Android listings moved on October 1 — the App Store build is 1.4.47, released that day, and its release note mentions only "Improvements to Chat, Voice and Imagine", while the Play listing was updated the same day. The change is pushed from the backend rather than built into a binary. Because the rollout is being reported rather than documented, the precise scope is harder to establish than for the Bedrock listing, which has a dated AWS post behind it, and the reporting has already drifted: the accounts tracking this now describe Grok 4.7 as the base model across all modes — Fast, Expert, Build and Heavy — a list that does not match what grok.com actually serves. Read it as the trackers' own reporting rather than as a vendor description. Selected cybersecurity partners get invite-only access to the model's red-team capabilities.
One correction to the record this piece kept. The parameter count that circulated for two months — roughly 2.1 trillion, about 40% above Grok 4.6's 1.5 trillion — is still not on any specification sheet. xAI has not published a parameter count for Grok 4.7, and Artificial Analysis lists the model size as undisclosed. The figure was always a founder's claim, and the launch did not convert it into a specification.
The benchmark table is xAI's own — here is what is not
Every figure in the list below comes from the vendor's announcement, and none of it has been independently reproduced. Read it with that label attached: these are xAI's numbers on xAI's chosen evaluations, published on the day of release, and the first week of any launch is when vendor benchmarks are least reliable. The comparison columns are the vendor's own too — Grok 4.6 (high), GPT-5.6 Sol (max) and Claude Fable 5.1 (max), each run at the vendor's chosen effort level.
• CursorBench 4.0 — Grok 4.7 46.3%, Grok 4.6 40.4%, GPT-5.6 Sol 41.7%, Claude Fable 5.1 51.8%. Vendor-reported.
• DeepSWE v1.1 — Grok 4.7 71.0% at high effort, Grok 4.6 65.2%, GPT-5.6 Sol 72.7%, Claude Fable 5.1 70.0%. Vendor-reported.
• Terminal-Bench 4.0 — Grok 4.7 37.6%, Grok 4.6 20.3%, GPT-5.6 Sol 37.3%, Claude Fable 5.1 57.9%. Vendor-reported, and revised: the Grok 4.7 row read 38.0% in the launch-day table and reads 37.6% on the vendor's page today.
• EEBench — Grok 4.7 64.0%, Grok 4.6 53.0%, GPT-5.6 Sol 39.4%, Claude Fable 5.1 56.4%. Vendor-reported.
• Harvey Legal Agent — Grok 4.7 19.6%, Grok 4.6 15.8%, GPT-5.6 Sol 2.5%, Claude Fable 5.1 6.7%. Vendor-reported.
• HealthBench Professional — Grok 4.7 56.7%, Grok 4.6 48.5%, GPT-5.6 Sol 60.5%, Claude Fable 5.1 62.1%. Vendor-reported.
• AA Briefcase v1.1 — Grok 4.7 1,657, Grok 4.6 1,546, GPT-5.6 Sol 1,487, Claude Fable 5.1 1,678. This is the one row in the table that is no longer vendor-only: Artificial Analysis' own AA-Briefcase board publishes the same 1,657 for Grok 4.7 at xhigh effort and 1,546 for Grok 4.6 at high. The comparison columns are still the vendor's choice of models and effort levels.
• GDPval Elo — Grok 4.7 1,695, Grok 4.6 1,605, GPT-6 Astra 1,542, Claude Fable 5.1 1,735. Vendor-reported.
The honest reading of that set is not "Grok 4.7 wins". It beats Grok 4.6 on all eight, which is the minimum a successor owes you. Against the field it is a mixed card: ahead of GPT-5.6 Sol on CursorBench, Terminal-Bench and EEBench, behind it on DeepSWE; behind Claude Fable 5.1 on CursorBench, Terminal-Bench, HealthBench and both Elo-style measures, ahead of it on EEBench and the Harvey legal agent evaluation. It also illustrates how much of a benchmark result is the effort level — DeepSWE's 71.0% is a high-effort number inside a table otherwise quoted at xhigh.
Safety is the one place the vendor's claims are unusually specific. xAI says Grok 4.7 was built on an entirely new safeguard stack and calls it the strongest model the company has tested for refusals and jailbreak resistance. On LatchBio's biosafety benchmark it reports 62.4%; on HackerBench v0.3 it reports letting through 3.3% of risky dual-use prompts while rarely blocking legitimate security work. Those are vendor-reported evaluations on the vendor's own release, and no third party has reproduced them.
The first numbers in this piece that are not the vendor's are the three Artificial Analysis published on September 21, and they disagree with each other in an informative way. On the Intelligence Index, Grok 4.7 scores 46 at xhigh effort on the v4.3.2 scale — 16th on that leaderboard — against 44 for Grok 4.6. That two-point gain is what Artificial Analysis says is enough to bring SpaceXAI into the index's top four labs. Read the placement precisely: it is a statement about a lab's best model, not about this one, and the model itself still sits four points below Claude Fable 5's 50 and seven below the 53 shared by Claude Fable 5.1 and GPT-6 Astra. A two-point gain is also within the range that shows up between effort levels of the same model, which is a reminder that a successor scoring above its predecessor is not the same as a successor changing what is possible. One further note on reading the number: the scale version matters, because Artificial Analysis restated its index and the 61/62/63 values that appear across most July and August coverage belong to the retired pre-September-2026 scale — they cannot be compared with these.
The Coding Agent Index is the second number, and it is the one that moved. Running in SpaceXAI's own Grok Build harness at xhigh effort, Grok 4.7 scores 56 against Grok 4.6's 47 — nine points, not two — which is fourth among models evaluated in their native harnesses, behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5, and ahead of GPT-5.6 Sol. The index is an equally weighted composite of DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA over 303 tasks, and all three components improved: DeepSWE from 65% to 73%, Terminal-Bench from 18% to 33%, SWE-Atlas-QnA from 58% to 63%. This is the first independent evidence that the fix xAI described — longer reinforcement learning weighted toward multi-hour tasks — did something, and it is the number a team choosing a coding agent should weigh hardest, because it is measured in the harness the model actually ships with rather than the vendor's own table.
The caveat is what the nine points cost. Artificial Analysis' own run generated 240 million output tokens on the Intelligence Index, against a median of 94 million across the models it tracks — more than double — and put the cost of evaluating one Index task at $3.74. At $6 per million output tokens, verbosity is the line item that decides whether an agentic loop is affordable, so a nine-point gain purchased with more than twice the median output is a real trade rather than a free upgrade. The same measurement also means the headline scores should not be read as one verdict: on a per-token basis Grok 4.7's advantage over Grok 4.6 is smaller than the index delta suggests, and on the general-knowledge evaluations that make up the Intelligence Index it is close to the same model with a substantially higher bill. The AA-Briefcase result in the next section is the exception to that, and it is a large one.

The second independent board: AA-Briefcase, and what half the cost per task buys
Artificial Analysis published a third number for Grok 4.7 the same day, and it is the one that speaks to knowledge work rather than code. On AA-Briefcase — its own board for long-horizon agentic knowledge work, built from four multi-week project scenarios and 91 graded deliverables — Grok 4.7 at xhigh effort scores 1,657 Elo, fourth on the board and behind only Anthropic entries: Claude Fable 5.1 at max effort (1,678), Claude Opus 5 at max effort (1,673) and Claude Fable 5.1 at xhigh (1,669). It is 16 Elo behind Claude Opus 5 at max effort and, at xhigh, 8 Elo ahead of Claude Opus 5 at xhigh (1,649). Read the placement precisely: "behind only Anthropic models" is the framing Artificial Analysis itself used, and it is accurate — but fourth is not second, and the three rows above it are all the same vendor.
The gain over the previous generation is the largest single move in the release. Against Grok 4.6 at high effort (1,546), Grok 4.7 at xhigh gains 111 Elo on AA-Briefcase — more than fifty times the two-point gain on the Intelligence Index. Artificial Analysis attributes it almost entirely to analytical quality: 1,994 Elo there against 1,690 for Grok 4.6, while presentation quality slips slightly, 1,499 against 1,519. That is a legible signature: the model reasons better about the analysis and formats the deliverable slightly worse. The same direction shows up in the figures Artificial Analysis quoted for AA-Briefcase-Lite, the public due-diligence scenario it released on Hugging Face — analytical quality up from 1,698 to 1,994, presentation down from 1,531 to 1,499. Two caveats on that comparison. AA-Briefcase-Lite is explicitly demonstrative: Artificial Analysis' own methodology page says the public fifth scenario "does not count toward official AA-Briefcase results." And the quality figures quoted for it match the xhigh rows published on the main board, so treat the Lite paragraph as an illustration of the direction of travel rather than a separate scoreboard.
The cost line is where this result changes a decision. AA-Briefcase's cost-per-task measure is the total cost of running the full submission divided by its 91 tasks; dividing Artificial Analysis' published totals out gives roughly $9.30 per task for Grok 4.7 at xhigh against roughly $17.79 for Claude Opus 5 at max effort — about half, for a 16-Elo gap. Artificial Analysis' own summary of the result is that Grok 4.7 ranks "just behind Opus 5 at ~50% of its Cost per Task." That is the same trade the rest of this piece describes, arriving at the same answer from a different direction: Grok 4.7 does not beat the frontier, it undercuts it. Two qualifications, both honest. The token bill is real — on AA-Briefcase-Lite, Artificial Analysis put the API cost of producing the example decks at about $8 for Grok 4.7 at xhigh against about $4.40 for Grok 4.6 at xhigh, so the successor costs roughly 80% more than the model it replaces on the same work. And the per-task figures are computed from token usage at list prices with a representative cache-hit rate, so they move with your caching: they are a model of a bill, not your bill.
Where Grok 4.7 sits against Grok 4.6 and the field
• Price per 1M tokens — Grok 4.7 $2 in / $6 out below 200K input, $4/$12 above it; Grok 4.6 identical.
• Context window — 500,000 tokens for both.
• Knowledge cutoff — May 2026 for Grok 4.7 versus February 1, 2026 for Grok 4.6.
• Reasoning effort — low / medium / high / xhigh for Grok 4.7 versus Grok 4.6's published levels.
• Independent score — 46 at xhigh effort versus 44, on Artificial Analysis' v4.3.2 Intelligence Index. Enough, Artificial Analysis says, to put SpaceXAI in the index's top four labs.
• Independent coding score — 56 versus 47 on the Artificial Analysis Coding Agent Index, each model in its native harness. Fourth among native-harness models, ahead of GPT-5.6 Sol.
• Independent knowledge-work score — 1,657 Elo at xhigh effort versus 1,546 for Grok 4.6 at high, on Artificial Analysis' AA-Briefcase board. Fourth overall, behind three Anthropic entries and ahead of Claude Opus 5 at xhigh.
• Cost per AA-Briefcase task — about $9.30 versus about $17.79 for Claude Opus 5 at max effort, on the same board. Roughly half the per-task bill for a 16-Elo gap.
• Output tokens per Index run — 240 million versus a 94 million median across the models Artificial Analysis tracks. The improvement is not free.
• Vendor coding eval — CursorBench 4.0 46.3% versus 40.4%.
• Serving speed — a fast variant at twice the output speed for twice the price, versus Grok 4.6's standard serving.
• Safety — a new safeguard stack, with a reported 62.4% on LatchBio's biosafety benchmark; Grok 4.6 shipped without that claim.
And the field, on the same scoreboard: Claude Fable 5.1 at 53, Claude Opus 5 at 51, GPT-5.6 Sol at 47, Kimi K3 at 44. Musk's own prediction — that Grok 4.7 "should be roughly on par with Opus 5.0, not 5.1" — now has a number under it, and the number landed where he said it would: below the frontier, not above it. The measured gap to Claude Fable 5.1 is seven points; the price gap runs the other way, with Claude Fable 5.1 listed at $10/$50 per million tokens against Grok 4.7's $2/$6 — five times the input price and more than eight times the output price. That is the actual trade a team is being asked to make, and it is a price-performance trade rather than a capability win. AA-Briefcase is the one board in this piece where the ordering changes: there Grok 4.7 at xhigh sits ahead of Claude Opus 5 at xhigh, 1,657 to 1,649, though still behind Opus 5 at max effort on 1,673 and behind Claude Fable 5.1 on 1,678. It is also worth setting the three independent scores side by side before drawing the line, because they do not point the same way: seven points behind on general intelligence, ahead of GPT-5.6 Sol on the coding agent index, 111 Elo clear of its predecessor on knowledge work, and paying for the gains in output tokens. Which of those facts decides your choice depends on whether the workload is a coding agent, a knowledge deliverable or a chat prompt, and that is a more useful split than a single rank.
What the leaks got right, and the one thing they did not settle
The artifacts this piece tracked through September were real, and the launch turns them into a test of what that class of evidence is worth. Here is the ledger.
• The catalog pull request got the price right. The September 21 submission adding Grok 4.7 to an open-source agent client's Zen and Go catalogs — opened 05:36 UTC, auto-closed by a compliance bot at 07:46 UTC over a missing pull-request template section, never merged — listed the model at Grok 4.6's published rates. It turned out to be exactly right: $2/$6, unchanged. But the mechanism matters more than the outcome. The contributor copied the rates from the model above it, and copying happened to be the correct guess. That is luck, not authority, and the same pull request's capability claims carried no information either. A proposal can be right and still not be evidence.
• The provisioning trail was genuine and was not the release. The grok-4.7 row on the Google Cloud Console's model quota page, reported by community trackers on September 16–17, and the dated grok-4-7-0907 build slug that surfaced in a Grok Bot configuration error, both pointed at a model being prepared. Neither had a date attached by anyone, and neither was the announcement. What they establish is that third-party infrastructure was being readied — which, in hindsight, was accurate and still not a schedule.
• The parameter count was never confirmed. Roughly 2.1 trillion, about 40% above Grok 4.6, repeated by Musk on September 2 — and still absent from every specification sheet at launch. This is the clearest case in the whole saga of a number that circulated widely enough to be treated as fact while never having a vendor source.

The card above records the pre-launch state as this piece last checked it: no model ID, no release date, no published benchmarks. Every row on it has since been overtaken by the September 21 announcement, and it is kept here because the gap between that card and the shipped model is the actual story of the last six weeks — a frontier launch that produced no confirmed specification at all until the day it shipped.
• The readiness caveat is the open question, and it now has partial evidence. Musk said in mid-September that Grok 4.7 "wraps up prematurely" on high-difficulty tasks and that its answer checking "is not strict enough," which he attributed to a reinforcement-learning penalty for long responses set too high, hedged with "possibly." The launch announcement does not address it. xAI's stated fix is indirect: longer reinforcement learning weighted toward multi-hour tasks and better self-verification is exactly the change you would make to address premature termination. The Coding Agent Index result is the first outside sign that it moved the needle — 56 against Grok 4.6's 47, with Terminal-Bench 4.0 nearly doubling from 18% to 33%, which is precisely the long-horizon, many-step end of the range. It is not a clean answer, because the same harness scores are also the ones buying their gains with far more output tokens. Whether the model still abandons long hard tasks is testable by anyone with API access, and it remains the first thing worth testing.
• The SpaceX corpus is still unverified. The reported training supplement — Starlink telemetry, Raptor combustion-chamber pressure logs, Starship re-entry telemetry, internal Slack threads, Jira tickets, GitHub repositories and ERP records — is community reporting of what Musk said, not a company disclosure. The launch announcement describes a larger base model and a longer reinforcement-learning run, and says nothing about the corpus. If it is in there, no specification sheet will confirm it; the only evidence will be whether the model does something on real engineering work that its peers cannot.
The timeline that missed four times, and what the market got right
Grok 4.7 was promised, in public, on five separate timelines, and the first four expired. On July 24–25 Musk said about four weeks, implying August 22. On August 12 he revised to "3 to 4 weeks," implying September 2–9. On September 2 he narrowed to roughly ten days, pointing at September 12. On September 11, the day that ran out, he said "a few more days." The fifth was not a window at all — a grok-4.7 row on a Google Cloud quota page — and it was the one closest to correct: the model shipped September 21, shortly after the last date anyone had committed to, and without a date attached to the artifact that preceded it.
The prediction markets called the calendar correctly and the model incorrectly. Kalshi's "SpaceXAI releases Grok 4.7 before 18 September?" stopped trading at 03:59 UTC on September 18 having last changed hands at 65¢, across 1,785 contracts traded and 1,211 open. Every settled contract on both major markets — Kalshi's August 14, 21, 28 and 31 and September 4, 8 and 11 windows, and Polymarket's September 12 and September 14 contracts — resolved No. Polymarket's "next Grok model (4.7 or higher) released by 18 September?" sat at 64% on September 15 after ranging between 42% and 88.5% across eleven days. The market was right to fade the tweet-driven spikes, and right that the September 18 window would close empty. It was three days early on the model, which is the one thing a dated contract cannot price.
The view counts are worth keeping as a calibration note. Musk's September 2 countdown, "comes out in 10 days," drew 10.29 million views and was wrong. His September 11 walk-back drew 2.05 million and was also wrong. His September 13–14 roadmap posts drew about 1.99 million and were the closest to right — he described Grok 4.8, a roughly 2.5-trillion-parameter model trained on a from-scratch C++ stack, as "a noticeable improvement" over Grok 4.7. Reach tracked confidence, not accuracy, through the whole of it.
Two of the roadmap items are now live questions rather than predictions. Grok 4.8 has no model ID, no pricing, no context window and no benchmark entry anywhere. And the framing that Grok 4.7 had been quietly "rebranded" into Grok 4.8 — one outlet's reading of Musk naming 4.8 while 4.7 was late — is settled the other way: 4.7 shipped as 4.7, at 46 on the Index against Grok 4.6's 44, which is a successor's increment and not a relaunch.
What this means for teams calling Grok today
The practical picture changed on September 21, and it changed in the least dramatic way possible: a new model ID with the same price, the same context window, a two-point Index gain, a nine-point coding-agent gain and a 111-point gain on knowledge work. Two changes since then matter more than they look. The Grok apps now hand the model to ordinary users rather than only to developers, and the catalog on this side now lists it — together those turn the routing layer described at the end of this piece from a plan into a default. If your cost model was built on Grok 4.6 at $2/$6, the per-token price is still correct for Grok 4.7 — but the token count is not. Artificial Analysis' own run burned 240 million output tokens on the Intelligence Index against a 94 million median across the models it tracks, so the same request can cost well over twice as much even though the rate card is identical, which is the kind of change that never shows up in a pricing table. Price a real agentic loop on output tokens, not on the per-million rate, before you conclude this release is free. If your reason to switch is the price-performance frontier the vendor is claiming, the strongest evidence for it is now AA-Briefcase rather than the vendor's table: fourth on that board at roughly half Claude Opus 5's cost per task, set against a seven-point Intelligence Index gap to Claude Fable 5.1 and a price gap running five-fold the other way on input and more than eight-fold on output. And if your workload is a coding agent specifically, the case is stronger than the Index suggests: 56 on the Coding Agent Index in the Grok Build harness, ahead of GPT-5.6 Sol, is a different league from 46 on general intelligence. Which of those matters more depends entirely on your workload, and nobody outside SpaceXAI has run Grok 4.7 on yours.
On OrcaRouter's catalog the Grok line now runs Grok 4.7 alongside Grok 4.6, Grok 4.5 and Grok 4.3, billed at the provider's list price with 0% markup — $2 per million input tokens and $6 per million output, cached input reads at $0.50, and the same step to $4 in and $12 out once a request crosses 200,000 input tokens that SpaceXAI charges. That is what passing the rate card through untouched buys: the per-token price in your cost model is the per-token price on your invoice, and every model in the family answers to one key, so moving a workload from Grok 4.6 to Grok 4.7 is a string change rather than a second contract. That page carries the 500,000-token context window and the same OpenAI-compatible surface — Chat Completions and Responses — that a Grok 4.6 integration already targets.

That routing layer is also the low-risk way to evaluate a model whose independent numbers landed the same day as the vendor's. The benchmark list above is still the vendor's own — its chosen evaluations, its chosen effort levels, its chosen comparison columns — and none of it has been reproduced. What has changed is that the three Artificial Analysis scores are not the vendor's, so the question has moved from "is any of this real" to "which of these results does my workload look like": the 46 that says mid-pack general intelligence, the 56 that says fourth among native-harness coding agents, or the 1,657 that says fourth on knowledge work at half the frontier's per-task cost. And the number that decides the economics is not a benchmark at all but the output tokens a run consumes, which only your own traffic can price — 240 million per Index run on Artificial Analysis' measurement against a 94 million median, and roughly $8 against $4.40 per set of example decks on its public AA-Briefcase-Lite scenario. Automatic failover lets you point real traffic at the new model and fall back to a provider that still answers if the token bill or the premature-abandonment behaviour turns out worse than advertised — which is the difference between testing an unproven model and betting a pipeline on it.
How to know first (the check that worked)
This section used to be a list of signals to watch while waiting. The wait is over, so here is the same list scored against what actually happened.
• The model list was the confirmation, and it moved. For six weeks the advice in this piece was to watch the vendor's model list rather than the tweets, because the moment Grok 4.7 was real the list would gain a model ID with pricing and a context window attached. That is what happened: grok-4.7 now appears with a 500K context window, $2/$6 pricing below 200K input and $4/$12 above it, a May 2026 knowledge cutoff and four effort levels. Every Musk window, every cadence argument and every market date failed to move that list; the release moved it.
• Third-party catalogs lead the announcement, and they are proposals rather than deployments. The September 21 catalog pull request was the clearest version of this yet, and its closing is instructive: a contributor prepared for a model, a bot closed the change over a missing template section two hours later, and the model shipped two days after that. Read that class of artifact as intent with a timestamp. It is genuinely upstream of the announcement, and it is not availability.
• Watch the OrcaRouter model catalog. It has now moved. For the whole of September the advice in this piece was that a grok-4.7 model page on orcarouter.ai would be the "you can call it today through us" signal, because a model is listed only once a provider has onboarded it. The catalog now carries Grok 4.7 at the vendor's rate with no markup, so the check a reader was told to run has an answer: it is routable through one API today.
• For consumer-side visibility, the Grok app has now put Grok 4.7 in front of users who will never open an API console. Read again on October 2, the picker data grok.com serves to a logged-out visitor still names the model on two of its four entries — Expert reads "Thinks hard · Grok 4.7" and Heavy reads "Team of Experts · Grok 4.7" — while Fast, which is the mode the app opens in by default, is described only as "Quick responses" and Auto as choosing between Fast and Expert. The four entries are Auto, Fast, Expert and Heavy. Build is not one of them: Grok Build is a separate terminal product on its own page, and that page is where the model's Build attribution actually comes from — it tells visitors to "install now and start building with Grok 4.7". So the claim that circulated on October 1 and was repeated on October 2 — that Grok 4.7 is now the base model across "Fast, Expert, Build and Heavy" — lists a mode the app does not serve and omits the one it opens in, and it is the trackers' reading rather than the vendor's, because SpaceXAI has posted nothing about this rollout at all. The equivalent Grok 4.5 step got a blog post of its own; this one has a product that changed quietly underneath a news page that did not.
The state of play
Grok 4.7 shipped on September 21, 2026, at $2 per million input tokens and $6 per million output tokens with a 500,000-token context window and a May 2026 knowledge cutoff. What has changed in the eleven days since is the availability map rather than the model: Amazon Bedrock added it on September 28, it began rolling out inside Grok's own web and mobile apps on October 1 and was still rolling out there on October 2, and it is now listed on OrcaRouter at SpaceXAI's own rate with no markup. Its independent scores arrived the same day as the launch and still disagree: 46 at xhigh effort on Artificial Analysis' v4.3.2 Intelligence Index, two points above Grok 4.6 and enough to put SpaceXAI in the index's top four labs; 56 on the Coding Agent Index in the Grok Build harness, nine points above Grok 4.6 and fourth among native-harness models ahead of GPT-5.6 Sol; and 1,657 Elo on AA-Briefcase, 111 points above Grok 4.6 and fourth on the board behind three Anthropic entries, at roughly half Claude Opus 5's cost per task. Every benchmark in the vendor's launch table is the vendor's own and unreproduced, with one exception: Artificial Analysis' own AA-Briefcase board publishes the same 1,657, and the vendor has since revised its own Terminal-Bench row from 38.0% down to 37.6%. The coding and knowledge-work gains are real and they cost output tokens — 240 million per Index run against a 94 million median, and about $8 against $4.40 per set of example decks on AA's public due-diligence scenario — which is the number that decides whether this release is cheap. The ~2.1-trillion-parameter figure that circulated for two months still has no vendor source, and the premature-abandonment caveat was never directly addressed, though the Terminal-Bench jump from 18% to 33% in the Grok Build harness is the first outside evidence that the fix worked. Both signals this piece told readers to watch have fired: the vendor's model list gained grok-4.7 with a price and a context window attached, and the catalog here lists it at the same rate. So the winning move is simpler than it was in September — Grok 4.6 at $2/$6 was the answer, and Grok 4.7 at $2/$6 is the same answer with a newer model ID, better coding and knowledge-work scores, a much larger token bill, and a consumer app that names it on its two hardest modes.
Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
