
What Is Grok 4.7? Specs, Price, Context Window and Benchmarks
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 300 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 112 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Grok 4.7 is SpaceXAI's current flagship model: a reasoning model that takes text and images in, returns text, carries a 500,000-token context window, and lists at $2.00 per million input tokens and $6.00 per million output tokens with reasoning effort selectable across four levels. It succeeds Grok 4.6, matches it on price and context exactly, and is aimed at the long-horizon work — multi-hour coding, agentic tool loops, professional knowledge work — that the company now sells its whole product line around. This page is the reference entry for it: what the model is, where it sits, what it accepts, what it costs, and then two things kept deliberately apart — what SpaceXAI says about its own model, and what an independent evaluator measured.
Why this is a reference page and not a launch story. Grok 4.7 shipped on September 21, 2026, and that launch is already covered on this blog (see the grok-4-7-release-date piece for the day-one reading of it), so the date below is used as a fact that dates the model, not as an event being reported again. The reader this page exists for is the one who types the model's bare name into a search box and wants the model, not the news: over the 28 days to September 25, 2026, our own Search Console data records 31,934 impressions on the exact query "grok 4.7" and 34,160 across every spelling that reduces to that name, landing at an average position of 7.8 with no dedicated page to land on. That is a measurement of what people type — it says nothing about how good the model is — and it is why the page starts with the plainest possible question. Every figure in this piece is labelled by where it came from: vendor-reported for anything from SpaceXAI's own docs, launch post or pricing page, independent for Artificial Analysis, and our own for what our catalogue and our search data show.
Where Grok 4.7 sits in the Grok line
The vendor now writes its name as SpaceXAI on both the developer docs and the news page, while the API itself is still served from api.x.ai and the model identifier is still a grok one. Grok 4.7 is the top of the text line and the vendor's own default recommendation: asked which model to choose, the models page answers that for everything except audio, image and video work — "including code" — you should use Grok 4.7, because "it is the most capable model we've built." Its siblings, at the vendor's own published rates per million tokens, are:
• Grok 4.6 — 500K context, $2.00 input / $0.50 cached / $6.00 output, shipped August 2026. The direct predecessor and the model 4.7 is priced against.
• Grok 4.5 — 500K context, $2.00 input / $0.30 cached / $6.00 output, shipped July 2026. Same list price as its successors but a cheaper cache rate, which is now the one concrete reason to still choose it.
• Grok 4.3 — 1M context, $1.25 input / $0.20 cached / $2.50 output. Half the price of the current flagship and twice its context window.
• Grok 4.20 (reasoning, non-reasoning and multi-agent variants) — 1M context, $1.25 input / $0.20 cached / $2.50 output.
• Grok Build 0.1 — 256K context, $1.00 input / $0.20 cached / $2.00 output. A coding specialist rather than a general model, and the default model of the Grok Build agent is Grok 4.7, not this.
So the shape of the current line is a capability-versus-window trade: the flagship gives you 500K of context at $2/$6, and the 4.3/4.20 tier gives you a full million tokens at roughly half the input price and less than half the output price. Nothing in the vendor's documentation suggests Grok 4.7 is a different architecture from Grok 4.6 — what the release describes is a larger base model, a longer reinforcement-learning run weighted toward tasks that take many hours to finish, better self-verification, better long-context handling, and training on the Grok Bot harness so the model behaves better in conversational and general knowledge work. An improvement claim, not a new design. SpaceXAI publishes no parameter count for Grok 4.7 on either its model page or its pricing table, and this page will not guess one.
When it shipped, and what shipped with it
The launch post is dated September 21, 2026, and the developer model page carries the same date as its last update, which is as precise as this gets — SpaceXAI does not publish a release timestamp beyond the day. What shipped that day was not just a model identifier. The grok-4.7 slug went live on both the Responses API and Chat Completions with a 500,000-token context window, a May 2026 knowledge cutoff, four reasoning levels including a new top setting, a Fast variant restricted to two first-party surfaces, a US-only regional endpoint, and list pricing identical to Grok 4.6's. It was simultaneously made the default model of the Grok Build coding agent and made available in Cursor on all plans. If you want the day-one reaction rather than the specification, the grok-4-7-release-date piece covers that ground; this page takes the same date as a fact about the model and moves on to what it will actually do for you.
The envelope: what you can send, and how much of it
All of the following is from SpaceXAI's own model page and models page, read on September 28, 2026, unless labelled otherwise.

• Context window — 500,000 tokens. Same as Grok 4.6, half of what Grok 4.3 and the 4.20 family offer.
• Knowledge cutoff — May 2026, three months later than Grok 4.6's. Anything more recent requires a search tool switched on; the model has no realtime awareness by default.
• Input — text and image. Images may be JPG/JPEG or PNG, up to 20 MiB each, and the vendor documents no cap on the number of images in a request; image and text may appear in any order.
• Output — text only. There is no image, audio or video generation in this model; those are separate models in the SpaceXAI line.
• Output ceiling — the vendor states there is no text output limit. Our own OrcaRouter model card for grok-4.7 lists a maximum output of 450,000 tokens, and the two figures do not agree. We are flagging it rather than quietly picking one: SpaceXAI's documentation is the authority on SpaceXAI's model, and the honest reading is that the vendor does not publish a text output cap.
• Reasoning control — reasoning_effort takes low, medium, high (the vendor's default) or xhigh. xhigh is new in this generation; Grok 4.6 also offered four levels, and the benchmark figures published for 4.7 are quoted at xhigh.
• Tools — function calling, web search, X search and code execution, all documented as server-side tools.
• Structured output — the vendor documents a structured-outputs capability alongside text generation; our own page exposes it as response_format and structured_outputs.
• Parameters our page accepts — include_reasoning, logprobs, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_logprobs and top_p. One caveat worth knowing before you build on logprobs: SpaceXAI's models page states that logprobs and top_logprobs are not supported by grok-4.20 and newer models and "will be silently ignored if set", which covers Grok 4.7. Our page listing them is not a promise that the upstream returns them.
• Endpoints — the OpenAI-compatible Chat Completions API and the Responses API.
• Prompt caching — the vendor recommends setting prompt_cache_key on the Responses API (the x-grok-conv-id header on Chat Completions) so a conversation's requests land on the same server; without it, the docs warn, you often pay full input price on a cache-cold server. Cached input is billed at $0.50 per million tokens below the 200K threshold.
• Encrypted reasoning — on the Responses API, Grok 4.7 always returns reasoning.encrypted_content, even when you do not ask for it, so multi-turn calls keep the model's reasoning without extra configuration. Chat Completions is unchanged.
The rate card, in full
Pricing is tiered by prompt size, and the threshold applies to the whole request once it is crossed, not just to the tokens above the line. Every figure here is vendor-published, read from SpaceXAI's pricing page on September 28, 2026:
• Standard, prompts under 200K tokens — $2.00 input / $0.50 cached input / $6.00 output per million tokens.
• Long context, prompts at or above 200K tokens — $4.00 input / $1.00 cached input / $12.00 output per million tokens.
• Grok 4.7 Fast — the same model on faster infrastructure at double the standard rates: $4.00 / $1.00 / $12.00 below 200K, and $6.00 / $1.50 / $18.00 above it. It is available only through Cursor and Grok Build, it is not on the public API at all, and Grok Build's free tier does not include it.
• US regional endpoint — https://us.api.x.ai/v1 keeps inference in the United States and bills at 1.1×, which works out to $2.20 / $0.55 / $6.60 per million below 200K and $4.40 / $1.10 / $13.20 above. Only grok-4.7 and grok-4.6 are served there today.
Read against the rest of the market, the interesting thing is not the headline rate — $2 in and $6 out is exactly what Grok 4.6 charged and what Grok 4.5 still charges — it is the cache discount, which the vendor publishes at 75% off input for repeated prompt prefixes. For a long agent loop that re-sends the same system prompt and file context every turn, that discount is the difference between the list price and the bill. It is also the one line where Grok 4.7 is worse than its older sibling: Grok 4.5 caches at $0.30 per million against Grok 4.7's $0.50.
Grok 4.7 is on OrcaRouter as grok/grok-4.7 at SpaceXAI's own list price with 0% markup — the provider price passed through, so a vendor price change is live here the same day rather than after a repricing pass. That matters more than usual on a model with a two-tier prompt threshold and a cache rate, because both of those move what you actually pay. It is one key and one OpenAI-compatible endpoint for the whole Grok line, so switching a workload from Grok 4.7 to Grok 4.5 or 4.3 to test the cost difference is a string change rather than a second contract.
What SpaceXAI claims for it
Everything in this section is vendor-reported. None of it has been independently reproduced, and the effort setting matters — the table on the launch post quotes Grok 4.7 at xHigh, Grok 4.6 at High, and the two comparison models at Max, which is not a matched comparison.
• CursorBench 4.0 — 46.3% for Grok 4.7 at xHigh, against 40.4% for Grok 4.6 at High. The launch post's own framing is that this is where Grok 4.7 is "at the frontier in price-performance".
• Terminal-Bench 4.0 — 37.6%, against 20.3% for Grok 4.6. The largest single delta in the release, and the one that matches the "longer RL on harder tasks" claim.
• DeepSWE v1.1 — 71.0%, marked in the vendor's own table as a high-effort score, against 65.2% for Grok 4.6.
• AA Briefcase v1.1 — 1,657, against 1,546 for Grok 4.6. Note the collision of names here: this is the vendor quoting a benchmark that Artificial Analysis also runs, and the independent result is discussed in the next section.
• EEBench — 64.0%, against 53.0% for Grok 4.6.
• Harvey Legal Agent Benchmark — 19.6%, against 15.8% for Grok 4.6.
• HealthBench Professional — 56.7%, against 48.5% for Grok 4.6.
• GDPval — 1,695 Elo at xhigh, against 1,605 for Grok 4.6 at high.
The safety claims are vendor-reported too, and unusually specific: SpaceXAI says Grok 4.7 was built on a new safeguard stack, that it tops LatchBio's biosafety benchmark at 62.4%, and that on its own HackerBench v0.3 it allows only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. The company also says it has begun giving selected cybersecurity partners invite-only access to the model's red-team capabilities. Treat all of that as the vendor's own testing of its own model — it is a claim set specific enough to be checked later, which is not the same as checked now. The one-line positioning on the launch post, "twice as fast, at half the price of comparable models", is marketing about rivals rather than a measurement of Grok 4.7, and the independent speed figure below does not support the "twice as fast" half of it.

What an independent evaluator measured
Artificial Analysis is the one source in this story with no stake in the model, and its Grok 4.7 page is explicitly the xhigh configuration. Read on September 28, 2026:
• Intelligence — 46 on the Artificial Analysis Intelligence Index, index version v4.3.2, which is 21st of 211 models on that board and well above the index's median of 26. That index is a composite of ten evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1.
• Cost — $3.74 to run the Intelligence Index once, with the same evaluator publishing the cache discount at 75% and an input/output rate matching the vendor's ($2.00 / $6.00). Its own comment is that the model is "reasonably priced when comparing to other models of similar price".
• Verbosity — 240M output tokens generated across the Intelligence Index, against a median of 88M. Artificial Analysis calls this "very verbose", and it is the single most useful number here for anyone budgeting: a model that thinks in longer chains costs more per task even at a fixed token rate.
• Speed — 82.6 output tokens per second, 83rd of 211 on that board, which the evaluator nonetheless labels "slower than average". Two caveats before you carry this number anywhere. Artificial Analysis changed its default benchmarking workload to a 10,000-token input to better reflect production use, and its published speed reading for this model has moved materially between snapshots, so treat it as a snapshot of one workload rather than a constant. On the API-provider side the same evaluator has only one provider listed for the model — SpaceXAI itself — so its measured latency and blended price are first-party figures, not a market spread.
Our own catalogue carries three further Artificial Analysis results for Grok 4.7, evaluated September 21, 2026: 43.1% on Humanity's Last Exam, 76.7% on the long-context recall evaluation, and 57.4% on SciCode. Those are labelled independent because they come from Artificial Analysis, and they are worth reading alongside the vendor's self-verification claims rather than as confirmation of them.
The honest comparison note: SpaceXAI's numbers and Artificial Analysis' numbers were produced on different harnesses at different effort settings, and the two have not been run head to head in a matched configuration. Where the same benchmark name appears on both sides — AA Briefcase, Terminal-Bench 4.0 — that is a shared name, not a shared run. Anyone who quotes a vendor figure next to an independent figure as if they were one measurement is doing the arithmetic the evidence does not support.
Where you can actually call it
Grok 4.7 has been callable since launch day through SpaceXAI's own API at api.x.ai, using the Responses API or Chat Completions, and through the US-only regional endpoint at a 10% premium. It is the default model of the Grok Build coding agent and is available in Cursor on all plans, with the Fast variant — the same model at double the rate — restricted to those two surfaces and absent from the public API. Beyond first-party access the vendor describes it as available through third-party coding harnesses, model routers and cloud platforms.
On our side specifically: grok/grok-4.7 is live in the OrcaRouter catalogue today, priced at xAI's own list rates with 0% markup and callable on either endpoint through the same key as the rest of the Grok line. If you are evaluating it against Grok 4.6 — same price, same window, a different model underneath — routing both through one endpoint with an automatic failover chain is the cheap way to find out which one your workload actually prefers, rather than committing a production path to the newer model on the strength of a vendor table.

What to watch, given how thin the record still is. SpaceXAI has published no parameter count, and the output-ceiling question is genuinely open — the vendor's "no text output limit" against the 450K our own page shows is a discrepancy nobody has resolved yet. There is no matched-harness head-to-head against the models it is sold against. And the speed reading above is the number most likely to move, because it already has. Those are the three places this page will need editing as the record fills in, and it is better to say so than to pretend the specification is settled.
Grok 4.7 is routable on OrcaRouter today, and so are Grok 4.6, Grok 4.5, Grok 4.3 and the Grok 4.20 family. OrcaRouter passes the provider's list price through at zero markup, so SpaceXAI's two-tier prompt threshold and its cache rate land on our side the same day they change on theirs, and testing a workload against the older, cheaper Grok models is a string change rather than a second contract.
Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
