A generated hero title card reading 'GPT-6 Luna vs Gemini 3.1 Pro', subtitled 'A preview that has run seven months against a GA model that shipped three weeks ago.', with chips showing Luna at $0.10/$0.50 per 1M and GA since September 22, Gemini 3.1 Pro at $2.00/$12.00 per 1M and in preview since February 19, and index scores of 37 against 30, with the OrcaRouter logo bottom-right.
Guides & Insights

GPT-6 Luna vs Gemini 3.1 Pro: A Seven-Month-Old Preview at Twenty Times the Input Price

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The frontier reasoning model has been labelled "Preview" since Gemini 3.1 Pro launched on February 19, 2026. Seven months later it is still a preview, still at $2.00 per million input tokens and $12.00 per million output, and still what the product page calls the company's frontier reasoning model. On September 22, 2026, the vendor shipped GPT-6 Luna as a generally available small tier at $0.10 and $0.50. Read the two rate cards side by side and the newer, cheaper, GA model is twenty times cheaper on input and twenty-four times cheaper on output — and it scores higher on the independent index that both vendors' marketing departments care about.

That last clause is the part worth slowing down for. GPT-6 Luna measures 37 on the Artificial Analysis Intelligence Index at maximum effort. Gemini 3.1 Pro Preview measures 30 on the same revision. A model sold as a high-volume efficiency play is seven points ahead of a model sold as a frontier reasoning system, at a twentieth of the input price. This is not a case of a cheap model catching up to an expensive one. On this particular axis the cheap model is simply ahead, and the expensive one has been in preview for seven months.

The natural conclusion — that Gemini 3.1 Pro is overpriced — is not quite right either, and the rest of this piece is about the three places where Google's model earns its rate card, because they are real and they are not small.

The five things that decide it

Before the detail, the short version:

• GPT-6 Luna wins on price by roughly 20x on input and 24x on output, on cached input by a wide margin, and on the general-purpose index by seven points at matched max effort.

• Gemini 3.1 Pro wins on input modality — it accepts audio and video, which GPT-6 Luna does not — and on having been reachable for seven months, which for a preview is a long production track record.

• The output ceilings are not close in either direction: GPT-6 Luna emits up to 128,000 tokens, Gemini 3.1 Pro up to 65,000.

• Both carry a long-context tier boundary that reprices the whole request, at 272,000 input tokens for GPT-6 Luna and around 200,000 for Gemini 3.1 Pro.

• The default-effort trap applies to Luna only: its 37 is a max-effort score and its default medium setting scores 29, which is below Gemini 3.1 Pro's 30.

Two rate cards and the arithmetic between them

One line per dimension, both sides on each:

• Input per 1M — GPT-6 Luna $0.10 vs Gemini 3.1 Pro $2.00

• Output per 1M — GPT-6 Luna $0.50 vs Gemini 3.1 Pro $12.00

• Cached input — GPT-6 Luna $0.01 per 1M, a 90% discount vs Gemini 3.1 Pro $0.20 per 1M, a 90% discount

• Long-context boundary — GPT-6 Luna above 272K input tokens vs Gemini 3.1 Pro above roughly 200K input tokens

• Long-context effect — GPT-6 Luna 2x input and cache, 1.5x output, on the whole request vs Gemini 3.1 Pro $4.00 in / $18.00 out, also on the whole request

• Context window — GPT-6 Luna 1,050,000 tokens vs Gemini 3.1 Pro 1M tokens

• Maximum output — GPT-6 Luna 128,000 tokens vs Gemini 3.1 Pro 65,000 tokens

• Input modality — GPT-6 Luna text and image vs Gemini 3.1 Pro text, image, audio, video and files

• AA Intelligence Index — GPT-6 Luna 37 at max effort, 29 at its default vs Gemini 3.1 Pro 30

• Output speed — GPT-6 Luna 153.9 tokens/sec vs Gemini 3.1 Pro roughly 115-143 tokens/sec depending on the snapshot

• Reasoning off switch — GPT-6 Luna none through max vs Gemini 3.1 Pro reasoning-first, no non-reasoning mode exposed

• Status — GPT-6 Luna generally available since 2026-09-22 vs Gemini 3.1 Pro in preview since 2026-02-19

A generated single-panel scoreboard card headed 'Two rate cards, seven months apart' with six rows: GPT-6 Luna $0.10 in / $0.50 out per 1M and GA since September 22, 2026; Gemini 3.1 Pro $2.00 in / $12.00 out per 1M and in preview since February 19, 2026; Intelligence Index Luna 37 at max and 29 at default against Gemini 3.1 Pro's 30; input modality Luna text and image against Gemini 3.1 Pro's text, image, audio and video; maximum output Luna 128K tokens against Gemini 3.1 Pro's 65K; and a long-context clause that reprices the whole request on both. Footer: 'Index figures per Artificial Analysis; prices per OpenAI and Google.'

The interesting row is not the price. It is the pair of rows at the bottom. Gemini 3.1 Pro's preview label is not a technicality — it changes what Google owes you. A preview model can be changed, repriced, or withdrawn without the deprecation notice a GA model gets, and pricing that is unchanged seven months in is a promise nobody has made in writing. Everything below about Google's model being the safer production bet has to be read against that caveat, because "seven months of stability" and "seven months without a stability guarantee" are the same seven months.

Where Google's model earns the extra twenty times

Three places, and the first is the one that ends the comparison for some teams outright.

Modality. Gemini 3.1 Pro takes audio and video input. GPT-6 Luna takes text and image. If your pipeline ingests call recordings, screen captures, or video, there is no version of this matchup where the price difference matters, because one of the two models cannot do the job. That is not a benchmark gap that narrows with a point release; it is a capability that is either present or absent, and it is the single strongest reason Google's rate card survives contact with a real procurement.

Output ceiling, in the other direction. The output caps cut the opposite way from the input modality, and this one favours OpenAI. GPT-6 Luna emits up to 128,000 tokens in one response; Gemini 3.1 Pro up to 65,000. If you are generating long structured artifacts — a full document, a large refactor, a multi-file patch — Google's ceiling is roughly half of OpenAI's, and a pipeline that truncates at 65,000 tokens is broken regardless of what it pays per token.

Multimodal frontier work. Gemini 3.1 Pro's positioning is as a reasoning model that happens to be multimodal, and the practical consequence is that tasks which require reasoning over a diagram, a chart, or a screen recording land on it rather than on a text-first small tier. GPT-6 Luna accepts images, so the boundary is not audio and video alone — it is that Google's model is built for visual and audio reasoning as a primary use case, and OpenAI's is built for throughput.

Where the cheap tier actually loses, and it is not where you expect

The effort default is the trap in this pairing, and it cuts harder here than it would against a weaker opponent. GPT-6 Luna's headline index score of 37 is measured at maximum reasoning effort. Its default is medium, and at medium it scores 29 — one point below Gemini 3.1 Pro's 30. So the seven-point lead that this article opened with is a comparison between one model's ceiling and another model's ceiling, and a team that installs GPT-6 Luna and leaves the configuration alone is running a model that is fractionally behind Google's, seven months older, still in preview, at a twentieth of the price.

That is still, for most workloads, the right trade — a point of index is not a capability difference and a 20x price gap is. But it is a different trade from the one the rate cards advertise, and it is invisible unless you go looking for the effort parameter.

The second place the cheap tier loses is long-context arithmetic. Both models reprice the entire request once it crosses a boundary rather than surcharging the excess, and both boundaries are low enough to matter: 272,000 input tokens on GPT-6 Luna, roughly 200,000 on Gemini 3.1 Pro. GPT-6 Luna's clause doubles input and cache and lifts output by half. Gemini 3.1 Pro's moves the whole call to $4.00 in and $18.00 out. Neither is a marginal surcharge, and on a model sold on price the clause is the price.

The third is verbosity, which is the cost line that never appears on a rate card. Artificial Analysis measured GPT-6 Luna generating 150 million output tokens across the index run against a median of 85 million — the model is chatty, and at $0.50 per million output that is $75 of tokens where a concise model would have spent $42.50. It is still an order of magnitude cheaper than the alternative. It is also the reason cost-per-task numbers and cost-per-token numbers tell different stories, and the reason to measure your own output volume before you model the saving.

What a migration between them looks like

Google's model is reachable through the vendor's own API and several third-party platforms; GPT-6 Luna is reachable through OpenAI's own API, and as of this writing neither is the simpler half of the pairing to standardise on. Both expose an OpenAI-compatible chat completions surface, which means the call shape is not the problem — the parameters are. GPT-6 Luna's effort ladder, its 272,000-token clause, and its requirement that function calling on Chat Completions run with reasoning effort set to none are all behaviour that Gemini 3.1 Pro does not share, and a port that changes only the model string will produce a working call with the wrong cost profile.

Gemini 3.1 Pro Preview is on OrcaRouter at Google's list price, under the pass-through pricing that means a Google rate change is live on our side the same day. GPT-6 Luna is not in our catalogue as of this writing. For a team running both — Google's model for anything that needs audio or video, OpenAI's for high-volume text — that is one key for Gemini 3.1 Pro alongside whatever you already use for OpenAI, with the routing decision expressed as configuration rather than as two code paths. This is the pairing where that matters most, because the two models are not interchangeable at all: a route that has to choose between a multimodal preview and a text-only GA small tier is a route that will be re-decided every time either vendor ships.

A screenshot of the Artificial Analysis page for GPT-6 Luna (max), showing an Intelligence rank of 6th of 183 models, a Speed rank of 35th, a Cost rank of 20th and a Verbosity rank of 36th, input pricing of $0.10 and output of $0.50 per 1M tokens, an Intelligence Index score of 37 against a comparable-model median of 12, 150M output tokens generated during the index run, 153.9 tokens per second, and a total of $122.39 spent evaluating the model.

What would change this

The comparison as it stands is a snapshot of an unusual moment: a GA model from September beating a preview model from February on the general index at a twentieth of the input price, while losing on modality and on the maturity that a preview never quite earns. Three things would move it, and each is worth watching rather than guessing at.

The first is Gemini 3.1 Pro leaving preview. A GA release would bring a deprecation policy, a stable rate card, and very likely a price move — Google has held $2.00/$12.00 across seven months of preview while the rest of the market halved around it, and that restraint is more consistent with a launch price waiting for a GA date than with a considered position.

The second is whether OpenAI extends the Luna effort ladder or changes its default. The twenty-two-point spread between GPT-6 Luna's max-effort score and its default-effort score is the largest configuration sensitivity in this article, and a default move would change the model's effective capability for every caller who never touched the parameter.

The third is the modality gap. It is the only dimension here that is not a matter of degree, and it is the one that keeps this from being a straight price comparison. Until one of these models can do what the other cannot, the twenty-times gap is a real number pointing at two different jobs.

A screenshot of the OrcaRouter model page for Gemini 3.1 Pro Preview (google/gemini-3.1-pro-preview), showing the Vision, Audio, Tools, JSON and Reasoning badges, a 2026-02-19 catalogue date, 65K maximum output, list pricing of $2.00 input and $12.00 output per 1M tokens, a p50 time to first token of 10.60s, 111.8M tokens of traffic, and the description of the model as Google's frontier reasoning model for software engineering and agentic reliability.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily