Card del titolo hero con la scritta "GPT-6 vs Gemini 4 Argon", con un chip che riporta "Argon: implementazione graduale per i partner di sicurezza dal 2026-09-30", due chip dell'indice con la scritta "GPT-6 Astra 52.7" e "Gemini 4 Argon 52.6", e un piè di pagina che recita "Valori dell'indice secondo Artificial Analysis v4.3.2; specifiche di Argon dichiarate dal fornitore e non disponibili per l'acquisto." Il logo OrcaRouter è sovrapposto nell'angolo in basso a destra.
Guides & Insights

GPT-6 vs Gemini 4 Argon: quasi parità tra un modello a listino e uno in lista d'attesa

Autore

Elias Hawthorne

Data di pubblicazione

Ultimi modelli · 20Vedi tutti i modelli →
Benchmark: Artificial Analysis · aggiornato ogni giorno
Torna a tutti gli articoli

The closest model pair on the current frontier board is not a pair you can buy. Gemini 4 Argon scores 52.6 on Artificial Analysis's Intelligence Index v4.3.2. GPT-6 Astra, Ope​nAI's flagship, scores 52.7. That is a one-tenth-of-a-point gap across ten evaluations, and both figures come from the same revision of the same suite. It is the tightest top-of-table pairing in the set.

There is a symmetry in that number and none at all in availability. Goo​gle announced Gemini 4 Argon on 30 September 2026, at a stated introductory price of $2.00 per million input tokens and $10.00 per million output, rolling out first to a named cohort of cyber defenders through a program Goo​gle calls Fairwind. It has no published API model identifier, no general-availability date, and no endpoint a customer can call today. GPT-6 Astra has been purchasable since 3 September 2026, and as of 7 October 2026 the GPT-6 family is also what ChatGPT routes more than 1.2 billion weekly users to — with GPT-6 Sol, not Astra, powering the paid tiers of that rollout.

Quindi il confronto è reale, ma asimmetrico in un modo che cambia come dovrebbe essere letto. I numeri di Argon descrivono un modello in un rollout controllato. Quelli di GPT-6 descrivono un prodotto con un listino prezzi. Tutto ciò che segue vale con la clausola "se potessi acquistarlo", e nulla di ciò che segue risponde a "dovresti passare", perché non c'è ancora nulla a cui passare.

Che cosa si sa effettivamente della controparte che esiste

Inizia dal lato che esiste, perché il confronto ha solo una metà su cui si può agire. GPT-6 viene distribuito in tre livelli, e solo due di essi contano qui:

• GPT-6 Astra — the flagship, model id gpt-6-astra, released 3 September 2026, 1,050,000-token context, 128,000-token output, text/image/file input, effort from low through max, $10.00 per million input and $50.00 output, repricing to $20.00/$75.00 for the whole request above 272,000 input tokens
• GPT-6 Sol — the mid tier, released 22 September 2026, same context and output ceilings, $2.00/$10.00 with the $4.00/$15.00 long-context step, and the model Ope​nAI put under ChatGPT Plus, Pro, Business and Enterprise on 7 October 2026
• GPT-6 Luna — the budget tier at $0.10/$0.50, serving ChatGPT Free and Go

Argon's side of that list is one line long. Goo​gle's announcement gave an introductory price, a capability framing — real-world coding, enterprise knowledge work and cyber defense — and a rollout plan. It did not give a model string, a context window, a maximum output, a cached-input rate, a deprecation schedule or a date for wider access. The absence of a model identifier is the practical one: any code sample you see online quoting a Gemini 4 Argon model string is guessing, because there is no string to quote.

L'unica misurazione pienamente comparabile, e ciò che nasconde

Un'unica metrica ti consente di confrontare Argon con un modello GPT-6 senza alcun «se», perché Artificial Analysis li ha eseguiti entrambi: il costo di produzione di una risposta completa sulla suite di valutazione.

• Costo per attività di indicizzazione completata — Gemini 4 Argon 1,99 $ vs GPT-6 Astra 3,26 $, una forbice di 1,6× a favore di Argon
• Token di output generati nell'intera suite — Gemini 4 Argon 112,8 M vs GPT-6 Astra 108,8 M, praticamente alla pari
• Intelligence Index, v4.3.2 — Gemini 4 Argon 52,6 vs GPT-6 Astra 52,7, una parità assoluta
• Prezzo dichiarato — 2,00 $ in input e 10,00 $ in output per 1M su Argon, come tariffa introduttiva, rispetto alla tariffa standard di Astra di 10,00 $/50,00 $
• Finestra di contesto — GPT-6 Astra 1.050.000 token vs quella di Argon, non pubblicata

Rileggi quella prima riga, perché è l'unica sorpresa di questa pagina. Le tariffe principali di Argon sono un quinto di quelle di Astra, e il suo costo per attività sulla stessa suite è inferiore del 39%. Un modello che ottiene lo stesso punteggio del modello di punta, a un quinto della tariffa di input, sarebbe la notizia del trimestre, se la si può chiamare così.

A two-column comparison scoreboard for GPT-6 Astra and Gemini 4 Argon, showing GPT-6 Astra at an Intelligence Index of 52.7, $3.26 per index task and a 1,050,000-token context window, against Gemini 4 Argon at 52.6, $1.99 and dimensions marked not published, with prices of $10.00/$50.00 against a stated introductory $2.00/$10.00 and an availability row reading phased rollout only. A footer reads "Index figures per Artificial Analysis v4.3.2; Argon figures vendor-stated, no purchase path."

La seconda riga è il motivo per cui la prima è meno drammatica di quanto sembri, e va nella direzione opposta rispetto al solito argomento sulla verbosità. Argon e Astra hanno scritto quasi esattamente lo stesso numero di token per produrre quasi esattamente lo stesso punteggio. Non c'è alcun divario di verbosità a spiegare la differenza di costo — la differenza di 1,27 $ deriva dal listino prezzi, non dal fatto che un modello si dilunghi. Questo rende il confronto più pulito della maggior parte in questo batch, e rende la disponibilità mancante l'unica cosa che separa Argon da una raccomandazione immediata.

Goo​gle's own table is not the independent one

Goo​gle published a launch table for Argon comparing it against GPT-6 Astra across nineteen benchmarks. If you read that table alone, Argon wins fourteen, Astra wins four, and one is a tie. That is a striking result and it should be handled carefully: every number in it comes from Goo​gle, selected and ordered by Goo​gle, and none of it has been independently reproduced. It is Goo​gle's case for Argon, which is what a launch table is for, and it belongs in a comparison as a labelled vendor claim rather than as evidence.

Set it against the independent run and the picture changes character. On Artificial Analysis's suite the two are level at 52.6 and 52.7 — a much smaller advantage than a 14-to-4 sweep implies, on a suite neither vendor assembled. The honest summary is that Argon is plausibly Astra's equal and possibly its better on software engineering, that the vendor's own table imagines a wider gap than the independent one finds, and that until someone outside Goo​gle runs it on something other than Goo​gle's tasks, none of it settles.

A screenshot of the top of the Artificial Analysis leaderboard table under the Model, Context Window, Creator, Intelligence Index, Cost per Task, Tokens/s, First Chunk and Response headings, showing Claude Opus 5.5 (max with fallback) at 58 and $5.98, Claude Sonnet 5.5 at 56, Claude Fable 5.1 at 53, then GPT-6 Astra (max) at 53 and $3.26 on the row immediately above Gemini 4 Argon (high) at 53 and $1.99, with GPT-6 Astra (xhigh) and GPT-6.1 Sol (max) at 52 below them - the near-tie rendered as adjacent rows at the same index.

Perché un modello non ancora rilasciato merita comunque una sezione

Because Goo​gle's rollout pattern is itself information, and it points at where the next Pro-tier release is going. Goo​gle spent 2026 shipping Flash-tier models — the 3.5, 3.6 and 3.8 Flash line, the Lite variants, a cybersecurity-tuned 3.8 Flash — while the Pro tier sat on Gemini 3.1 Pro Preview, a model that has been in preview since February 2026 with no general-availability commitment and no shutdown date. Argon is the first movement at the top of the line in seven months, and it went out first to defenders rather than to developers.

That sequencing is a coherent choice — cybersecurity is the workload where a frontier model with agentic tool use and long-horizon autonomy has the clearest, most measurable value, and it is also the workload where a vendor wants a controlled cohort before general release. It is not evidence that the model is unready. It is evidence that Goo​gle is treating general availability as a later decision rather than a launch-day one, which is exactly what the missing model identifier and the missing date say in a different way.

Per chi sviluppa, la conseguenza è semplice: non puoi pianificare tenendo conto di Argon. Puoi pianificare tenendo conto di GPT-6 Astra o GPT-6 Sol, perché entrambi hanno identificatori, endpoint, listini prezzi e un fornitore che ha già pubblicato una politica di dismissione per la linea. Un modello senza identificatore è un modello per cui non puoi scrivere un adattatore.

Cosa cambierebbe questo confronto

Three things, and any one of them turns the article above into a settled question. The first is a model identifier — a real string served from Goo​gle's API, because that is the point at which an adapter becomes writable and an integration estimate becomes meaningful. The second is a general-availability date, which is what distinguishes a preview from a product and is the one date Argon's announcement omitted entirely. The third is independent benchmarking on tasks Goo​gle did not choose; a suite assembled by the vendor is a claim, and a suite assembled by someone else is a measurement.

Until all three exist, Argon's role in a comparison is as a ceiling rather than as an option. Its numbers tell you where the Gemini Pro tier is heading and give you a sense of how much price room Goo​gle has at the top of the line. They do not tell you what to build on.

Cosa fare mentre Argon aspetta

Se il motivo per cui sei qui è che i numeri di Argon sembrano buoni, la mossa utile è testare la stessa classe di carico di lavoro sul modello che puoi effettivamente chiamare oggi, e per l'ingegneria del software e il lavoro agentico a lungo orizzonte questo significa GPT-6 Astra. Si trova nel catalogo di OrcaRouter con lo stesso schema di prezzi pass-through di Google — i $10,00/$50,00 del fornitore con il livello long-context da $20,00/$75,00, a 0% di markup, così un cambiamento delle tariffe del fornitore ti raggiunge lo stesso giorno anziché al ciclo di fatturazione successivo. Anche GPT-6 Sol a $2,00/$10,00 è disponibile, e per la maggior parte dei carichi di lavoro è l'acquisto migliore: ottiene 47,6 contro il 52,7 di Astra, e costa circa un terzo per attività completata.

Il livello di routing è ciò che rende sostenibile il calendario di distribuzione di un fornitore. Quando Argon ottiene un identificatore, non deve per forza diventare un progetto di migrazione: una regola di routing può inviargli una quota del tuo traffico, oppure usare la configurazione di fusione dei modelli per eseguire un panel e confrontare le risposte, mentre GPT-6 Astra o Sol resta l'impostazione predefinita alle sue spalle. Il failover automatico copre il caso che qui conta di più — un modello in una distribuzione controllata è esattamente il tipo di route che può essere soggetta a limiti di frequenza o ritirata senza preavviso, e un percorso di fallback fa sì che diventi una richiesta lenta invece di un'interruzione del servizio.

A screenshot of the OrcaRouter model page for openai/gpt-6-astra, showing the OpenAI vendor label, a 2026-09-04 release date, a 1,050,000-token context window, a 128K-token maximum output, text, image and file input, and pricing of $10.00 per million input tokens and $50.00 per million output tokens.

不该做的,是围绕一个没有端点的模型的发布表来启动迁移计划。与 GPT-6 Astra 的差距只有 0.1 分,价格差异也真实存在,但这两者都不值得为一个你无法向其发送请求的东西重建集成。

Confrontati in questo articolo1

Rilevato da questo articolo · Benchmark: Artificial Analysis · aggiornato ogni giorno