主视觉标题卡显示“GPT-6 对比 Gemini 4 Argon”,一个标签显示“Argon:自 2026-09-30 起分阶段向安全合作伙伴推出”,两个指数标签分别显示“GPT-6 Astra 52.7”和“Gemini 4 Argon 52.6”,页脚显示“指数数据依据 Artificial Analysis v4.3.2;Argon 规格由厂商提供,且无法购买。”OrcaRouter 徽标合成于右下角。
Guides & Insights

GPT-6 与 Gemini 4 Argon:一个在价目表上,一个在等待名单上,两者近乎打成平手

作者

Elias Hawthorne

发布日期

最新模型 · 20查看全部模型 →
基准测试:Artificial Analysis · 每日更新
返回全部文章

The closest model pair on the current frontier board is not a pair you can buy. Gemini 4 Argon scores 52.6 on Artificial Analysis's Intelligence Index v4.3.2. GPT-6 Astra, Ope​nAI's flagship, scores 52.7. That is a one-tenth-of-a-point gap across ten evaluations, and both figures come from the same revision of the same suite. It is the tightest top-of-table pairing in the set.

There is a symmetry in that number and none at all in availability. Goo​gle announced Gemini 4 Argon on 30 September 2026, at a stated introductory price of $2.00 per million input tokens and $10.00 per million output, rolling out first to a named cohort of cyber defenders through a program Goo​gle calls Fairwind. It has no published API model identifier, no general-availability date, and no endpoint a customer can call today. GPT-6 Astra has been purchasable since 3 September 2026, and as of 7 October 2026 the GPT-6 family is also what ChatGPT routes more than 1.2 billion weekly users to — with GPT-6 Sol, not Astra, powering the paid tiers of that rollout.

所以这个比较是真实的,但它的不对称方式会改变它应有的解读方式。Argon 的数字描述的是一个处于受控发布中的模型。GPT-6 的数字描述的则是一个附带价目表的产品。以下所有内容都带有“如果你能买到的话”这一前提,而以下内容都没有回答“你是否应该切换”,因为目前还没有可供切换的东西。

关于存在的对应物,实际上已知些什么?

从存在的那一侧开始,因为对比中只有一半是可行动的。GPT-6 以三个层级推出,而这里只有其中两个重要:

• GPT-6 Astra — the flagship, model id gpt-6-astra, released 3 September 2026, 1,050,000-token context, 128,000-token output, text/image/file input, effort from low through max, $10.00 per million input and $50.00 output, repricing to $20.00/$75.00 for the whole request above 272,000 input tokens
• GPT-6 Sol — the mid tier, released 22 September 2026, same context and output ceilings, $2.00/$10.00 with the $4.00/$15.00 long-context step, and the model Ope​nAI put under ChatGPT Plus, Pro, Business and Enterprise on 7 October 2026
• GPT-6 Luna — the budget tier at $0.10/$0.50, serving ChatGPT Free and Go

Argon's side of that list is one line long. Goo​gle's announcement gave an introductory price, a capability framing — real-world coding, enterprise knowledge work and cyber defense — and a rollout plan. It did not give a model string, a context window, a maximum output, a cached-input rate, a deprecation schedule or a date for wider access. The absence of a model identifier is the practical one: any code sample you see online quoting a Gemini 4 Argon model string is guessing, because there is no string to quote.

唯一完全可比的度量,以及它隐藏了什么

一项指标让你能把 Argon 与 GPT-6 模型直接比较,不附带任何“如果”;Artificial Analysis 对两者都进行了测试:在评估套件上生成一个完整答案的成本。

• 每项已完成索引任务的成本——Gemini 4 Argon $1.99,GPT-6 Astra $3.26,Argon 以 1.6 倍的差距占优
• 整套测试生成的输出 token 总量——Gemini 4 Argon 112.8M,GPT-6 Astra 108.8M,基本持平
• 智能指数 v4.3.2——Gemini 4 Argon 52.6,GPT-6 Astra 52.7,难分伯仲
• 标价——Argon 每 1M 输入 $2.00、输出 $10.00,为限时引入价,而 Astra 的标准价为 $10.00/$50.00
• 上下文窗口——GPT-6 Astra 1,050,000 个 token,而 Argon 未公布

再读一遍第一行,因为这是本页唯一的意外之处。Argon 的标价只有 Astra 的五分之一,而且它在同一套测试上的每任务成本低 39%。一个得分与旗舰模型相同、输入费率却只有五分之一的模型,本该是本季度的大新闻——如果这还能算作大新闻的话。

A two-column comparison scoreboard for GPT-6 Astra and Gemini 4 Argon, showing GPT-6 Astra at an Intelligence Index of 52.7, $3.26 per index task and a 1,050,000-token context window, against Gemini 4 Argon at 52.6, $1.99 and dimensions marked not published, with prices of $10.00/$50.00 against a stated introductory $2.00/$10.00 and an availability row reading phased rollout only. A footer reads "Index figures per Artificial Analysis v4.3.2; Argon figures vendor-stated, no purchase path."

第二行说明了为什么第一行没有看起来那么戏剧化,而且它的指向与通常关于冗长的论点相反。Argon 和 Astra 生成的 token 数量几乎完全相同,以产生几乎完全相同的得分。并没有冗长程度上的差距可以解释成本差异——这 1.27 美元的差距来自费率表,而不是某个模型啰嗦。这使得它比这批中的大多数比较都更清晰,也使得可用性缺失成为 Argon 与一个明确推荐之间唯一的障碍。

Goo​gle's own table is not the independent one

Goo​gle published a launch table for Argon comparing it against GPT-6 Astra across nineteen benchmarks. If you read that table alone, Argon wins fourteen, Astra wins four, and one is a tie. That is a striking result and it should be handled carefully: every number in it comes from Goo​gle, selected and ordered by Goo​gle, and none of it has been independently reproduced. It is Goo​gle's case for Argon, which is what a launch table is for, and it belongs in a comparison as a labelled vendor claim rather than as evidence.

Set it against the independent run and the picture changes character. On Artificial Analysis's suite the two are level at 52.6 and 52.7 — a much smaller advantage than a 14-to-4 sweep implies, on a suite neither vendor assembled. The honest summary is that Argon is plausibly Astra's equal and possibly its better on software engineering, that the vendor's own table imagines a wider gap than the independent one finds, and that until someone outside Goo​gle runs it on something other than Goo​gle's tasks, none of it settles.

A screenshot of the top of the Artificial Analysis leaderboard table under the Model, Context Window, Creator, Intelligence Index, Cost per Task, Tokens/s, First Chunk and Response headings, showing Claude Opus 5.5 (max with fallback) at 58 and $5.98, Claude Sonnet 5.5 at 56, Claude Fable 5.1 at 53, then GPT-6 Astra (max) at 53 and $3.26 on the row immediately above Gemini 4 Argon (high) at 53 and $1.99, with GPT-6 Astra (xhigh) and GPT-6.1 Sol (max) at 52 below them - the near-tie rendered as adjacent rows at the same index.

为什么未发布的模型仍然值得专门用一节来介绍

Because Goo​gle's rollout pattern is itself information, and it points at where the next Pro-tier release is going. Goo​gle spent 2026 shipping Flash-tier models — the 3.5, 3.6 and 3.8 Flash line, the Lite variants, a cybersecurity-tuned 3.8 Flash — while the Pro tier sat on Gemini 3.1 Pro Preview, a model that has been in preview since February 2026 with no general-availability commitment and no shutdown date. Argon is the first movement at the top of the line in seven months, and it went out first to defenders rather than to developers.

That sequencing is a coherent choice — cybersecurity is the workload where a frontier model with agentic tool use and long-horizon autonomy has the clearest, most measurable value, and it is also the workload where a vendor wants a controlled cohort before general release. It is not evidence that the model is unready. It is evidence that Goo​gle is treating general availability as a later decision rather than a launch-day one, which is exactly what the missing model identifier and the missing date say in a different way.

对于构建者来说,后果很简单:你无法围绕 Argon 做规划。你可以围绕 GPT-6 Astra 或 GPT-6 Sol 做规划,因为两者都有标识符、端点、费率卡,以及一个已经针对该产品线发布了退役政策的供应商。一个没有标识符的模型,就是一个你无法为其编写适配器的模型。

什么会改变这个比较?

Three things, and any one of them turns the article above into a settled question. The first is a model identifier — a real string served from Goo​gle's API, because that is the point at which an adapter becomes writable and an integration estimate becomes meaningful. The second is a general-availability date, which is what distinguishes a preview from a product and is the one date Argon's announcement omitted entirely. The third is independent benchmarking on tasks Goo​gle did not choose; a suite assembled by the vendor is a claim, and a suite assembled by someone else is a measurement.

Until all three exist, Argon's role in a comparison is as a ceiling rather than as an option. Its numbers tell you where the Gemini Pro tier is heading and give you a sense of how much price room Goo​gle has at the top of the line. They do not tell you what to build on.

Argon 等待期间该做什么

如果你来到这里是因为 Argon 的数字看起来不错,那么更有用的做法,是用你如今真正能调用的模型去测试同一类工作负载;而对于软件工程和长周期智能体任务来说,那就是 GPT-6 Astra。它就在 OrcaRouter 的目录里,采用 Google 自己的直通定价模式——厂商价 $10.00/$50.00,长上下文档为 $20.00/$75.00,0% 加价,因此厂商费率一变,当天就能传导到你这里,而不用等到下一个计费周期。定价 $2.00/$10.00 的 GPT-6 Sol 也在其中,而且对大多数工作负载来说它更值得买:它的得分是 47.6,而 Astra 是 52.7,并且每完成一个任务的成本大约只有三分之一。

路由层正是让供应商的发布计划变得可承受的关键。当 Argon 真的拿到标识符时,它不必变成一项迁移工程:一条路由规则可以把一部分流量发给它,或者用模型融合配置来运行一个评审组并对比答案,而 GPT-6 Astra 或 Sol 仍作为其背后的默认选项。自动故障转移覆盖的正是这里最关键的情形——处于受控发布中的模型,恰恰是最容易被限流或在毫无预警下被撤下的那类路由,而一条回退路径意味着它会变成一次慢请求,而不是一次服务中断。

A screenshot of the OrcaRouter model page for openai/gpt-6-astra, showing the OpenAI vendor label, a 2026-09-04 release date, a 1,050,000-token context window, a 128K-token maximum output, text, image and file input, and pricing of $10.00 per million input tokens and $50.00 per million output tokens.

不该做的是,围绕一个没有端点的模型的发布表去启动迁移计划。与 GPT-6 Astra 的差距只有十分之一分,价格差异也确实存在,但两者都不值得为一个你无法向其发送请求的东西重建集成。

本文中的对比1

根据本文内容识别 · 基准测试:Artificial Analysis · 每日更新