一张生成的标题卡,上面写着“AREX-2 vs Gemma 4 12B——其中一个才是模型”,分为左右两个面板。左侧面板标题为 Gemma 4 12B,列出:2026 年 6 月 3 日发布;11.95B 参数,稠密模型;256K 上下文,Apache 2.0;今天即可下载。右侧面板标题为 AREX-2,列出:仓库创建于 2026 年 9 月 29 日;未上传权重;没有模型卡,没有许可证标签;无法下载。页脚写着“一列是模型,另一列是时间戳。”OrcaRouter 徽标合成在右下角。
Guides & Insights

AREX-2 与 Gemma 4 12B:其中之一是模型

作者

Gideon Frost

发布日期

最新模型 · 20查看全部模型 →
基准测试:Artificial Analysis · 每日更新
返回全部文章

A comparison between Gemma 4 12B and AREX-2 has a property no other matchup on this blog has: one of the two columns is empty because the model on that side does not exist yet. Gemma 4 12B is a shipped artifact — Goog​le DeepMind published it on June 3, 2026 as an 11.95-billion-parameter dense open-weights model under Apache 2.0, with a full model card, a 256K-token context window, native audio and image input, and three and a half months of independent testing behind it. AREX-2 is a Hugging Face repository that BAAI created on September 29, 2026 at 17:56 UTC and has not yet put anything into: no weights, no card, no licence tag, no evaluation, no announcement.

这种不对称并不是跳过比较的理由。它本身就是比较。值得回答的问题不是这些当中哪个更好——在 AREX-2 里有文件之前,没有什么能回答这一点——而是第二代 BAAI 研究智能体究竟会不会与 12B 边缘模型竞争,或者这两者是否占据着技术栈中永不相交的位置。搞错这一点才是代价高昂的错误,因为正是这个错误会让团队为错误的东西做预算。

交付侧,以其自身标准

Gemma 4 12B is the small member of Goog​le's fourth-generation open family, and its defining design choice is architectural: it drops the separate vision encoder entirely and feeds raw image patches and audio waveforms straight into the transformer. That is not a benchmark trick. It is what makes the model genuinely portable — roughly 27 GB of BF16 weights that quantise below 7 GB, and a card that names 16 GB of VRAM as the deployment target. On a laptop or a single mid-range card, this is a model you can actually hold.

The capability profile follows from that. Goog​le's own card — vendor-reported, and worth labelling as such — lists DocVQA 94.9 and InfoVQA 88.4 for document understanding, MMLU-Pro 77.2, GPQA Diamond 78.8, AIME 2026 77.5 and LiveCodeBench v6 72.0 in thinking mode. Artificial Analysis, which is independent, scores the reasoning configuration at an Intelligence Index of 14, measures it at 114 tokens per second, and prices a typical served call at $0.10 per million input tokens and $0.30 per million output. Our own catalogue entry for the larger Gemma 4 31B carries 85.7 on GPQA Diamond. The shape of that is a small generalist that is unusually strong on documents and images, reasonable on reasoning, and nowhere near a frontier model on hard multi-step work.

Its job description is therefore concrete: local or single-tenant inference, document and screenshot understanding, modest agentic loops, and anything where the data cannot leave the building. It is not a deep-research engine and Goog​le does not sell it as one.

A screenshot of the Artificial Analysis model page for Gemma 4 12B (Reasoning), showing an Artificial Analysis Intelligence Index of 14, a speed rank of 19 of 142, $0.10 per million input tokens and $0.30 per million output, a 256K-token context window, text, image, speech and video input with text output, and the note that it is among the leading models in intelligence but somewhat expensive for its size. The page header reads 'Released June 2026' and 'Open weights model'.

另一侧是名称和时间戳

本周关于 AREX-2 一切可证实的信息,用一句话就能说完。它是 Hugging Face 上 BAAI 组织下的一个仓库,创建于 2026 年 9 月 29 日,其中仅有一个 .gitattributes 文件,别无他物——零字节的模型数据存储、零下载量,没有模型卡,没有许可声明,没有任何提供方部署。BAAI 本身也未发布任何公告。

真正让它值得写下来的是它得名于的那个家族。BAAI 的 AREX 系列于 2026 年 7 月 23 日发布,包含两个模型:AREX-Base,一个基于 Qwen3.5-122B-A10B 的 1220 亿参数混合专家模型,其中 100 亿为活跃参数;以及 AREX-Turbo,一个基于 Qwen3.5-4B 构建的稠密 4B 模型。两者均采用 Apache 2.0 许可。两者都是深度研究智能体,而非聊天模型——内循环负责收集证据并生成带有置信度数值的候选答案,外循环则对照原始约束检查该答案,并可改进或重启整个轨迹。根据 BAAI 公布的数据,AREX-Base 在 BrowseComp 上达到 82.5,在 GAIA 上达到 85.4,在 DeepSearch QA 上达到 89.9;AREX-Turbo 在同样三项上分别达到 70.7、81.6 和 78.5。

所有这些数据都来自厂商自己,没有任何一个经过独立复现。它们涉及的还是另一个模型。AREX-2 没有任何数据,而“2”也无法说明规模、主干或架构——这一产品线的第二代可能是一个更大的智能体、一个更小的蒸馏版,或者一个更换了主干的重新训练框架。

这两个到底会不会碰面?

正是在这里,比较变得比表面上更有用。以 AREX-2 将会演变成什么的两种合理解读为例。

如果它是一个第二代研究代理,且规模哪怕只是接近 Base 的程度,那么它和 Gemma 4 12B 根本不会竞争。一个驱动多源搜索的 122B 混合专家模型是一项托管式、按 token 计费、受网络约束的服务;Gemma 4 12B 则是你自己下载并运行的检查点。二者之间的抉择不是质量比较,而是关于你的工作负载是研究循环还是推理端点的决定。

如果 AREX-2 最终被归为小规模档位——是 4B Turbo 的后继者,而非 122B Base 的后继者——那么两者确实同处一个货架,这种比较才真正成立。那个版本的 AREX-2 参数量大约只有 Gemma 4 12B 的三分之一,专门针对工具驱动的搜索而非多模态广度进行优化,并且会在 BrowseComp 和 GAIA 上而非 DocVQA 上接受评估。两个小模型,一个为维持长期研究轨迹而优化,另一个为在普通硬件上进行观察和阅读而优化。构建文档流水线的团队不应关心前者;构建研究智能体的团队不应关心后者。

• 已存在的内容 — Gemma 4 12B 今日即可下载,并附有完整的模型卡片。AREX-2 是一个其中没有任何文件的代码仓库。

• 参数 — 11.95B 稠密,对应 Gemma 4 12B。AREX-2 未予说明,仅可从产品名称推断。

• 上下文——Gemma 4 12B 为 256K tokens(262,144 个位置)。AREX-2 未知。

• 模态——Gemma 4 12B 支持文本、图像和音频输入。AREX-2 的情况尚不明确;第一代 AREX 为文本加工具调用。

• 许可证——Gemma 4 12B 采用 Apache 2.0,卡片上已注明。AREX-2 未注明;AREX-Base 和 AREX-Turbo 曾采用 Apache 2.0。

• 用途——在自有硬件上可部署的多模态推理,对比根据该系列模型的证据、针对托管端点进行的长时程搜索与证据聚合。

A generated two-column card headed 'Two deployment shapes, and only one of them exists'. It contrasts where Gemma 4 12B runs (local or single-tenant; about 27 GB BF16, under 7 GB quantised, 16 GB VRAM target; no rate card or per-token meter) against where AREX-Base runs (a 122B mixture-of-experts, a cluster or a hosted endpoint billed per token), names tokens per search step times steps per trajectory as the cost driver for the AREX shape, and notes that AREX-2 is neither shape because there are no files in the repository. A footer reads 'Is this a quality comparison? Not until one side has a model in it.' The OrcaRouter logo is composited in the bottom-right corner.

真正具有可比性的部分:它们各自运行的位置

由于能力比较无从进行,部署层面的比较才是实事求是的比较,而两者差距悬殊。

Gemma 4 12B 在你部署它的地方运行。没有费率表,没有按 token 计量的仪表,你和模型之间也没有提供商——成本就是硬件和电力,而故障模式则取决于你自己的容量规划。相比之下,当 AREX-Base 于 7 月发布时,它是一个 122B 的专家混合模型:你要么在集群上部署它,要么按 token 租用它,而这个系列自己的 4B Turbo 之所以存在,正是因为这种选择是有代价的。如果第二代沿用同样的形态,AREX-2 的经济性将是每次查询消耗的 token 数乘以一条轨迹所经历的搜索步数的函数——对于深度研究智能体来说,这个数字要比文档阅读模型大得多,因为智能体在每次迭代时都会重新读取不断增长的上下文。

这正是路由层不再是一个抽象概念、而开始变成一项具体账单条目的地方。如果你一边针对托管模型运行研究型 agent,一边保留本地的 Gemma 4 12B 处理文档工作,那么通常情况下这就是两套集成。通过一个位于 200 多个模型之前的 API——供应商的标价原样透传、上面不加任何加价,因此供应商调价当天就能落地——它们就变成了一个密钥、一张账单和一个客户端,而一条故障转移规则就能把请求从某个当时状态不佳的供应商那里转走,你的 agent 循环对此毫无察觉。我们今天已经路由 Gemma 4 26B-A4B 和 Gemma 4 31B;我们不路由 12B,AREX 系列的任何模型也都不在我们的目录中,所以如果这两者中有一个正是你需要的模型,那它得来自其供应商自己的分发渠道。

本周做什么

如果你需要一个可以自行运行的紧凑型多模态模型,这个决定从来不必等 AREX-2。Gemma 4 12B 已经发布、有文档、经过独立评分并且可部署,而另一边的对比栏是空的。不要仅凭一个预留名称就去购买任何东西。

如果你正在构建深度研究智能体,而 AREX 系列正是你真正想要的,那么要关注的不是名字,而是仓库树:该仓库中是否有 safetensors,模型卡上写的参数量是多少,以及它被打上哪个许可证标签。第一代采用 Apache 2.0 发布;如果第二代也如此,那么问题就变成托管问题,而不是许可问题。在此之前,AREX-Base 是你能实际运行的 AREX 模型,而且自 7 月起就已可用。

关于本文名义上要做的这个对比,有一点必须直白地说清楚:一个 VS 页面,一边是代码仓库里的一次提交,另一边是已发布的模型——这根本算不上对决,而是一份时间表。这份时间表说的是:Gemma 4 12B 现在已经可用,而 AREX-2 则完全不可用。