A generated title card reading 'AREX-2 vs Gemma 4 12B - One of These Is a Model', with two panels. The left panel, headed Gemma 4 12B, lists: shipped 3 June 2026; 11.95B params, dense; 256K context, Apache 2.0; download it today. The right panel, headed AREX-2, lists: repo created 29 Sep 2026; no weights uploaded; no card, no licence tag; not downloadable. A footer reads 'One column is a model. The other is a timestamp.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

AREX-2 vs Gemma 4 12B: One of These Is a Model

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

A comparison between Gemma 4 12B and AREX-2 has a property no other matchup on this blog has: one of the two columns is empty because the model on that side does not exist yet. Gemma 4 12B is a shipped artifact — Goog​le DeepMind published it on June 3, 2026 as an 11.95-billion-parameter dense open-weights model under Apache 2.0, with a full model card, a 256K-token context window, native audio and image input, and three and a half months of independent testing behind it. AREX-2 is a Hugging Face repository that BAAI created on September 29, 2026 at 17:56 UTC and has not yet put anything into: no weights, no card, no licence tag, no evaluation, no announcement.

That asymmetry is not a reason to skip the comparison. It is the comparison. The question worth answering is not which of these is better — nothing can answer that until AREX-2 has files in it — but whether a second-generation BAAI research agent would even be competing with a 12B edge model in the first place, or whether these two occupy positions in a stack that never touch. Getting that wrong is the expensive mistake, because it is the mistake that leads a team to budget for the wrong thing.

The shipped side, on its own terms

Gemma 4 12B is the small member of Goog​le's fourth-generation open family, and its defining design choice is architectural: it drops the separate vision encoder entirely and feeds raw image patches and audio waveforms straight into the transformer. That is not a benchmark trick. It is what makes the model genuinely portable — roughly 27 GB of BF16 weights that quantise below 7 GB, and a card that names 16 GB of VRAM as the deployment target. On a laptop or a single mid-range card, this is a model you can actually hold.

The capability profile follows from that. Goog​le's own card — vendor-reported, and worth labelling as such — lists DocVQA 94.9 and InfoVQA 88.4 for document understanding, MMLU-Pro 77.2, GPQA Diamond 78.8, AIME 2026 77.5 and LiveCodeBench v6 72.0 in thinking mode. Artificial Analysis, which is independent, scores the reasoning configuration at an Intelligence Index of 14, measures it at 114 tokens per second, and prices a typical served call at $0.10 per million input tokens and $0.30 per million output. Our own catalogue entry for the larger Gemma 4 31B carries 85.7 on GPQA Diamond. The shape of that is a small generalist that is unusually strong on documents and images, reasonable on reasoning, and nowhere near a frontier model on hard multi-step work.

Its job description is therefore concrete: local or single-tenant inference, document and screenshot understanding, modest agentic loops, and anything where the data cannot leave the building. It is not a deep-research engine and Goog​le does not sell it as one.

A screenshot of the Artificial Analysis model page for Gemma 4 12B (Reasoning), showing an Artificial Analysis Intelligence Index of 14, a speed rank of 19 of 142, $0.10 per million input tokens and $0.30 per million output, a 256K-token context window, text, image, speech and video input with text output, and the note that it is among the leading models in intelligence but somewhat expensive for its size. The page header reads 'Released June 2026' and 'Open weights model'.

The other side, which is a name and a timestamp

Everything confirmable about AREX-2 this week fits in one sentence. It is a repository under the BAAI organisation on Hugging Face, created on 29 September 2026, containing a .gitattributes file and nothing else — zero bytes of stored model data, zero downloads, no model card, no licence declaration, no provider deployment. BAAI itself has announced nothing.

What makes it worth writing down is the family it is named after. BAAI's AREX line shipped on 23 July 2026 with two models: AREX-Base, a 122-billion-parameter mixture-of-experts with 10 billion active parameters based on Qwen3.5-122B-A10B, and AREX-Turbo, a dense 4B built on Qwen3.5-4B. Both are Apache 2.0. Both are deep-research agents rather than chat models — an inner loop that gathers evidence and produces a candidate answer with a confidence figure, and an outer loop that checks that answer against the original constraints and can refine or restart the whole trajectory. On BAAI's published numbers, AREX-Base reaches 82.5 on BrowseComp, 85.4 on GAIA and 89.9 on DeepSearch QA; AREX-Turbo reaches 70.7, 81.6 and 78.5 on the same three.

All of those figures are the vendor's own and none have been independently reproduced. They are also about a different model. AREX-2 has no figures, and the "2" tells you nothing about size, backbone or architecture — a second generation in this line could be a bigger agent, a smaller distillation, or a retrained framework with the backbone swapped out.

Do these two even meet?

Here is where the comparison gets more useful than it looks. Take the two plausible readings of what AREX-2 becomes.

If it is a second-generation research agent at anything like the Base's scale, then it and Gemma 4 12B never compete. A 122B mixture-of-experts driving multi-source search is a hosted, per-token-billed, network-bound service; Gemma 4 12B is a checkpoint you download and run yourself. The decision between them is not a quality comparison, it is a decision about whether your workload is a research loop or an inference endpoint.

If AREX-2 turns out to be the small tier — the successor to the 4B Turbo rather than the 122B Base — then the two do share a shelf, and the comparison becomes a real one. That version of AREX-2 would be roughly a third of Gemma 4 12B's parameter count, specialised for tool-driven search rather than multimodal breadth, and would be measured on BrowseComp and GAIA rather than on DocVQA. Two small models, one optimised for holding a long research trajectory, the other for seeing and reading on modest hardware. A team building a documentation pipeline should not care about the first; a team building a research agent should not care about the second.

• What exists — Gemma 4 12B is downloadable today with a full card. AREX-2 is a repository with no files in it.

• Parameters — 11.95B dense for Gemma 4 12B. Not stated, and inferable only from a product name, for AREX-2.

• Context — 256K tokens (262,144 positions) for Gemma 4 12B. Unknown for AREX-2.

• Modality — text, image and audio in for Gemma 4 12B. Unknown for AREX-2; the first AREX generation was text-plus-tool-use.

• Licence — Apache 2.0 for Gemma 4 12B, stated on the card. Not stated for AREX-2; AREX-Base and AREX-Turbo were Apache 2.0.

• What it is for — deployable multimodal inference on your own hardware versus, on the evidence of the family, long-horizon search and evidence aggregation against a hosted endpoint.

A generated two-column card headed 'Two deployment shapes, and only one of them exists'. It contrasts where Gemma 4 12B runs (local or single-tenant; about 27 GB BF16, under 7 GB quantised, 16 GB VRAM target; no rate card or per-token meter) against where AREX-Base runs (a 122B mixture-of-experts, a cluster or a hosted endpoint billed per token), names tokens per search step times steps per trajectory as the cost driver for the AREX shape, and notes that AREX-2 is neither shape because there are no files in the repository. A footer reads 'Is this a quality comparison? Not until one side has a model in it.' The OrcaRouter logo is composited in the bottom-right corner.

The part that is actually comparable: where each one runs

Since the capability comparison is unavailable, the deployment comparison is the honest one to make, and it is not close.

Gemma 4 12B runs where you put it. There is no rate card, no per-token meter and no provider between you and the model — the cost is the hardware and the electricity, and the failure mode is your own capacity planning. When AREX-Base shipped in July, by contrast, it was a 122B mixture-of-experts: a model you serve on a cluster or rent by the token, and the family's own 4B Turbo exists precisely because that choice has a price. If the second generation follows the same shape, AREX-2's economics will be a function of tokens consumed per query multiplied by the number of search steps a trajectory takes — a number that is much larger for a deep-research agent than for a document-reading model, because the agent re-reads a growing context on every iteration.

That is where a routing layer stops being an abstraction and starts being a line item. If you end up running a research agent against a hosted model while keeping a local Gemma 4 12B for the document work, those are two integrations in the ordinary case. Through one API in front of 200-plus models — with provider list price passed through and no markup added on top, so a vendor's price change lands the same day — they are one key, one bill and one client, and a failover rule can move a request off a provider that is having a bad hour without your agent loop knowing. We route Gemma 4 26B-A4B and Gemma 4 31B today; we do not route the 12B, and nothing from the AREX line is on our catalogue either, so if either of those is the model you need, it comes from its vendor's own distribution.

What to do this week

If you need a compact multimodal model you can run yourself, the decision was never waiting on AREX-2. Gemma 4 12B is shipped, documented, independently scored and deployable, and the comparison column on the other side is empty. Buy nothing on the strength of a reserved name.

If you are building a deep-research agent and the AREX line is what you actually want, the thing to watch is not the name but the tree: whether safetensors appear in that repository, what the card says the parameter count is, and which licence tag lands on it. The first generation shipped Apache 2.0, and if the second does too, the question becomes a hosting question rather than a licensing one. Until then, AREX-Base is the AREX model you can actually run, and it has been available since July.

The one thing worth saying plainly about the comparison that this article is nominally about: a VS page where one side is a repository commit and the other is a shipped model is not a matchup, it is a schedule. The schedule says Gemma 4 12B is available now and AREX-2 is not available at all.