A generated title card reading 'AREX-2 vs MathForm-8B - Two Ways to Be Checked', with two panels. The left panel, headed MathForm-8B, lists: 8B dense, shipped Aug 2026; writes Lean 4 theorem statements; checked by the Lean compiler; Pass@8 88.06% syntax-checked. The right panel, headed AREX-2, lists: repo created 29 Sep 2026; no files, no card; nothing to download or run; no figures exist to check. A footer reads 'One is graded by a compiler. The other has nothing to grade yet.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

AREX-2 vs MathForm-8B: A Proof Checker and a Search Loop Are Different Kinds of Verification

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

MathForm-8B and AREX-2 are the two most rigorously verifiable models in this series, and they verify in opposite ways. MathForm-8B is OpenBMB's 8-billion-parameter autoformalizer, released in August 2026: give it a mathematics problem written in plain English and it produces a formally correct Lean 4 statement of that problem, which a proof assistant's compiler then checks. AREX-2 is a Hugging Face repository that BAAI created on September 29, 2026 at 17:56 UTC, holding a .gitattributes and no weights, no model card, no licence tag and no announcement — a reserved name for a probable second generation of BAAI's AREX deep-research agents, whose first generation verifies itself by re-checking its own answer against the constraints it was given.

So the interesting question here is not which model is better. It is what it means for a model to be checkable, because one of these two is checked by a compiler that does not care what anyone believes, and the other is checked, if the family precedent holds, by a loop that the same model is running. That distinction survives even though one side of the matchup has not shipped, and it is the reason these two belong in the same conversation at all.

The side that exists, and the reason it is checkable

MathForm-8B has a narrow job description, and it is worth stating precisely because the narrowness is the point. It does not solve mathematical problems and does not claim to. It was fine-tuned on FormalVerse, a corpus of roughly 367,000 verified Lean 4 examples, and its training reward came from the Lean compiler rather than from a human preference model. Given an informal statement, it emits the imports, types and theorem header — leaving the proof obligation itself open — so that a compiler can confirm the formalization says what the informal problem said.

Its published figures, all vendor-reported by OpenBMB and none independently reproduced, are an average Pass@8 of 88.06% under a syntax check and 72.37% under a stricter consistency check across six benchmarks, dropping to 63% and 37% on the hardest FATE-H and FATE-X subsets. Read those as what they are: a model measured against a formal system, where a wrong answer is a failure to compile rather than a disagreement about quality. That is a rare property. Most benchmark claims on this blog rest on a grader that has to be trusted; this one rests on arithmetic that a machine performs.

The side that does not exist yet, and why its verification is different

AREX-2's only confirmable facts are the org, the name and the timestamp. Nothing has been uploaded and BAAI has said nothing. The family behind the name does exist, though, and it shipped on 23 July 2026 as AREX-Base — a 122-billion-parameter mixture-of-experts with 10 billion active parameters on a Qwen3.5-122B-A10B base — together with AREX-Turbo, a dense 4B. Both Apache 2.0, both agents rather than chat models.

The AREX design is a two-loop research framework, and the outer loop is a verification step. An inner loop gathers evidence from search and browsing, integrates it, and produces a candidate answer with an attached confidence figure. The outer loop then checks that candidate against the original constraints and decides: accept it, refine it, or discard the trajectory and restart. Between iterations the model maintains a context block of verified findings, open candidates, unresolved constraints and the next plan, which is what lets it follow a long line of enquiry without losing the thread. BAAI reports 82.5 on BrowseComp, 85.4 on GAIA and 89.9 on DeepSearch QA for the Base, and 70.7 / 81.6 / 78.5 for the Turbo — the vendor's own numbers, unreproduced.

That is a real and useful form of self-checking. It is also categorically weaker than a compiler. When the AREX outer loop decides an answer satisfies its constraints, the judgement comes from a language model reading the constraints. When Lean accepts a MathForm-8B formalization, the judgement comes from a type checker implementing a fixed calculus. Neither one is free of error — a formalization can be valid Lean that captures the wrong theorem — but only one of them can be wrong without anybody noticing, and the difference is not a matter of degree.

• Availability — MathForm-8B is downloadable from OpenBMB with a published card. AREX-2 is a repository with no files in it.

• Parameters — 8B dense for MathForm-8B. Unknown for AREX-2; the family spans a 122B mixture-of-experts and a dense 4B.

• What it produces — Lean 4 theorem statements for MathForm-8B, with the proof left open by design. Unknown for AREX-2; the first generation produced a researched answer with a confidence figure.

• How it is checked — the Lean compiler, for MathForm-8B. A verification loop inside the same model, on the evidence of AREX-Base.

• Licence — OpenBMB's own terms for MathForm-8B; nothing declared yet for AREX-2. The first AREX generation was Apache 2.0.

• Headline figures — 88.06% syntax-checked and 72.37% consistency-checked Pass@8 for MathForm-8B, vendor-reported. None exist for AREX-2.

A screenshot of the Hugging Face model card for openbmb/MathForm-8B, showing the OpenBMB author line, tags for Text Generation, Transformers, Safetensors, the openbmb/FormalVerse training dataset and arXiv 2608.14221, and the opening of the card describing MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement, with the note that the model is trained on FormalVerse through supervised fine-tuning followed by reinforcement learning using Lean compilation and semantic-consistency feedback.A generated card headed 'Two ways a model can be checked' with rows contrasting MathForm-8B (8B dense, shipped August 2026; emits a Lean 4 theorem statement with the proof obligation left open; checked by the Lean compiler, a type checker implementing a fixed calculus; published figures Pass@8 88.06% syntax-checked and 72.37% consistency-checked, vendor-reported, with FATE-H and FATE-X at 63% and 37%) against AREX-2 (repository created 29 Sep 2026 with no weights, card or licence tag; nothing to check yet), and a row noting the AREX-Base outer loop re-reads the agent's own answer against its constraints - a model reading a model - with BrowseComp 82.5, GAIA 85.4 and DeepSearch QA 89.9 for the 122B Base, unreproduced. A footer reads 'A compiler fails a wrong answer. A verification loop has to notice one.' The OrcaRouter logo is composited in the bottom-right corner.

Two jobs that a stack can hold at once

These models do not overlap in function, and it is worth saying so before anyone reads a "versus" as a choice. MathForm-8B is a front-end for a proof assistant. Its output is a theorem statement that a mathematician or an automated prover then works on. Its natural place is inside a verification pipeline for mathematics, formal methods, or specification work, where the value is that a downstream tool can consume the output without trusting the model.

AREX-2, if it continues its family, is a back-end for questions that have no compiler. "Which of these three filings is inconsistent with the other two" has no formal system to check it against, and the only available defence is gathering more evidence and re-examining your own reasoning — which is what the AREX outer loop is. A team that needs both would run them in different places: formalization for the part of the problem that can be made formal, and a search-and-verify agent for the part that cannot.

That split is also the honest answer to which of these you can act on this week. MathForm-8B is shipped, documented and usable now. AREX-2 is a name.

What the routing story looks like when verification is the subject

Because the two are used so differently, the infrastructure question splits as well.

For MathForm-8B the workload is a burst of short, deterministic generations whose output goes straight into a compiler. What matters is that the model is reachable and that a failed call is retried, because a formalization pipeline that drops a request looks identical to a formalization pipeline that failed — and only one of those is a fact about the model. Automatic failover across providers is the specific feature that keeps those two apart. OrcaRouter carries neither MathForm-8B nor AREX-2 today, so the OpenBMB weights come from OpenBMB's own distribution and any hosted endpoint is somebody else's; what we offer is the routing in front, with provider list price passed through and nothing added per token.

For an agent like AREX-Base the shape is harder and the case for routing is stronger. A deep-research trajectory makes dozens of model calls per query, each one re-reading a context that grows as evidence accumulates, and the failure modes compound: a provider that times out on step nineteen of twenty-five does not degrade the answer, it produces a confident wrong one. That is the argument for putting the loop behind one API for 200-plus models with a failover rule that decides at runtime, so a single bad provider is a retry rather than a bad citation.

The one thing to watch, and the one thing not to assume

Three facts would answer most of the open questions about AREX-2, and all three are visible from outside: whether files appear in that repository, what parameter count the card claims, and whether it carries the Apache 2.0 tag its two predecessors did. Until they do, the only defensible statement about AREX-2 is that BAAI reserved the name on 29 September 2026 and has announced nothing.

The thing not to assume is that a second generation inherits the first one's shape. A "2" is a product name, not an architecture. It could be bigger than the 122B Base, or it could be a small distillation of the same framework — and given that the family already shipped both a quality tier and a serving-cost tier at once, both readings are plausible. What is not plausible, at this moment, is publishing a specification for it.