
AREX-2 vs MathForm-8B: A Proof Checker and a Search Loop Are Different Kinds of Verification
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 507 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 194 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1143 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 55 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 106 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
MathForm-8B and AREX-2 are the two most rigorously verifiable models in this series, and they verify in opposite ways. MathForm-8B is OpenBMB's 8-billion-parameter autoformalizer, released in August 2026: give it a mathematics problem written in plain English and it produces a formally correct Lean 4 statement of that problem, which a proof assistant's compiler then checks. AREX-2 is a Hugging Face repository that BAAI created on September 29, 2026 at 17:56 UTC, holding a .gitattributes and no weights, no model card, no licence tag and no announcement — a reserved name for a probable second generation of BAAI's AREX deep-research agents, whose first generation verifies itself by re-checking its own answer against the constraints it was given.
So the interesting question here is not which model is better. It is what it means for a model to be checkable, because one of these two is checked by a compiler that does not care what anyone believes, and the other is checked, if the family precedent holds, by a loop that the same model is running. That distinction survives even though one side of the matchup has not shipped, and it is the reason these two belong in the same conversation at all.
The side that exists, and the reason it is checkable
MathForm-8B has a narrow job description, and it is worth stating precisely because the narrowness is the point. It does not solve mathematical problems and does not claim to. It was fine-tuned on FormalVerse, a corpus of roughly 367,000 verified Lean 4 examples, and its training reward came from the Lean compiler rather than from a human preference model. Given an informal statement, it emits the imports, types and theorem header — leaving the proof obligation itself open — so that a compiler can confirm the formalization says what the informal problem said.
Its published figures, all vendor-reported by OpenBMB and none independently reproduced, are an average Pass@8 of 88.06% under a syntax check and 72.37% under a stricter consistency check across six benchmarks, dropping to 63% and 37% on the hardest FATE-H and FATE-X subsets. Read those as what they are: a model measured against a formal system, where a wrong answer is a failure to compile rather than a disagreement about quality. That is a rare property. Most benchmark claims on this blog rest on a grader that has to be trusted; this one rests on arithmetic that a machine performs.
The side that does not exist yet, and why its verification is different
AREX-2's only confirmable facts are the org, the name and the timestamp. Nothing has been uploaded and BAAI has said nothing. The family behind the name does exist, though, and it shipped on 23 July 2026 as AREX-Base — a 122-billion-parameter mixture-of-experts with 10 billion active parameters on a Qwen3.5-122B-A10B base — together with AREX-Turbo, a dense 4B. Both Apache 2.0, both agents rather than chat models.
The AREX design is a two-loop research framework, and the outer loop is a verification step. An inner loop gathers evidence from search and browsing, integrates it, and produces a candidate answer with an attached confidence figure. The outer loop then checks that candidate against the original constraints and decides: accept it, refine it, or discard the trajectory and restart. Between iterations the model maintains a context block of verified findings, open candidates, unresolved constraints and the next plan, which is what lets it follow a long line of enquiry without losing the thread. BAAI reports 82.5 on BrowseComp, 85.4 on GAIA and 89.9 on DeepSearch QA for the Base, and 70.7 / 81.6 / 78.5 for the Turbo — the vendor's own numbers, unreproduced.
That is a real and useful form of self-checking. It is also categorically weaker than a compiler. When the AREX outer loop decides an answer satisfies its constraints, the judgement comes from a language model reading the constraints. When Lean accepts a MathForm-8B formalization, the judgement comes from a type checker implementing a fixed calculus. Neither one is free of error — a formalization can be valid Lean that captures the wrong theorem — but only one of them can be wrong without anybody noticing, and the difference is not a matter of degree.
• Availability — MathForm-8B is downloadable from OpenBMB with a published card. AREX-2 is a repository with no files in it.
• Parameters — 8B dense for MathForm-8B. Unknown for AREX-2; the family spans a 122B mixture-of-experts and a dense 4B.
• What it produces — Lean 4 theorem statements for MathForm-8B, with the proof left open by design. Unknown for AREX-2; the first generation produced a researched answer with a confidence figure.
• How it is checked — the Lean compiler, for MathForm-8B. A verification loop inside the same model, on the evidence of AREX-Base.
• Licence — OpenBMB's own terms for MathForm-8B; nothing declared yet for AREX-2. The first AREX generation was Apache 2.0.
• Headline figures — 88.06% syntax-checked and 72.37% consistency-checked Pass@8 for MathForm-8B, vendor-reported. None exist for AREX-2.


Two jobs that a stack can hold at once
These models do not overlap in function, and it is worth saying so before anyone reads a "versus" as a choice. MathForm-8B is a front-end for a proof assistant. Its output is a theorem statement that a mathematician or an automated prover then works on. Its natural place is inside a verification pipeline for mathematics, formal methods, or specification work, where the value is that a downstream tool can consume the output without trusting the model.
AREX-2, if it continues its family, is a back-end for questions that have no compiler. "Which of these three filings is inconsistent with the other two" has no formal system to check it against, and the only available defence is gathering more evidence and re-examining your own reasoning — which is what the AREX outer loop is. A team that needs both would run them in different places: formalization for the part of the problem that can be made formal, and a search-and-verify agent for the part that cannot.
That split is also the honest answer to which of these you can act on this week. MathForm-8B is shipped, documented and usable now. AREX-2 is a name.
What the routing story looks like when verification is the subject
Because the two are used so differently, the infrastructure question splits as well.
For MathForm-8B the workload is a burst of short, deterministic generations whose output goes straight into a compiler. What matters is that the model is reachable and that a failed call is retried, because a formalization pipeline that drops a request looks identical to a formalization pipeline that failed — and only one of those is a fact about the model. Automatic failover across providers is the specific feature that keeps those two apart. OrcaRouter carries neither MathForm-8B nor AREX-2 today, so the OpenBMB weights come from OpenBMB's own distribution and any hosted endpoint is somebody else's; what we offer is the routing in front, with provider list price passed through and nothing added per token.
For an agent like AREX-Base the shape is harder and the case for routing is stronger. A deep-research trajectory makes dozens of model calls per query, each one re-reading a context that grows as evidence accumulates, and the failure modes compound: a provider that times out on step nineteen of twenty-five does not degrade the answer, it produces a confident wrong one. That is the argument for putting the loop behind one API for 200-plus models with a failover rule that decides at runtime, so a single bad provider is a retry rather than a bad citation.
The one thing to watch, and the one thing not to assume
Three facts would answer most of the open questions about AREX-2, and all three are visible from outside: whether files appear in that repository, what parameter count the card claims, and whether it carries the Apache 2.0 tag its two predecessors did. Until they do, the only defensible statement about AREX-2 is that BAAI reserved the name on 29 September 2026 and has announced nothing.
The thing not to assume is that a second generation inherits the first one's shape. A "2" is a product name, not an architecture. It could be bigger than the 122B Base, or it could be a small distillation of the same framework — and given that the family already shipped both a quality tier and a serving-cost tier at once, both readings are plausible. What is not plausible, at this moment, is publishing a specification for it.
