
TwIL-LM3-Pro vs MathForm 8B: Statement and Entailment Are Different Halves of Formalisation
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 118 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 213 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
TwIL-LM3-Pro and MathForm-8B are aimed at the same broad problem — getting a machine to produce formal, checkable text rather than plausible prose — and they are solving different halves of it. MathForm-8B is OpenBMB's 8-billion-parameter autoformalizer, released in August 2026: given a mathematics problem in plain language it emits the Lean 4 statement of that problem, leaving the proof obligation open for a compiler to check. TwIL-LM3-Pro is webAI's 3.66B formal-logic model from September 2026, and its target outputs are entailment labels, first-order-logic translations, semantic parses and Lean proof critiques. One produces the thing to be checked. The other tells you whether the thing you have entails what you claim.
Because the two overlap on exactly one lane — Lean formalisation — a comparison page is tempting and misleading. The more useful question is what each was trained to be rewarded for, because the reward signal is where the two designs diverge most sharply and where the smarter choice actually lives.
Statement versus entailment
MathForm-8B is trained on FormalVerse, a Lean 4 corpus the OpenBMB paper describes as roughly 367,000 verified examples, through supervised fine-tuning followed by reinforcement learning. The pipeline around it retrieves relevant definitions and existing formalisations from Mathlib before generation, then refines using compiler diagnostics and a semantic-consistency check. Its output is a formal statement — imports, types, theorem header — which an external compiler accepts or rejects. The model card is explicit that it does not solve the problem and does not claim to; the proof is somebody else's job.
TwIL-LM3-Pro comes at the same domain from the logic side. Its pipeline was aimed at six objectives: first-order-logic translation, entailment labelling, semantic parsing, Lean formalisation, Lean proof critique, and rule induction. Where MathForm is built to emit a statement, TwIL-LM3-Pro is built to make judgements about statements — is this entailed, is this parse correct, is this proof attempt sound. It also formalises, and that is the one place the two models are genuinely comparable rather than merely adjacent.
Reward from a compiler versus reward from a scorer
This is the design difference that matters, and both vendors document it.
MathForm's training reward came from Lean compilation and semantic-consistency feedback. A compiler is an oracle: it either accepts the formalisation or it does not, and there is no partial credit and nothing to argue with. That is the cleanest reward signal in this corner of machine learning, and it is why autoformalization benchmarks are reported as pass rates — the metric is downstream of a program that does not care what anyone believes.
TwIL-LM3-Pro's final stage was an entropy-weighted GRPO run against a programmatic verifier, with the model's own card describing partial credit for loose matches and token-F1 so that prompt groups where everything failed still produce gradient. Its headline macro gate is the equal-weight mean of four bounded classification lanes plus rule induction, and in that gate the multiple-choice and procedural lanes are credited as `max(exact_match, loose_match)` — a rule the card states plainly, because a correct answer formatted differently is a formatting artefact rather than a reasoning failure.
Neither design is worse. They are tuned for different consequences. A compiler-graded reward produces a model you can point at a hard formalisation problem and trust the output of, because the output is verifiable. A scorer-graded reward with partial credit produces a model that can be trained on fuzzier judgements — entailment, critique, rule induction — where no compiler exists to adjudicate and the alternative is not training on those lanes at all. The cost is that the resulting numbers are only as good as the harness, and webAI says so: its strict-7 row, which strips out every trace of loose-match credit, exists precisely so a reader can see the harsher figure alongside the headline one.
The scores are not on the same ruler
Both models publish Lean-adjacent numbers and they cannot be subtracted. MathForm's paper reports an average Pass@8 of 88.06% under Syntax Check and 72.37% under Consistency Check across six Lean benchmarks, with 63% and 37% consistency-check pass rates on the harder FATE-H and FATE-X subsets, and claims to outperform several specialised 32B autoformalizers. TwIL-LM3-Pro's card reports `lean_formalize` token-F1 of 0.5092 and `lean_critic` accuracy of 0.7950 on its own Track A harness.
Pass@8 under a compiler check and token-F1 against a reference string measure different things. Pass@8 gives the model eight attempts and asks whether one compiled and stayed consistent; token-F1 scores how closely a single generation matched a reference. A model can score well on one and poorly on the other, and nothing in either document licenses reading the two side by side.
The honest summary of the benchmark record is that MathForm-8B's evidence comes from a paper with a compiler in the loop and TwIL-LM3-Pro's comes from a vendor harness with a scorer in the loop, and no third party has run both through one rubric. Treat any page that puts "88.06%" next to "0.5092" as though they were a scoreline as a page that has not read the definitions.
Specifications, side by side
• Parameters — TwIL-LM3-Pro 3.66B dense vs MathForm-8B 8B, roughly 2.2 times the size.
• Base model — `ibm-granite/granite-4.2-3b` vs `Qwen/Qwen3-8B`.
• Training data — a synthetic formal-logic corpus covering six objectives vs FormalVerse, roughly 367,000 verified Lean 4 examples.
• Reward — a programmatic verifier with partial credit and loose-match tolerance vs Lean compilation and semantic-consistency feedback.
• Primary output — entailment labels, FOL translations, semantic parses, Lean statements and proof critiques vs a Lean 4 statement for a compiler to check.
• Context window — 131,072 tokens for TwIL-LM3-Pro, measured inside 8,192 tokens; MathForm-8B is served with generation budgets up to 16,384 tokens in its own usage example.
• Licence — webAI Non-Commercial License ver. 1.0 vs Apache 2.0.
• Quantised footprint — TwIL-LM3-Pro ships a 2.09 GiB Q4_K_M GGUF for CPU or 4 GB of VRAM; MathForm-8B publishes no GGUF in its own repository.
The licence decides the production question again
MathForm-8B is Apache 2.0. You can fine-tune it, redistribute it, and serve it inside a commercial product. TwIL-LM3-Pro is non-commercial: research and personal use are permitted, and revenue-generating deployment requires a separate agreement with webAI, whose commercial route is an enterprise arrangement rather than a published rate card.
For a research group or a student, that difference is irrelevant — both are ungated downloads on Hugging Face. For a company, it is decisive, and it points at MathForm-8B for any production autoformalization work regardless of which model scores better on a lane neither vendor measured the same way.
Where a router sits in this pipeline
Neither TwIL-LM3-Pro nor MathForm-8B is a route in the OrcaRouter catalogue, and the reasons differ: one is non-commercially licensed with no hosted endpoint, the other is published as an open checkpoint for self-hosting. The routing question in a formalisation pipeline is the layer around them. Building a FormalVerse-style corpus means generating thousands of informal problems, filtering them for well-formedness, and writing the semantic-consistency checks that grade the output — all text calls, and all the kind of work where you want to swap the model generating bulk data for the model judging it without a second contract. On one OpenAI-compatible endpoint in front of 200-plus models at provider list price with 0% markup added, that swap is a config change, and a vendor price cut on the generation model lands the same day. Automatic failover is load-bearing in a corpus build specifically because a job that dies two-thirds of the way through a generation pass costs the whole run.

The division of labour
If the job is turning textbook mathematics into compilable Lean statements at scale, and the result has to ship in a commercial product, MathForm-8B is the tool and the Apache 2.0 licence is why. If the job is deciding whether a body of text entails a claim, labelling entailment, parsing semantics, or critiquing a proof attempt, TwIL-LM3-Pro covers ground MathForm was never aimed at — and its non-commercial licence is a boundary on where that capability can be used, not a mark against the capability itself.
The combination that makes sense is a two-stage pipeline rather than a choice: one model emits the formal statement, the other judges it. Both have published the pipeline that produced them, which is unusual at this size and is the reason this comparison can be made at all.


