
TwIL-LM3-Pro: A 3.66B Formal-Logic Model That Ships by Email Request
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 118 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 213 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
TwIL-LM3-Pro is a 3.66-billion-parameter reasoning model from webAI, built specifically for formal logic rather than for general conversation, and it has been on Hugging Face since 3 September 2026. It is not a new release, and the framing circulating around it this week — that webAI "has released a 3.66B model" — is a month late to the weights. What actually changed in the last few days is smaller and more useful: the checkpoint was revised on 30 September, and on 1 October webAI published a LiteRT-LM build of the model for on-device Android and iOS inference. So the interesting question is not what TwIL-LM3-Pro is, but whether you can get it, and the answer depends on which of three distribution channels you are standing in — because one of them is a form.
That distinction matters more than it normally would, because TwIL-LM3-Pro is not a product you can simply call. It is a checkpoint you download, run on your own hardware, and accept the terms of a licence that forbids commercial use. The comparison its own model card invites — against Qwen3-8B, a model with more than twice the parameters — is real, but it is a comparison of vendor-run benchmarks on a vendor-built harness, and the model card says so itself in more places than most. This piece reads that card closely and separates what webAI measured from what anyone else has.
What the model actually is
TwIL-LM3-Pro is a dense Granite decoder — `GraniteForCausalLM`, 40 layers, hidden size 2,560, 40 attention heads against 8 key-value heads — derived from IBM's `granite-4.2-3b`. The 3.66B parameter count is the real total, 3,659,737,600, and the context window is the base model's 131,072 tokens carried through unchanged. It emits a `<think>…</think>` block before answering and was evaluated with greedy decoding at a 2,048-token budget.
The post-training stack is the part webAI actually wants you to look at, and it is unusually well documented for a model this size: a rank-64 LoRA supervised fine-tune on a synthetic formal-logic corpus, multipath distillation to produce several reasoning traces per prompt, parameter-space checkpoint fusion rather than taking the final checkpoint, a WiSE-FT interpolation at α = 0.15 merged back toward the pretrained base with TIES and DARE-SLERP, and finally an entropy-weighted GRPO stage — webAI calls it MGPO — against a programmatic verifier, run through step 2580, which is the checkpoint that shipped.
The published artifact is the merged policy rather than a LoRA adapter, which is the practical detail: you can run it as a standalone model, in bf16 at 6.82 GiB, or as a Q4_K_M GGUF at 2.09 GiB on a CPU or 4 GB of VRAM. That is the "runs on a laptop" claim, and it is the one claim in the whole card that is trivially verifiable and therefore worth trusting.
The Qwen3-8B comparison, read carefully
The headline webAI puts on the model is that TwIL-LM3-Pro reaches the top of its own Track A summary rows at 3.66B parameters. The four summary rows are a macro gate of 0.5539, strict-7 of 0.2879, a six-lane average of 0.5389, and macro_primary of 0.5875. Against Qwen3-8B, the macro gate is 0.5539 versus 0.5336 — and the model card itself writes the sentence that a marketing page would have deleted: "the gate lead over Qwen3-8B is 0.020 — smaller than the sampling noise at n = 200 per lane — so read that one as parity, not a win."
That is the single most important line on the page. webAI's own framing of the headline result is that it is a tie. Where the comparison is not a tie:
• Strict-7 — TwIL-LM3-Pro 0.2879 vs Qwen3-8B 0.2093, and this row uses no loose-match credit anywhere, so it cannot be produced by formatting luck.
• Strict MCQ — TwIL-LM3-Pro 0.4100 vs Qwen3-8B 0.0000, the widest lane gap of the eleven rows reported.
• Rule induction — TwIL-LM3-Pro 0.4195 vs Qwen3-8B 0.3680.
• Corpus fit — the lowest `lm_corpus` perplexity of any arm at 2.3130, with Qwen3-8B at 2.5478.
• Parameters — 3.66B vs roughly 8B, which is the entire basis for the "less than half the parameters" claim.
Two caveats ride on that table and webAI states both. Track A generations from TwIL-LM3-Pro average 1,902 tokens and 24.2% of them hit the length cap, against 564 tokens for its own TwIL-LM3 predecessor; the card's word for the resulting gap is "indicative rather than exact." And the throughput figures in the table were measured in a later session on a newer vLLM (0.19.1) than the comparison columns (0.11.2), so the token-per-second rows are comparable with each other and only approximately with the rest.
Where it does not win
On the held-out suite the model card is blunt: "TwIL-LM3-Pro does not lead it." The 10-dataset macro is 0.7901, statistically level with its own base model at 0.7942, and well behind Qwen3-8B at 0.8493 and gpt-oss-120b at 0.8689. Its 14-dataset macro of 0.7425 beats its base by 0.0093, and that entire gain comes from BBH-logic (0.9540 against the base's 0.9013) — strip that one row out and the model averages 0.7263 across the remaining thirteen, below VibeThinker-3B's 0.7350.
There are three genuinely weak spots inside the specialisation, all disclosed: FOL translation exact match at 0.0100, `procedural` at 0.1200 strict, and semantic-parse exact match at 0.0000. Rule induction parses only 56.5% of its outputs. And the model is verbose by construction — about 792 tokens per Track B answer against 482 for its predecessor — which is a cost, not a capability.
None of this is a criticism of the release. It is what a well-written model card looks like, and it is why the numbers here can be quoted at all.
Three ways to get it, and one of them is a form
The distribution situation is the actual news of the week and it is worth stating precisely.
The weights are on Hugging Face, ungated, under the webAI Non-Commercial License ver. 1.0 — which permits research and personal use and requires a separate agreement with webAI for revenue-generating deployment. That licence is the binding constraint on the whole release, and it is why TwIL-LM3-Pro is not a model most product teams can simply adopt.
webAI says the family has passed half a million downloads in a month, a figure the company published in its own announcement and that has not been independently audited. Downloads of a free non-commercial checkpoint are not a proxy for production use, and webAI does not present them as one.
The vendor's own platform, meanwhile, is not a public endpoint. webAI describes itself as an enterprise platform that brings AI to your data, with a local-first model application, and its commercial route is an enterprise arrangement rather than a published per-token rate card. TwIL-LM3-Pro is also absent from the large third-party inference platforms that index open checkpoints — we checked, and it is not in the OrcaRouter catalogue either — so the practical answer to "where do I call this from an API" is that right now you mostly do not; you download it.
The third channel is the new one. The `TwIL-LM3-Pro-LiteRT-LM` repository appeared on 1 October 2026, carrying `.litertlm` artifacts and tagged for Android and iOS — webAI's own LiteRT-LM runtime packaging of the same weights. Whether that build inherits the non-commercial licence is the question to answer before planning around it, because the parent repository does.
Why a formal-logic specialist at this size is worth watching
The reason this release keeps resurfacing is not the benchmark table. It is that webAI published the whole pipeline, ran the identical recipe on five different base models, and reported what happened. The family comparison table records Track A macro-gate movement of +0.012 to +0.145 depending on the base, with the held-out suite staying inside ±0.013 — the paper-trail version of "this generalises." The one coefficient chosen per model is the merge weight, swept on held-out performance: 0.15 for SmolLM3, 0.20 for Granite and the TwIL base, 0.50 for VibeThinker-3B.
That is a reusable recipe for turning a small general model into a narrow specialist without destroying what it already knew, published with the negative results attached. For anyone who needs entailment labels, Lean statement formalisation, or semantic parses rather than fluent prose, a 2.09 GiB GGUF that runs offline with no per-call fee and no network dependency is a different proposition from any hosted model, whatever the Elo table says.
The open question is whether the weights ever appear under a licence that permits commercial use, and whether the LiteRT-LM port ships with the same restriction. Until then TwIL-LM3-Pro is a research artifact with an unusually honest model card — which is not nothing.

If you need this capability in production today
For a commercial deployment, the licence rather than the benchmarks is the blocker, and the practical substitute is a hosted model you can route to. That is the layer OrcaRouter operates: one OpenAI-compatible endpoint in front of 200-plus models at provider list price with no markup added, so a vendor price cut is live the same day rather than after a contract cycle. For the handful of places a small formal-logic checkpoint genuinely fits — offline batch labelling, an air-gapped entailment pass, a preprocessing step whose output is a label rather than prose — the honest split is to run the local model where the workload is text-to-label, and route the surrounding reasoning, prompt expansion and QA to hosted models on the same key.
One caveat worth repeating in this section specifically: with automatic failover you can also point at a model before you trust it on a production path, and roll back by editing a routing rule rather than redeploying. That is the pattern that fits a model like this when its commercial status is unresolved.

What to watch
Three things would change this story. A licence change, or a clearly licensed LiteRT-LM port, would move TwIL-LM3-Pro from research artifact to deployable component. Appearance on a major hosted inference platform would give it a real API. And an independent evaluation — anyone other than webAI running the Track A harness — would tell us whether the strict-7 lead over Qwen3-8B survives being measured by someone who did not build the harness. None of the three has happened yet, and the model card is honest enough that you can see exactly what would need to hold for them to matter.

