A generated title card headed 'TwIL-LM3-Pro vs Qwen 3.8', subtitled 'The Comparison the Card Actually Ran', with two rounded cards reading 'TwIL-LM3-Pro - 3.66B, local, non-commercial' and 'Qwen3.8-Max - hosted, 1M context, $2 / $6'. A footer line reads 'webAI measured against Qwen3-8B, not Qwen3.8-Max.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

TwIL-LM3-Pro vs Qwen 3.8: A Mismatch, and the Comparison That Actually Matters

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

TwIL-LM3-Pro and Qwen3.8-Max do not belong on the same page, and the reason is not that one is better. It is that the published case for webAI's 3.66B formal-logic model is built on a comparison against a Qwe​n model — and the Qwe​n model in question is not the one the headline names. webAI's Track A and Track B tables measure TwIL-LM3-Pro against Qwen3-8B, a dense 8B open-weights model from an earlier generation, and pay Qwen3.8-Max no attention at all: not because TwIL-LM3-Pro would win, but because Qwen3.8-Max is Aliba​ba's flagship hosted system, a sparse mixture-of-experts model reported at 2.4 trillion parameters with roughly 95 billion active per token, a one-million-token multimodal context window, and a rate card of $2.00 per million input tokens and $6.00 per million output. Putting a 2.09 GiB laptop GGUF next to that is a category error dressed up as a versus page.

So this piece does two things. It corrects the comparison the search results get wrong, and then it answers the version of the question a reader probably actually has: given that a frontier model is one API call away and a formal-logic specialist runs on your own machine, when does each one earn its place?

The model the card actually measured

The relevant comparison is against Qwen3-8B, and webAI's own summary of it is more modest than the secondhand write-ups suggest. TwIL-LM3-Pro's in-domain macro gate is 0.5539 against Qwen3-8B's 0.5336, and the model card contains the sentence that a marketing page would have removed: the gate lead is 0.020, "smaller than the sampling noise at n = 200 per lane — so read that one as parity, not a win."

Where the comparison is not a coin flip: strict-7, 0.2879 against 0.2093, a row that uses no loose-match credit anywhere and therefore cannot be manufactured by formatting. Strict multiple-choice, 0.4100 against 0.0000. Rule induction, 0.4195 against 0.3680. And the parameter ratio — 3.66B against roughly 8B — which is the whole "less than half the size" claim.

On the held-out suite webAI is blunter still: "TwIL-LM3-Pro does not lead it." The ten-dataset macro is 0.7901, level with its own Granite base and behind Qwen3-8B's 0.8493. A small specialist that ties a general model twice its size on formal logic and loses to it on general capability is the expected shape, and reporting it that way is what makes the strict-7 figure credible.

What Qwen3.8-Max actually is

Qwen3.8-Max is not an iteration of Qwen3-8B. It is Aliba​ba's premium tier, and on the OrcaRouter catalogue page it appears as `qwen/qwen3.8-max` with a release date of 3 August 2026, a one-million-token context window, text, image and video input, text output, and full support for thinking mode, function calling, built-in tools and structured outputs. It is served across OpenAI-compatible, Anthropic-compatible and native dashboards in five regions, and thinking is toggled with the `enable_thinking` parameter. The rate is $2.00 per million input tokens and $6.00 per million output.

A dated snapshot, `qwen/qwen3.8-max-0902`, was published on 2 September 2026 with the same capabilities and the same pricing, published specifically so behaviour can be pinned for reproducible runs. If you are building against Qwen3.8-Max for anything long-lived, that snapshot identifier is the one to put in the config rather than the moving alias.

Note what is absent from that description: a parameter count you can download, a licence file, and any path to running it yourself. Qwen3.8-Max is a hosted product. Journalistic coverage at launch reported open weights promised for a following week, and techsy's framing — "2.4T params, 1M context, no weights yet" — is the state a reader should verify rather than assume, because a promise reported at launch is not a published checkpoint.

Four dimensions, and none of them is close

• Parameters — TwIL-LM3-Pro 3.66B dense, all active vs Qwen3.8-Max reported at 2.4T total in a sparse mixture-of-experts design with roughly 95B active per token. Three orders of magnitude on the total, two on the active.

• Where it runs — TwIL-LM3-Pro downloads as bf16 at 6.82 GiB or a 2.09 GiB Q4_K_M GGUF for CPU or 4 GB of VRAM vs Qwen3.8-Max, which exists only as a hosted endpoint.

• Context — TwIL-LM3-Pro 131,072 tokens, inherited from its Granite base and never validated beyond 8,192 in its own evaluation vs Qwen3.8-Max at 1,000,000 tokens.

• Modalities — text in, text out vs text, image and video in, text out.

• Licence — webAI Non-Commercial License ver. 1.0, with a separate agreement required for revenue-generating use vs a hosted commercial API with no weights to license.

• Cost — TwIL-LM3-Pro: no per-token fee and no hosted endpoint, against Qwen3.8-Max's $2.00 per million input tokens and $6.00 per million output, with a cached-input rate around $0.25 per million.

A 3.66B model is cheap because it is small, and it is small because it does one thing. Qwen3.8-Max is expensive because it is general and it is hosted. Neither rate is a verdict.

The question behind the question

Strip the naming confusion away and an engineering question remains, and it is a good one: if a frontier hosted model can reason about formal logic, why run a narrow specialist at all?

Three reasons hold up. The first is the licence and the network: a non-commercial research group, an air-gapped environment, or a batch labelling pass that has to run on a laptop with no connectivity cannot call a hosted endpoint, and a 2.09 GiB GGUF answers that constraint directly. The second is cost structure over volume — a formal-logic labelling job over a million short inputs pays per token on a hosted frontier model and pays nothing but wall-clock on a local one. The third is the output shape: TwIL-LM3-Pro was tuned to emit machine-checkable structures rather than prose, and for entailment labels and semantic parses that is a smaller pipeline than prompting a general model and parsing its answer.

Against that, Qwen3.8-Max's case is simple. It sees images and video, it holds a million tokens, it handles tools and structured output through a documented API, and it has published benchmarks and independent evaluation behind it. Nothing about a narrow specialist threatens that; the two are answering different constraint sets.

The stack that uses both

The pattern that survives contact with a real pipeline is two-tier, and it is not exotic. A frontier hosted model reads the messy input — a scanned textbook page, a mixed-modality source, a long specification — and turns it into clean structured text. That text goes to a local formal-logic model whose entire job is to emit a label, a parse or a Lean critique on hardware you own. The verdict on whether that second tier is worth its maintenance cost is the strict-7-style comparison, not the gate, because a specialist earns its place on the lanes where the metric has no room for charitable scoring.

There is also a cue worth taking from webAI's own card: the pipeline was run on five different base models, with only the merge coefficient changing per model, and webAI reports the outcome for each including the ones that moved least. That is a stronger claim about a recipe than any single benchmark line, and it is the kind of evidence that tells you a method generalises rather than that one checkpoint got lucky.

What is routable here, and what is not

This is the one place the two models sit in the same sentence honestly. Qwen3.8-Max is a route: it is served through OrcaRouter as `qwen/qwen3.8-max`, and because the platform passes provider list price through with 0% markup added, the $2.00/$6.00 rate is the vendor's own and a vendor price cut is live the same day rather than after a contract cycle. The dated `qwen/qwen3.8-max-0902` snapshot is routable alongside it, so pinning a reproducible model id is a config choice rather than a separate vendor relationship.

TwIL-LM3-Pro is not a route, and it would be dishonest to imply otherwise. It is non-commercially licensed with no published hosted endpoint, and the current way to use it is to download the checkpoint and run it yourself. What the router does for a pipeline built around it is the surrounding text work — the prompt expansion, the corpus generation, the filtering pass — which runs on the same OpenAI-compatible endpoint against 200-plus models, so swapping the model that generates evaluation data for the model that grades it costs nothing but a config edit, with automatic failover to keep a long batch alive when a provider hiccups.

The most defensible use of the two together is exactly that: a routable frontier tier doing general reasoning and structured extraction, a downloaded specialist doing the formal-logic judgement, and a single key for everything in between.

Which model to compare, and when

If you are evaluating small formal-logic models, the Qwe​n worth putting beside TwIL-LM3-Pro is Qwen3-8B, and webAI has already published that comparison with the caveats attached. If you are choosing where to run a production workload, Qwen3.8-Max is a different decision entirely — a hosted frontier system with a rate card and no download — and the right question is not which is better but which constraint binds: connectivity and licence on one side, context, modality and scale on the other. Most real systems end up needing both answers.

A generated two-column scoreboard headed 'TwIL-LM3-Pro vs Qwen3.8-Max - the scoreboard' with six shared dimension rows. TwIL-LM3-Pro: Parameters 3.66B dense; Where it runs local download; Context 131,072 tokens; Input text only; Licence non-commercial; Cost no per-token fee. Qwen3.8-Max: Parameters 2.4T total, sparse MoE; Where it runs hosted endpoint; Context 1,000,000 tokens; Input text, image, video; Licence hosted commercial API; Cost $2.00 in / $6.00 out. A footer line reads 'Qwen3.8-Max specs per its OrcaRouter catalogue page; TwIL figures vendor-reported by webAI.' The OrcaRouter logo is composited in the bottom-right corner.A screenshot of the OrcaRouter model page for qwen/qwen3.8-max, showing the Qwen author line and release date 2026-08-03, the Vision, Tools, JSON and Reasoning capability badges, the vendor description positioning Qwen3.8-Max against GPT-5.5, Claude Opus 4.7 and Gemini 3.1 Pro, the /v1/chat/completions, /v1/messages and /v1/responses wire formats, and the $2.00 input and $6.00 output rates.A screenshot of the OrcaRouter model page for qwen/qwen3.8-max-0902, headed 'Qwen3.8 Max (0902)', showing the dated snapshot identifier, the Vision, Tools, JSON and Reasoning badges, text plus image and video input, and a 131K maximum output length field.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily