A hero title card for 'AstaBrief 8B vs Qwen3-8B: What an Ai2 Fine-Tune Adds to Its Own Base Model', naming the shared Qwen3 architecture and the 87.0 versus 77.3 SQABench-CS2 averages, with the OrcaRouter logo in the corner.
Guides & Insights

AstaBrief 8B vs Qwen3-8B: What an Ai2 Fine-Tune Adds to Its Own Base Model

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

There is no architecture story here, and that is the point. AstaBrief 8B is a fine-tune of Qwen/Qwen3-8B, and the two models share a configuration down to the last attention head: Qwen3ForCausalLM, 36 layers, hidden size 4096, 32 attention heads with 8 key-value heads, a 151,936-token vocabulary, and a 40,960-token position ceiling. Everything that separates Ai2's 2026-10-02 release from the April 2025 base is post-training. The interesting question is therefore not which is better in general — it is what a specific, well-funded post-training run buys you on top of a base model you can already download for free, and what it costs you in return.

Ai2's answer is a report writer that produces a cited multi-section synthesis in about 51 seconds, against 178.5 seconds for the Claude-powered pipeline it replaced inside the Asta platform — roughly 3.5x faster, with Ai2 claiming report quality held on the metrics it tracked. Qwen3-8B's answer is a generalist with hybrid thinking and non-thinking modes, a hundred-plus languages, and a year and a half of downstream use. Both are Apache-2.0. The trade is narrow capability plus speed against broad capability plus flexibility, and the data below makes the shape of it fairly concrete.

The architecture row is a no-op, so the prompt is not

Because the checkpoint geometry is identical, every serving intuition you have about Qwen3-8B transfers unchanged to AstaBrief 8B. It is a dense 8B in BF16, it fits on a 24GB card, it serves on vLLM or Transformers with no custom kernels, and it inherits the base's 40K position ceiling — neither model is a long-context model, and the 32K trained window extends with YaRN exactly as the base does. Deployment planning for the fine-tune is deployment planning for the base. That is worth saying plainly because it is not always true of fine-tunes, and it means the only thing you are evaluating is behavior.

Behavior is where the lock-in lives. AstaBrief's card carries a warning in bold: the model was fine-tuned against one specific prompt format, published as sft_prompt.txt in the AstaBrief_prompts dataset, and Ai2 says using a different prompt or interaction format may lead to degraded or inconsistent behavior. That prompt is not a casual chat template. Read it and you find a rigid contract — a bracketed query slot, a section_references block of quotes keyed by identifier, explicit citation syntax requiring the reference key to follow the sentence it supports, a designated marker for model knowledge that must not be combined with another source, and an instruction to open each section with a two-sentence TLDR. Qwen3-8B will accept anything you type. AstaBrief 8B is tuned to one shape of input and will underperform if you hand it another.

What 47,000 examples and 6,000 preferences bought

The fine-tune is entirely post-training, and the recipe is deliberately conventional: supervised fine-tuning followed by offline direct preference optimization, no reinforcement learning. Ai2 started from real queries submitted through its Asta system, filtered 90,000 research-focused prompts for quality, relevance, and privacy, generated report targets with the multi-step ScholarQA pipeline using Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1, and kept 47,000 examples. The preference stage then built roughly 6,000 pairs, with GPT-4.1 and DeepSeek-R1 judging and only agreed-upon pairs surviving.

Measured on SQABench-CS2 — 100 user-written computer science research questions — against the base model, here is what that bought, with every figure vendor-reported and none independently reproduced:

• Overall average — AstaBrief 8B 87.0 vs Qwen3-8B 77.3, a 9.7-point gain.

• Ingredient recall — 90.2 vs 77.8.

• Citation precision — 90.5 vs 76.2.

• Citation recall — 78.2 vs 64.6.

• Answer precision — 89.0 vs 90.6, a 1.6-point loss against the base.

The last row is the one that should shape expectations. Fine-tuning for report quality did not uniformly improve everything; it traded a small amount of per-paragraph relevance for large gains in coverage and citation grounding. If your task rewards relevant paragraphs above all else, the base model is not obviously the worse choice, and the fine-tune is not automatically your answer.

Two other things the money did not buy. AstaBrief 8B has no reported score on DeepScholarBench above Ai2's own alternatives — its 53.50 sits behind Asta ScholarQA at 60.25 and DR-Tulu-8B at 56.26 — and its human validation is fourteen questions from three researchers, on which DR-Tulu won overall preference. A specialist model can lose to a general-purpose research model on the specialist's own benchmark, and Ai2 published that.

A two-column scoreboard headed 'AstaBrief 8B vs Qwen3-8B — the scoreboard'. The AstaBrief column reads 'identity: fine-tune of Qwen3-8B', 'job: cited report writing', 'context: 40K inherited', 'prompt: one fixed format', 'licence: Apache 2.0', 'evidence: vendor-reported SQA-CS2'; the Qwen column reads 'identity: the base model', 'job: generalist, thinking modes', 'context: 40K, YaRN to 131K', 'prompt: open', 'licence: Apache 2.0', 'evidence: 11M downloads, established'. A footer reads 'AstaBrief SQA-CS2 average 87.0 vs base 77.3 per Ai2; unreproduced.'

What the base still does better

The comparison is not one-directional, and the generalist side of it is easy to undersell.

Qwen3-8B ships with hybrid thinking and non-thinking modes, so it can spend tokens on reasoning when a problem warrants it and answer directly when it does not. AstaBrief 8B was trained for one-shot report writing and inherits none of that deliberate reasoning behavior as a product feature. Qwen3-8B also carries the base's multilingual coverage — more than a hundred languages — while AstaBrief was trained on a corpus explicitly filtered to English scientific queries, so its competence outside that lane is unknown rather than merely untested.

Then there is the instruction surface. A generalist is what you want when the task will change next week, when you need tool calling, when the output schema is yours rather than Ai2's, or when you are prototyping and do not yet know what the final prompt looks like. And on the raw provenance question — the one that matters when someone asks why you trust the model — Qwen3-8B has had more than eleven million downloads and a year and a half of independent use. AstaBrief 8B has had a launch week.

What AstaBrief 8B has that the base cannot match is the citation discipline. Citation precision and recall are the two metrics where the fine-tune's lead is largest and most consistent with the training story — the strongest single filter in Ai2's data pipeline was citation density, the share of statements carrying at least one citation, and the resulting model reflects that priority. Getting a generalist to cite inline in a fixed key format, reliably, without dropping support for claims, is prompt engineering you write and maintain. Ai2's fine-tune makes that the model's default behavior.

A screenshot of the Hugging Face model card for Qwen/Qwen3-8B showing the TextGeneration tag, the apache-2.0 licence line, the qwen3 tag, the 8B parameter size in BF16 and a downloads-last-month figure of 11,020,586.

The Qwen line has moved on since April 2025

One more complication, and it is a practical one: Qwen3-8B is a year and a half old, and comparing a new fine-tune to it is not the same as comparing it to a viable incumbent.

Qwen's line now runs through Qwen3.5, Qwen3.6, Qwen3.7, and Qwen3.8, and the current generation is reachable without downloading anything. Qwen3.8 27B is a live route in our catalogue with a 262,144-token context window at $0.33 per million input tokens and $2.40 per million output; Qwen3.8 Flash carries a million-token window at $0.15 and $0.47; Qwen3 VL 8B Instruct covers the multimodal case at $0.18 and $0.70 per million with a 131,072-token window and has been routed since October 2025. Qwen3-8B itself is not a route we serve, and neither is AstaBrief 8B — both are download-and-run weights, and the honest framing is that the base belongs on your hardware and the modern Qwen generation does not have to.

That gives you a third arm for the evaluation Ai2 could not run. Put AstaBrief 8B on your own GPU, stand the Qwen3-8B base up beside it to confirm the reported deltas, and point the same harness at a hosted Qwen3.8 model for scale. All three arms are one endpoint away, which is what makes this a configuration exercise rather than a procurement project — one API across 200-plus models, provider list price passed through at 0% markup so a vendor price cut is live on your side the same day, automatic failover, and a routing DSL if you want several models answering one question together. The self-hosted arms stay self-hosted for the data-sensitivity reasons; everything hosted stays swappable, which matters when the next Qwen generation ships in a quarter.

How to decide

Take AstaBrief 8B if your product is cited literature synthesis and you want the citation behavior to be the model's default rather than your prompt's responsibility. The base-model geometry means zero serving surprises, the Apache-2.0 license and published SFT and DPO mixes make it auditable, and its lead on citation precision and recall is the largest and most explainable result in Ai2's tables.

Stay on Qwen3-8B, or move to a current Qwen3.x, if your tasks are varied, if you need thinking mode or tools or multilingual coverage, if you are still exploring the prompt, or if you would rather not be locked to a sixteen-thousand-character instruction block you did not write. The base lost on coverage and citation grounding, not on relevance, and it retained something the fine-tune gave up.

The thing to avoid is treating the 87.0-versus-77.3 average as the whole story. Ai2's own table shows the fine-tune giving up answer precision to gain citation discipline, its own lab notes that most of the training and evaluation was completed in 2025 and has not been re-run against the current frontier, and the base model it improves on is no longer the generation you would pick for anything except this comparison. What the fine-tune demonstrably adds is a shape of output; whether that shape is worth the narrowness depends entirely on whether it matches the shape of your product.