
AstaBrief 8B: Ai2's Open Report Writer Runs 3.5x Faster Than the Pipeline It Replaces
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 216 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 111 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 42 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 106 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 214 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The headline number from Ai2's AstaBrief 8B announcement is a stopwatch reading, not a benchmark score: across the full Asta pipeline, the new model generates a cited research report in an average of 51.1 seconds, against 178.5 seconds for the Claude-powered path it now sits beside. That is about 3.5x faster, and it comes from a single architectural decision — write the whole report in one pass instead of summarizing, clustering, and then writing section by section. AstaBrief 8B is an 8B open-weights model, Apache-2.0 licensed, fine-tuned from Qwen3-8B, that takes a research question plus retrieved literature excerpts and returns a report with inline citations. It shipped into Ai2's Asta platform as "Fast mode" on 2026-10-02, and the weights, the training data, and an example local workflow went public the same day.
The repository is older than the announcement, and that matters
One detail in this release is worth stating plainly, because it changes how you should read everything else: the Hugging Face repository at allenai/AstaBrief_8B carries an internal creation timestamp of 2026-02-09, and the companion DPO dataset dates to November 2025. The October 2 announcement is real — the blog post, the AI2 tweet, and the model card all landed that day — but the repository history is longer than the news cycle suggests.
What happened on October 2 is that Ai2 shipped and documented the thing, and touched the repo to match: three commits landed on the model card between 16:18:06 and 16:18:20 UTC on release day, adding a response template to the chat template and dropping a no-op token. That is housekeeping on a card, not a retraining. Ai2 is also explicit in the blog that "most of the training and evaluation described was completed in 2025," and that the proprietary models used to generate training data and as comparison points "reflect the frontier at the time." The lab says it has not rerun the full evaluation against today's frontier models.

So this is a GA release of work that was largely finished in 2025 — a launch, with a launch's documentation and a launch's platform integration, on top of a checkpoint that has existed in the lab's account for months. Nothing in that is dishonest, and the model is new to everyone outside Ai2. But it does mean the comparison numbers below describe a 2025 frontier, and Ai2 says so itself.
What AstaBrief 8B actually does
Ai2's Asta platform has a report-generation feature where researchers bring a substantial question and a body of literature, and get back a cited synthesis they treat as a working artifact. That feature previously ran entirely on what Ai2 calls Thinking mode: a multi-step pipeline that summarizes retrieved snippets, clusters them thematically, and writes the answer section by section, backed by Claude models. Thinking mode is thorough and slow.
AstaBrief 8B is the Fast mode. Given the same question and the same retrieved excerpts, it writes the final report in one forward pass. Ai2's claim is the interesting part: bypassing snippet summarization, clustering, and the section-by-section loop did not cost report quality, at least on the measures the lab tracked. The model is not a search engine and it does not retrieve — it is the writer at the end of a retrieval pipeline, and it expects to be handed the evidence.
That makes it a component rather than a product. It is also why Ai2 shipped an example workflow alongside the weights: a small library that turns your own PDFs into reports, in the ai2-scholarqa-lib repository, gives you a starting point for local generation.
The training recipe is deliberately boring, and that is the point
Ai2's own prior work, DR Tulu, showed that reinforcement learning can improve long-form report generation in open-weights models. For AstaBrief the team considered that path and chose not to take it, on the grounds that RL is unstable and expensive and they wanted something easier to debug. What they built instead is supervised fine-tuning followed by offline direct preference optimization. Two stages, both conventional.
The SFT stage started from real queries submitted through Ai2's Asta system rather than synthetic prompts. From the collected logs the team filtered for quality, relevance, and privacy — stripping bot and beta-tester traffic, dropping queries too short to be meaningful, and running an LLM pass to catch non-English queries, non-scientific requests, and prompts containing personal information. That left a pool of 90,000 research-focused queries. Full-report targets for those queries were generated with the multi-step ScholarQA pipeline, using Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1 as backing models. After quality filtering: 47,000 usable training examples.
The DPO stage needed something different — pairs of reports with a preference between them. One report per query came from the ScholarQA pipeline; the competing report came from feeding ScholarQA's retrieved excerpts to a different model, drawn from o3, o4-mini, DeepSeek-V3, or DeepSeek-R1. Two judges, GPT-4.1 and DeepSeek-R1, picked a winner for each pair, and only pairs where both judges agreed survived. Ai2 reports 95% agreement between the LLM judges and human preferences. The final preference set came to roughly 6,000 examples. Both the SFT mix and the DPO mix are on Hugging Face for inspection.

The most transferable finding is also the least glamorous. Early SFT runs wrote better reports but lagged on answer precision and citation quality, so the team tested four statistics-based filters on the synthetic training data: output-to-input token ratio, average retrieval relevance of cited papers, citation density (the share of statements carrying at least one citation), and citation diversity. The strongest gains came from filtering out synthetic reports with low citation density — and more aggressive filtering, filter combinations, and learning-rate sweeps did not add meaningful gains. Ai2's conclusion is worth carrying beyond this model: specialization was not a matter of pouring more scientific text into pretraining, but of the composition and quality of post-training data.
The scoreboard Ai2 published, and how to read it

Every figure below is vendor-reported. It comes from Ai2's model card and blog, produced by the lab's own evaluation harness, and none of it has been independently reproduced. The primary target was SQABench-CS2, a set of 100 user-written computer science research questions, tracked on four metrics: ingredient recall (coverage of necessary content), answer precision (whether each paragraph is relevant), citation precision (whether each citation supports its claim), and citation recall (whether claims are fully supported by the citations given).
• Overall average — AstaBrief 8B 87.0, its own SFT checkpoint 83.7, base Qwen3-8B 77.3
• Ingredient recall — AstaBrief 8B 90.2, SFT 85.2, Qwen3-8B 77.8
• Answer precision — AstaBrief 8B 89.0, SFT 90.4, Qwen3-8B 90.6
• Citation precision — AstaBrief 8B 90.5, SFT 87.7, Qwen3-8B 76.2
• Citation recall — AstaBrief 8B 78.2, SFT 71.3, Qwen3-8B 64.6
Read the rows together and the DPO stage looks targeted rather than uniformly better. Overall average rises, ingredient recall and both citation metrics rise sharply, and answer precision slips fractionally below both the SFT checkpoint and the base model. If your use case is relevance-per-paragraph above all else, the fine-tune is not automatically your answer.
On secondary evaluations Ai2 tested the final model against its own ScholarQA pipeline and against DR-Tulu-8B. On SQABench-CS2 dev, Asta ScholarQA scored 87.6, DR-Tulu-8B 86.5, and AstaBrief 8B 86.3; on the test split, ScholarQA 86.2, DR-Tulu-8B 88.8, AstaBrief 87.0. On DeepScholarBench, a 63-query long-form synthesis benchmark, ScholarQA scored 60.25, DR-Tulu-8B 56.26, and AstaBrief 53.50 — the one row where the open report writer trails both alternatives. In two LLM-judged pairwise comparisons against the Claude-powered pipeline, AstaBrief took 55% of dev comparisons and 72% of test comparisons. A separate human study covered only 14 questions from three researchers, and on overall preference DR-Tulu won, though two of the three researchers preferred AstaBrief on citation accuracy. Fourteen questions from three people is a signal, not a verdict, and Ai2 presents it that way.
The lab also published early usage from the Asta integration: among 374 users who tried Fast mode, 29.1% used it on two or more days, users generated 3.67 report threads on average, 23% never switched back to Thinking mode, and 18% moved between the two modes depending on the task. Positive feedback arrived at 84.2% for Fast mode against 85.2% for Thinking mode — close enough that Ai2 characterizes the feedback as too sparse to draw strong conclusions from. That is the honest framing, so keep it.
Running it yourself
AstaBrief 8B is Apache-2.0 and based on Qwen/Qwen3-8B, so it lands in the single-GPU class rather than the datacenter class: a dense 8B model on the Qwen3 architecture, 36 layers, hidden size 4096, 32 attention heads with 8 key-value heads, a 151,936-token vocabulary, and the base's 40,960-token position ceiling. In BF16 it fits comfortably on a 24GB card, and the card's own example uses vLLM with temperature 0.7, top-p 0.95, and a 4096-token output cap.
One operational warning is printed in bold on the model card and is easy to skip past: the checkpoint was fine-tuned against one specific prompt format, published as sft_prompt.txt in the AstaBrief_prompts dataset. Use a different prompt or interaction format and Ai2 expects degraded or inconsistent behavior. If you evaluate this model with your own harness and get poor results, check the prompt before you conclude anything about the weights.
The card is also candid about scope. AstaBrief is intended for research and educational use, not as a general-purpose assistant, and there is no hosted inference endpoint from Ai2 — you download it, or you use it inside Asta. No provider in our catalogue serves it, including us, so there is no route to point at; the honest way to use it is on your own hardware or behind your own firewall, which is precisely the case Ai2 makes for open weights in the first place.
Where a router fits into a report pipeline
AstaBrief occupies one slot in a longer chain. Something retrieves the literature, something decides which excerpts are worth passing down, and something may judge or verify the finished report. Those calls are ordinary hosted-model calls, and they are the part of the stack where you actually want a choice of vendors: one API covering 200-plus models, provider list price passed through at 0% markup so a vendor's price cut shows up on your bill the same day, automatic failover if a provider degrades, and a routing DSL if you want to compose several models into a single call. Self-host the writer because that is the part with data-sensitivity implications and a fixed cost; route the orchestration because that is the part where flexibility is worth more than ownership.
There is also a cheap experiment here. Ai2's own comparison set is Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1 — deliberately a 2025 frontier, and the lab says it has not rerun it. Putting AstaBrief's output next to today's hosted models on your own questions is a one-key exercise if both arms are reachable through the same endpoint, and it answers the question the blog post leaves open.
What would change the picture
Three things would turn this from a well-documented launch into a settled result. An independent reproduction of the SQABench-CS2 numbers, since every figure above is self-reported. A rerun against current frontier models, since the comparison set is a year stale by the lab's own admission. And a larger human evaluation than fourteen questions from three researchers — Ai2's own blog closes by arguing that the field needs evaluations going beyond citation support to ask whether a model preserves the evidentiary scope of its sources, including whether it quietly turns sample-specific findings into broad generalizations. That is a fair critique of the whole evaluation category, and it applies to this release too.
For now, the practical read is simple. If you generate cited reports from literature and you want the writer on your own hardware, at an Apache-2.0 license, with the training data open, AstaBrief 8B is the most directly purpose-built open option anyone has published, and the 3.5x wall-clock reduction is the reason it exists. If your work depends on the sharpest possible synthesis, note that Ai2's own table puts DR-Tulu ahead on overall human preference and ScholarQA ahead on DeepScholarBench — and that Thinking mode is still there, in the same product, for exactly that reason.
