Open-sourcing AstaBrief, the fast report-generation model in Asta

2026年10月2日 15:19
站內 AI 整理稿

Back to Articles Open-sourcing AstaBrief, the fast report-generation model in Asta Enterprise Article Published October 2, 2026 Upvote 2 Kyle Wiggers Ai2Comms Follow allenai 🤗 Model | 📊 Data Language models can already help researchers search the literature, synthesize evidence, and work through complex questions.

But scientific work places particular demands on these models—answers need to stay grounded in evidence, the models need to preserve what the evidence actually supports rather than quietly broadening a study’s conclusions, and researchers need to be able to verify the final outputs.

We see that in how scientists use Asta, our agentic platform for scientific work.

Instead of simple keyword searches, users often bring substantial context and many constraints—for example, asking Asta to compare approaches across a body of literature while accounting for a particular method, population, or setting.

Many also return to generated reports later, treating them as working research artifacts rather than one-off answers.We wanted to help scientists generate cited reports faster, with a model they could download and run themselves.

To do that, we tested whether a small, open model trained specifically for scientific report generation could match the report quality of the proprietary models we were using, while reducing generation time and serving costs.

We built AstaBrief 8B, a model that turns a research question and retrieved literature excerpts into a cited report.

AstaBrief is available in Asta’s Generate a report feature today as Fast mode alongside Claude-powered Thinking mode, and we’re also open-sourcing it and the training data so others can study, reproduce, and build on our approach.

Developing AstaBrief required tens of thousands of real research queries, citation-focused filtering, preference data, and a redesigned report-generation pipeline that writes the full report in one pass rather than section by section.

The result is nearly an order-of-magnitude reduction in report generation time compared to the proprietary models we tracked—across the full Asta pipeline, Fast mode averages 51.1 seconds per report compared with 178.5 seconds for Thinking mode, about 3.5× faster.

Together, those efficiency gains made AstaBrief a useful test case for a broader goal: building open language models that can be adapted to the specific demands of scientific work.

Open weights will also let institutions run AstaBrief on their own infrastructure, which is necessary when research questions reveal sensitive or unpublished work.

Alongside the model weights, we’re releasing an example workflow that researchers can adapt to create reports from their own PDFs, providing a starting point for local report generation This post covers how we trained AstaBrief, what we learned about grounding it in scientific evidence, and which parts of our approach we think can carry forward to future models for science.

Most of the training and evaluation described was completed in 2025, so the proprietary models used to generate training data and as comparison points reflect the frontier at the time.

We haven’t rerun the full evaluation against today’s frontier models; the results below are best read as evidence about the particular training and system design choices we tested.

Training the model Our goal with AstaBrief was to build an open-weights model with all the qualities that matter most for long-form scientific synthesis: answer quality, relevance, structure, and citation grounding.

We started from Qwen3-8B and focused most of our effort on the post-training data, evaluation, and surrounding report-generation scaffolding.Adapting general-purpose models for scientific work – and training new scientific models from scratch – is something we're exploring broadly across Ai2.

Through NSF OMAI, a U.S.national initiative led by Ai2 to build fully open AI infrastructure and models for scientific discovery, our researchers are working directly with scientific communities to understand what they need from future open models and where today's general-purpose models fall short.

That includes studying how needs differ across scientific fields and workflows, with more findings from that research to share in the future.

Recent work, including our DR Tulu, has shown that reinforcement-learning-based (RL) methods can improve long-form report generation for open-weights models, especially when judge models are involved in the training loop.

We considered that path for AstaBrief, but ultimately focused on a simpler recipe built around supervised fine-tuning (SFT) and direct preference optimization (DPO).RL-based training can be unstable and expensive.

We wanted to see how far we could push report generation quality with a cheaper, more operationally manageable setup—one that's also easier to debug and iterate on.That made the quality of the training data especially important.

Rather than relying on a more complex optimization method to compensate for noisy examples, we spent much of the project figuring out how to generate, select, and filter examples that actually demonstrated the report-writing behavior we wanted.

We also wanted AstaBrief to be faster so that users could get preliminary reports quickly that they could then iterate over in subsequent turns.

For speed improvements, we decided to train AstaBrief to directly generate the final report in one pass given a user query and relevant retrieved snippets, bypassing the expensive snippet summarization and clustering stages our Claude-based Thinking mode uses and not writing out the answer section-by-section.

Interestingly, we found it was possible to do so without sacrificing performance.

Collecting SFT training data The training pipeline began with real user queries submitted through the system described in our paper “Synthesizing scientific literature with retrieval-augmented LMs” and ScholarQA, the framework that now underpins Asta’s Generate a report feature.

Rather than training only on synthetic prompts or benchmark-style tasks, we wanted AstaBrief to learn from real queries from real scientists.Our research suggests that scientists often ask different things of language models than users do of general-purpose chatbots or traditional search tools.

In our analysis of hundreds of thousands of Asta queries, expert researchers frequently supplied substantial context, multiple constraints, and relationships between concepts rather than relying on short, keyword-style prompts.

More recent Asta user studies have also surfaced differences in how researchers want AI involved in their work—some are comfortable using models for ideation or experimentation, while others prefer a narrower role in synthesis, literature surveillance, or pattern-finding.

Across those differences, participants want clearer source traceability, more visibility into what a model is doing, and greater control over the context it uses.

We filtered the user logs we collected for quality, relevance, and privacy, stripping out beta-tester and bot traffic, dropping queries that were too short to be meaningful, and using an LLM-based filtering pass to catch non-English queries, non-scientific requests, and prompts containing personal information.

That left a pool of 90K research-focused queries.For SFT, we generated full-report target outputs from the filtered queries using the multi-step ScholarQA pipeline behind Asta's report generation.

The pipeline retrieved relevant literature, organized the material into sections, and used a backing report-generating model to synthesize the evidence into a cited report.We drew on a mix of proprietary systems: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1.

After quality filtering, this yielded 47K usable training examples.Creating DPO pairs DPO required a different kind of training data.Instead of a single target report per query, we needed pairs of reports with one preferred over the other.

We built those pairs from a separate subset of queries not used during SFT data generation.One report per query came from the existing ScholarQA pipeline, typically backed by Claude 3.5 Sonnet or 3.7 Sonnet.

The competing report was generated by feeding ScholarQA's retrieved literature excerpts to a different model: o3, o4-mini, DeepSeek-V3, or DeepSeek-R1, depending on the example.Two judge models – GPT-4.1 and DeepSeek-R1 – compared each pair and picked a winner.

We ensured that LLM judges were aligned with human preferences (95% agreement) and only kept pairs where both judges agreed, which gave us a cleaner preference set and cut much of the noise that typically shows up in preference data generated at scale.

After quality filtering, the final DPO dataset came to about 6K examples.Using multiple generators and requiring agreement between two judges gave us a relatively simple way to construct preference data without treating any single model’s output or judgment as ground truth.

Filtering data for better attribution Our main evaluation target was SQABench-CS2, a set of 200 user-written computer science research questions.We tracked four metrics throughout the development of AstaBrief: Rubric score, which measures how much necessary content is covered by the report.

Answer precision, which measures whether each paragraph is relevant to the question.Citation precision, which measures whether each citation supports the claim it's attached to.Citation recall, which measures whether the report's claims are fully supported by the citations provided.

For our final model, we also ran secondary evaluations: DeepScholarBench, a 63-query benchmark for long-form research synthesis built from recent ArXiv papers, and two separate pairwise evaluations against reports generated by the Claude-powered pipeline—an LLM-judged comparison on SQABench-CS2 and a small human study.

A report can sound polished and complete while meandering from the question or attaching citations to claims from which the underlying evidence doesn't follow.For scientific synthesis, we needed to measure those behaviors separately.

But citation support is only part of scientific faithfulness—a model can cite the right study and still make a stronger claim than the study itself supports.

This can happen in subtle ways, for example, turning a finding about a particular sample into a generic claim about an entire population, shifting a result reported in the past tense into a present-tense statement that sounds more universally true, or turning a descriptive finding into a recommendation for what clinicians, policymakers, or researchers should do.

Those kinds of generalizations are especially important for scientific report generation because each step can broaden the apparent scope of the evidence without introducing an obviously false statement.

A cited sentence may therefore be technically related to its source while still overstating what researchers actually established.

Our development metrics focused primarily on relevance, coverage, and citation grounding; a richer evaluation of scientific report writers should also test whether they preserve the scope and strength of the claims in their sources.

Our first SFT runs improved overall content quality, but they still lagged behind our Claude-powered report generation pipeline on answer precision and citation quality.

In other words, the model got better at writing reports, but it still wasn’t grounded in evidence as consistently as we needed for scientific synthesis.That pushed us to spend more time on data quality.

We tested four statistics-based filters to identify weaker synthetic training examples: Output-to-input token ratio.Answers with very high ratios were often noisy because they were generating a lot of text from too little evidence.Citation relevance.

For each synthetic report in the training set, we averaged the retrieval relevance scores of its cited papers.Low averages suggested the report was relying

Related

相關文章

IT之家AI Agent

Meta 旗下 AI 智能體 Muse 將登陸智能眼鏡平臺,可代用戶完成各種任務

作者:漾仔 責編:漾仔 評論: 10 月 2 日消息,Meta 宣佈旗下 AI 智能體 Muse 將於近期登陸智能眼鏡,用戶無需拿出手機,只需通過語音下達指令,Muse 就能在後臺代用戶完成一系列任務。Meta 表示,Muse 基於 Muse Spark AI 模型運行,其核心運行環境是一臺部署在雲端的私有持久化 Linux 虛擬機,配備獨立的網頁瀏覽器、文件系統和終端。

剛剛
鈦媒體AI Agent

AI殺不死諮詢公司

防冷塗的蠟2026.10.02 16:50 · 來自浙江全文4693字00:00 / 14:23下一代諮詢公司真正值錢的能力是什麼?文 | 防冷塗的蠟10月1日,埃森哲公佈了最新財年報告。公司2026財年總收入約742億美元,按本幣計算增長5%;第四季度收入186.

7 分鐘前
IT之家AI Agent

OpenAI 通報逾百家第三方機構,自家智能體存在失控風險

作者:清源 責編:清源 評論: 10 月 2 日消息,當地時間 9 月 30 日晚,OpenAI 披露,旗下 AI 智能體或曾擅自嘗試繞過安全防護,或對百餘家機構的系統造成負面影響。OpenAI 此次披露的信息顯示,已知的 AI 智能體失控行為涉及範圍進一步擴大。

5 小時前
IT之家AI Agent

衝刺感恩節前掛牌,消息稱 Anthropic 尋求最早 11 月中旬上市

作者:清源 責編:清源 評論: 10 月 2 日消息,據彭博社今天(2 日)援引知情人士消息稱,在推遲原定計劃後,Anthropic 尋求最早於 11 月中旬上市。知情人士稱,Anthropic 最早可能在 11 月 9 日當週正式啟動 IPO 推介,並爭取在 11 月 26 日感恩節前掛牌交易。

6 小時前
IT之家AI Agent

加州檢察長向 OpenAI 發出傳票,調查 AI 網絡安全風險

作者:清源 責編:清源 評論: 10 月 2 日消息,據路透社今天(2 日)凌晨報道,加利福尼亞州總檢察長羅布 · 邦塔辦公室宣佈,邦塔本人已向 OpenAI 發出調查傳票,要求其就 AI 模型涉及的網絡安全事件和風險提供更多信息。邦塔上個月宣佈,加州司法部已就“Hugging Face 事件”正式展開調查。

8 小時前
Hugging Face BlogAI Agent

AutoSynthData: Generating Training Data for Enterprise Agents

Back to Articles AutoSynthData: Generating Training Data for Enterprise Agents Enterprise Article Published October 2, 2026 Upvote - Esakkivel Esakkiraja esakkivel Follow ServiceNow-AI Shruthan Radhakrishna shruthan-r Follow ServiceNow-AI Denis Akhiyarov dtanow Follow ServiceNow-AI Sagar Davasam davasam Follow ServiceNow-AI Enterprises need agents that work well in their own environments.

12 小時前