Ai2 Releases AstaBrief 8B: Fast Open-Source Scientific Report…

At a Glance
| Item | Value |
|---|---|
| Publisher | Allen Institute for AI |
| Published | 2026-10-02 |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
Overview
The Allen Institute for AI (Ai2) has open-sourced “AstaBrief 8B", a high-speed report generation model for scientific research, along with its training data. This model is used as the “Fast mode" in the report generation feature of “Asta", an AI2 agent platform for scientific research.
AstaBrief is designed to generate academic reports with citations based on researcher questions and collected literature excerpt data. It is provided as an open-weight model that can be deployed and run locally on servers or infrastructure while maintaining the evidence-based answer quality and citation accuracy expected by researchers. By replacing the multi-step generation process incorporating conventional commercial models (Thinking mode) with a method that outputs the entire report in a single pass (one pass), it aims to reduce generation time and operational costs. Detailed information can be found on the official Ai2 blog (https://allenai.org/blog/astabrief).
Specifications
- Parameters: 8B
- Base model: Qwen3-8B
Performance
AstaBrief’s performance and evaluation results have been published by Ai2, the model’s release source. As the primary target for evaluation, “SQABench-CS2", a benchmark consisting of 200 research questions in computer science, was used. This evaluation tracked the following four metrics:
- “Rubric score", which measures how well the required content is covered
- “Answer precision", which measures whether each paragraph is relevant to the question
- “Citation precision", which measures whether each citation correctly supports the presented claim
- “Citation recall", which measures whether the claim is fully supported by the provided citation
In measurements conducted by Ai2 during development, AstaBrief showed competitive performance comparable in answer and citation quality to existing report generation pipelines based on Claude and “DR Tulu". During the SFT (Supervised Fine-Tuning) stage, it is reported that evidence-based generation accuracy was significantly improved by adopting a filtering method that removed samples with a low synthetic report “Citation density (the proportion of sentences with at least one citation)".
Additionally, as secondary evaluations, the “DeepScholarBench" long-form research synthesis benchmark comprising 63 questions created from recent ArXiv papers and a small-scale human study involving 14 questions evaluated by three scientific researchers were also conducted. While “DR Tulu" was superior in overall preference in this human study, two out of the three researchers rated AstaBrief higher than other systems on citation accuracy metrics.
In terms of report generation speed, in processing time across the entire Asta platform, the Claude-based “Thinking mode" takes an average of 178.5 seconds per item, whereas “Fast mode" using AstaBrief records an average of 51.1 seconds per item, achieving approximately a 3.5x speedup. Furthermore, regarding the feedback rate from 374 Asta users who actually tried “Fast mode", positive evaluations reached 84.2% (compared to 85.2% for Thinking mode), indicating satisfaction levels comparable to systems driven by commercial models. Note that these benchmark results and operational figures are all based on measurements conducted by the model release source itself rather than third-party verification.
Strengths and Use Cases
The primary strength of AstaBrief 8B lies in its ability to generate scientifically evidence-based reports at extremely high speeds. Unlike general chat models, it is designed to meet specific demands in scientific research. Specifically, it excels at creating long-form reports with appropriate citations for each claim from research questions and collected literature excerpts in a single pass.
Main intended use cases include:
– Scientific literature synthesis and summary: Synthesizing evidence from multiple literature sources for a research topic to create systematic reports.
– Literature surveys and pattern discovery: Conducting comparative analyses based on specific methods, target populations, and settings from large amounts of research data.
– Rapid creation of preliminary reports: Suitable for quickly drafting an overview before diving into “Thinking mode", which involves detailed thought processes.
– Use in privacy-focused research environments: Because it is released as open weights, it can be run on internal infrastructure or local environments without sending data to external APIs when handling confidential research or unpublished data.
This model focuses on maintaining scientific integrity. The development team prioritized “Citation density" during training data filtering to prevent the model from making unsupported claims. This measures the proportion of sentences in a report containing at least one citation, and eliminating synthetic data with low density suppresses the generation of weakly-supported text. It is also tuned to avoid scientific leaps, such as generalizing findings from specific samples to population-wide claims or describing past research results in the present tense as universal truths. This reduces the effort required for researchers to verify the final output and provides reports as reliable research artifacts.
The development process utilized tens of thousands of queries made by actual researchers on the Asta platform for training. This equips it with the ability to appropriately respond to researcher-specific inquiries involving complex contexts and constraints rather than short keyword searches. Moreover, adopting a pipeline that outputs the entire report at once rather than generating it section by section enables practical report creation speeds. According to Ai2’s user surveys, approximately 23% of users have completely migrated from traditional commercial model-based modes to this “Fast mode", demonstrating that the balance between speed and quality has reached a practical level.
How to Get It
AstaBrief 8B is available from the Hugging Face repository. It is distributed with open weights, and the training data is also released so that researchers and developers can reproduce and expand upon it in their own environments.
Downloads are performed using tools such as the Hugging Face CLI:
huggingface-cli download allenai/astabrief-8b
In addition to the model weights, Ai2 has also released a sample workflow for creating reports from the user’s own PDF files. This establishes an environment where local report generation can be started immediately. Because it uses Qwen3-8B as its base architecture, it is expected to work with many existing inference engines and frameworks.

