Ai2 open-sources AstaBrief, an 8B model that writes cited research reports
The Allen Institute for AI (Ai2) has released AstaBrief, an 8-billion-parameter model that takes a research question plus retrieved passages from the literature and writes a cited report. It now powers the Fast mode of the "Generate a report" feature in Asta, Ai2's platform for scientific work, next to a Thinking mode that runs on Claude. Ai2 is publishing the weights, the training data and an example workflow for generating reports from a researcher's own PDFs.
The training recipe is deliberately simple. Ai2 started from Qwen3-8B and considered reinforcement learning, which its earlier DR Tulu work used for long-form reports, but chose supervised fine-tuning followed by direct preference optimisation because it is cheaper and easier to debug. For fine-tuning, it filtered real user queries down to about 90,000 research questions, generated target reports for them with its existing ScholarQA pipeline backed by several proprietary models, and kept 47,000 after quality filtering. Preference pairs set ScholarQA's report against one written by another model from the same retrieved excerpts, with two model judges picking the winner.
The most useful finding, by Ai2's account, was about filtering. Several signals were tried, including how relevant the cited papers were and how many of the retrieved papers a report used. The biggest gain came from one simple measure: dropping training reports in which too few statements carried a citation. More aggressive filters and their combinations added little.
Speed is the headline number. Across the full Asta pipeline, Fast mode averages 51.1 seconds per report against 178.5 seconds for Thinking mode, about 3.5 times faster; Ai2 says report generation is nearly ten times faster than with the proprietary models it had been tracking. On SQABench-CS2, 200 computer-science questions written by users, Ai2 describes AstaBrief as competitive with the Claude-backed pipeline and DR Tulu on answer quality and on citation precision and recall.
Ai2 adds an important caveat itself: most of the training and evaluation was done in 2025, so the proprietary models used for training data and comparison are those of that time, and it has not rerun the evaluation against current frontier models.

Why it matters
A small model that runs locally changes who can use this kind of tool. Ai2 points to research questions that reveal unpublished work, which institutions may not want to send to an external API. The open training data matters as much as the weights: anyone checking whether the model widens a study's conclusions — a failure Ai2 discusses at length — can now look at what it was trained to imitate.