AstaBrief 8B is an open-weights model released to speed up generation of long-form, cited scientific reports within Asta. Built from Qwen3-8B and optimized with supervised fine-tuning (SFT) and direct preference optimization (DPO) rather than reinforcement learning, it generates full reports in one pass from a research question plus retrieved snippets, cutting end-to-end report time from 178.5s to 51.1s on average (about 3.5× faster) while remaining usable for sensitive on-premises workflows. Weights, training data, and an example local PDF-to-report workflow are published so institutions can run and adapt the system themselves.
Training prioritized real researcher behavior and citation grounding: 90K filtered user queries produced 47K SFT examples via the ScholarQA retrieval-and-synthesis pipeline (backed by a mix of proprietary models), and a cleaned DPO set of ~6K pairwise comparisons where GPT-4.1 and DeepSeek-R1 judges agreed. Development focused on relevance, coverage, citation precision/recall, and preserving the scope and strength of source claims, recognizing that correct citations can still enable subtle overgeneralization. Evaluations used SQABench-CS2, DeepScholarBench, LLM-judged comparisons, and a small human study; results show that careful data selection and system design can yield faster, citation-grounded report generation without complex RL training, though comparisons reflect the frontier as of 2025.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.