AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research
Published 18 Sept 2026arXiv:2609.16519
Updated 4 h ago · first seen 16 Sept 2026
paper_01M2PPQEKGMRY7FVGQXG4A2J6G
Abstract
Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that provide natural language access to scientific knowledge and research workflows. Researchers are exploring the viability of these systems as natural language interfaces for document search and for generating analysis code and pipeline components. At the same time, concerns about data privacy and control over research infrastructure have motivated interest in open-weight models and open-source deployments hosted within research institutions. In astronomy, this development follows a long history of computational infrastructure development, from archival databases and Structured Query Language (SQL)-based systems to large language model (LLM)-assisted research tools. This paper presents a domain-expert evaluation of faithfulness for AquiLLM, an open-weight, offline RAG-LLM platform designed to support scientific research groups in the use and preservation of tacit and formal knowledge. We define faithfulness as the extent to which generated responses remain grounded in retrieved scientific context without unsupported claims or omissions. We report results from an astronomy case study evaluating AquiLLM across retrieval and scientific analysis tasks. AquiLLM performs most reliably on explicit retrieval-oriented questions grounded in the RAG collection, while faithfulness degrades for queries requiring synthesis or ambiguity resolution. These results highlight both the promise and limitations of open-weight RAG-LLM systems for scientific research and demonstrate the importance of domain-expert evaluation beyond standard benchmark leaderboards.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 3
- Property changedPaperAquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research
AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research: arxiv announce type changed from new to replace
Arxiv announce typenew→replacearxiv - Property changedPaperAquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research
AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research: published at changed from 2026-09-16T04:00:00+00:00 to 2026-09-18T04:00:00+00:00
Published16 Sept 2026→18 Sept 2026arxiv - New paperPaperAquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research
New paper: AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research
arxiv
Sources
Sources 1
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.