Papers


AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research Bernie Boscoe, Tuan Do, Jack Stark, Srinath Saikrishnan, Vikram Seenivasan, PJ Allen, Morgan Himes, Jonathan Soriano, Andrew Lizarraga

Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrievalaugmented generation (RAG) systems that provide natural language access to scientific knowledge and research workflows. Researchers are exploring the viability of these systems as natural language interfaces for document search and for generating analysis code and pipeline components. At the same time, concerns about data privacy and control over research infrastructure have motivated interest in open-weight models and open-source deployments hosted within research institutions.

In astronomy, this development follows a long history of computational infrastructure development, from archival databases and SQL-based systems to LLM-assisted research tools. This paper presents a domain-expert evaluation of faithfulness for AquiLLM, an open-weight, offline RAG-LLM platform designed to support scientific research groups in the use and preservation of tacit and formal knowledge.

We define faithfulness as the extent to which generated responses remain grounded in retrieved scientific context without unsupported claims or omissions. We report results from an astronomy case study evaluating AquiLLM across retrieval and scientific analysis tasks. AquiLLM performs most reliably on explicit retrieval-oriented questions grounded in the RAG collection, while faithfulness degrades for queries requiring synthesis or ambiguity resolution. These results highlight both the promise and limitations of open-weight RAG-LLM systems for scientific research and demonstrate the importance of domain-expert evaluation beyond standard benchmark leaderboards.

Measuring What Matters in the Age of AI: The Research Software Metrics Landscape Addi Malviya Thakur, Reed Milewicz, Gregory Watson and Audris Mockus

DART: Distributed Assignment of Research Tasks for Heterogeneous Compute Environments Abdur Rouf, Fahad Ahmad Khan, Benjamin Keene, Shafaq Chaudhry and Murat Yuksel

Ten Practices for Refactoring Bioinformatics Pipelines to nf-core DSL2 on Shared Slurm HPC Nil Tianchen Mu, William Dizon, Glen Otero and Torey Battelle

Tapis AI Assistant: Increasing Research Impact through Traceable and Evaluated Agentic Tools Smruti Padhy, Anagha Jamthe and Joe Stubbs

Structuring agentic AI for HPC code modernization Anthony Marinov and Igor Sfiligoi

Josh: Efficient, Portable, and Scalable Cross-Disciplinary Vegetation Modeling via Domain-Specific Language A Pottinger, Nick Gondek, Lucia Layritz, Maya Zomer, Nicolas Graver, Amanda Anderson-You and Maya Weltman-Fahs

Take it on the road: How RCD Helps Field Researchers Report Culvert Condition even when Offline Jing Qi, William Cowen and Christian Darabos

Escaping the Enclave: Format-Faithful Synthetic Data for Open, Reproducible Medicare Claims Pipelines Pavel Belakurski, Dmitry Etin, Mark Chumack and Michael Bouzinier

Fabla: An Open-Source Voice-First EMA Platform for Clinical Research Santiago Arconada Alvarez, Tulika Banerjee, Wiza Munthali, Kennedy Linzie, Mabuchi Nyrienda, Hope Madziakapita, Morgan Greenleaf, Wilbur Lam and Deanna Kaplan

We Create Quality: Towards a Human-Centric Theory of Research Software Quality in the Age of AI Reed Milewicz, Connor Brynteson, Ella Luedeke and Italo Santos

AI-Augmented Research Software Engineering: Risks, Challenges, and Practices I Luk Kim, Xiao Liu, Jungha Woo, Elham J Barezi, Jorge Ivan Fuentes Rosado, Jaewoo Shin, Lan Zhao and Carol X. Song

A Generalized Methodology for Evaluating AI-Integrated Research Software: Lessons from a Smart Search Implementation for Cyber-Training Xiao Liu, Jungha Woo, Lan Zhao, Jaewoo Shin, Chimdia Kabuo and Carol Song