Article
Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models
2025-09-12
Abstract excerpt
<h4>Objectives</h4> To evaluate the performance of open and proprietary LLMs, with and without Retrieval-Augmented Generation (RAG), on cardiology board-style questions and benchmark them against the human average. <h4>Materials and Methods</h4> We tested 14 LLMs (6 open-weight, 8 proprietary) on 449 multiple-choice questions from the American College of Cardiology Self-Assessment Program (ACCSAP). Accuracy was...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 225d3f3b-16c7-5ee0-9eaa-07e44701c87d
- DOI
- 10.1101/2025.09.11.25335607
