Article
Automating Evaluation of LLM-generated Responses to Patient Questions about Rare Diseases
2025-10-07
Abstract excerpt
<h4>Objectives</h4> Patients with rare diseases often struggle to find accurate medical information, and large language model (LLM)-based chatbots may help meet this need. However, evaluating LLM-generated free-text answers typically requires physician review, which is time-consuming and difficult to scale. This study compared traditional natural language processing (NLP) metrics to emerging LLM-based evaluation...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- e411adb2-7967-54e3-b1a7-8d52e061007d
- DOI
- 10.1101/2025.10.06.25337181
