Article
Benchmarking Large Language Models for Replication of Guideline-Based PGx Recommendations
2025-05-15
Abstract excerpt
<title>Abstract</title> <p>We evaluated the ability of large language models (LLMs) to generate clinically accurate pharmacogenomic (PGx) recommendations aligned with CPIC guidelines. Using a benchmark of 599 curated gene–drug–phenotype scenarios, we compared five leading models, including GPT-4o and fine-tuned LLaMA variants, through both standard lexical metrics and a novel semantic evaluation framework (LLM Sc...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- e6fd0a80-a360-5033-9032-867862e5fb16
- DOI
- 10.21203/rs.3.rs-6630450/v1
