Back to search

Article

Benchmarking Large Language Models for Replication of Guideline-Based PGx Recommendations

2025-05-15

Abstract excerpt

<title>Abstract</title> <p>We evaluated the ability of large language models (LLMs) to generate clinically accurate pharmacogenomic (PGx) recommendations aligned with CPIC guidelines. Using a benchmark of 599 curated gene–drug–phenotype scenarios, we compared five leading models, including GPT-4o and fine-tuned LLaMA variants, through both standard lexical metrics and a novel semantic evaluation framework (LLM Sc...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
e6fd0a80-a360-5033-9032-867862e5fb16
DOI
10.21203/rs.3.rs-6630450/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Benchmarking Large Language Models for Replication of Guideline-Based PGx RecommendationsDOI 10.21203/rs.3.rs-6630450/v1
Select a neighboring publication to make it the new centre.