Article
CLEVER: Clinical Large Language Model Evaluationby Expert Review
2025-07-23
Abstract excerpt
<title>Abstract</title> <p>The proliferation of both general-purpose and healthcare-specific Large Language Models (LLMs) has intensified the challenge of rigorous benchmarking. Existing evaluation methods face critical limitations: data contamination undermines the validity of public benchmarks, self-preference biases LLM-as-a-judge approaches, and current tasks do not fully reflect real-world clinical applicati...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 8d8fe722-89e3-5af2-a4b0-ee653c160c74
- DOI
- 10.21203/rs.3.rs-6883531/v1
