Article
Context Matching is not Reasoning: Assessing Generalized Evaluation of Generative Language Models in Clinical Settings
2025-08-29
Abstract excerpt
<title>Abstract</title> <p>Current discussion surrounding the clinical capabilities of generative language models (GLMs) predominantly center around multiple-choice question-answer (MCQA) benchmarks derived from clinical licensing examinations. While accepted for human examinees, characteristics unique to GLMs bring into question the validity of such benchmarks. Here, we validate four benchmarks using eight GLMs,...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 9199db4c-9fa8-5bb4-ad55-7574c6465053
- DOI
- 10.21203/rs.3.rs-7325383/v1
