Back to search

Article

Context Matching is not Reasoning: Assessing Generalized Evaluation of Generative Language Models in Clinical Settings

2025-08-29

Abstract excerpt

<title>Abstract</title> <p>Current discussion surrounding the clinical capabilities of generative language models (GLMs) predominantly center around multiple-choice question-answer (MCQA) benchmarks derived from clinical licensing examinations. While accepted for human examinees, characteristics unique to GLMs bring into question the validity of such benchmarks. Here, we validate four benchmarks using eight GLMs,...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
9199db4c-9fa8-5bb4-ad55-7574c6465053
DOI
10.21203/rs.3.rs-7325383/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Context Matching is not Reasoning: Assessing Generalized Evaluation of Generative Language Models in Clinical SettingsDOI 10.21203/rs.3.rs-7325383/v1
Select a neighboring publication to make it the new centre.