Article
High Consistency, Limited Accuracy: Evaluating Large Language Models for Binary Medical Diagnosis
2025-12-09
Abstract excerpt
<h4>ABSTRACT</h4> <h4>Background</h4> Large Language Models (LLMs) have demonstrated impressive capabilities in medical knowledge tasks, achieving 60-80% accuracy on licensing examinations. However, their reliability and consistency in clinical diagnosis—critical for clinical trustworthiness–remain incompletely characterized. <h4>Objective</h4> To systematically evaluate the consistency and diagnostic accuracy...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- cd409f2b-6571-57ce-8f3e-ff599ce9bb33
- DOI
- 10.64898/2025.12.08.25341823
