Article
Medical errors in large language models revealed using 1,000 synthetic clinical transcripts
2026-03-25
Abstract excerpt
<h4>ABSTRACT</h4> Current clinical evaluations of large language models (LLMs) rely on datasets which fail to reflect real-world medical complexity. We developed a high-throughput simulation of patients presenting with headache and generated 1,000 doctor-patient transcripts, enabling an unprecedented mapping of clinical reasoning failures across a vast spectrum of demographic and clinical phenotypes. While GPT-5....
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- f33d50c5-67fe-500d-96b7-2cc4c099104f
- DOI
- 10.64898/2026.03.23.26349082
