Article
Evaluating Large Language Model Diagnostic Performance on JAMA Clinical Challenges via a Multi-Agent Conversational Framework
2025-08-24
Abstract excerpt
<h4>Background & Objective</h4> Standard clinical LLM benchmarks use multiple-choice vignettes that present all information up front, unlike real encounters where clinicians iteratively elicit histories and objective data. We hypothesized that such formats inflate LLM performance and mask weaknesses in diagnostic reasoning. We developed and evaluated a multi-AI agent conversational framework that converts JAMA Cl...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- b59a3ee6-ffc7-512e-ae0a-4bde3c7b6691
- DOI
- 10.1101/2025.08.20.25334087
