Article
Evaluating Diagnostic Accuracy and Clinical Reasoning of Multiple Large Language Models in Psychiatry
2026-02-09
Abstract excerpt
<h4>Summary</h4> <h4>Background</h4> Existing large language model (LLM) evaluations rely on accuracy benchmarks that fail to capture whether models reason well while making diagnoses. Studies that do analyse reasoning focus on post hoc explanations accompanying model outputs rather than distinct, clinician-visible artifacts such as detailed reasoning traces. This creates a translational gap in domains such as p...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 2e2392cb-804b-55d8-a3ae-f00a8d691436
- DOI
- 10.64898/2026.02.03.26345402
