Back to search

Article

Evaluating Diagnostic Accuracy and Clinical Reasoning of Multiple Large Language Models in Psychiatry

2026-02-09

Abstract excerpt

<h4>Summary</h4> <h4>Background</h4> Existing large language model (LLM) evaluations rely on accuracy benchmarks that fail to capture whether models reason well while making diagnoses. Studies that do analyse reasoning focus on post hoc explanations accompanying model outputs rather than distinct, clinician-visible artifacts such as detailed reasoning traces. This creates a translational gap in domains such as p...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
2e2392cb-804b-55d8-a3ae-f00a8d691436
DOI
10.64898/2026.02.03.26345402
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Evaluating Diagnostic Accuracy and Clinical Reasoning of Multiple Large Language Models in PsychiatryDOI 10.64898/2026.02.03.26345402
Select a neighboring publication to make it the new centre.