Back to search

Article

Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive Benchmark

2024-04-25

Abstract excerpt

The adoption of large language models (LLMs) to assist clinicians has attracted remarkable attention. Existing works mainly adopt the closeended question-answering (QA) task with answer options for evaluation. However, many clinical decisions involve answering open-ended questions without pre-set options. To better understand LLMs in the clinic, we construct a benchmark ClinicBench . We first collect eleven exist...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
24f394de-f4a5-5442-abc4-59672db9d94f
DOI
10.1101/2024.04.24.24306315
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive BenchmarkDOI 10.1101/2024.04.24.24306315
Select a neighboring publication to make it the new centre.