Article
Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive Benchmark
2024-04-25
Abstract excerpt
The adoption of large language models (LLMs) to assist clinicians has attracted remarkable attention. Existing works mainly adopt the closeended question-answering (QA) task with answer options for evaluation. However, many clinical decisions involve answering open-ended questions without pre-set options. To better understand LLMs in the clinic, we construct a benchmark ClinicBench . We first collect eleven exist...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 24f394de-f4a5-5442-abc4-59672db9d94f
- DOI
- 10.1101/2024.04.24.24306315
