Back to search

Article

Benchmarking the Confidence of Large Language Models in Clinical Questions

2024-08-11

Abstract excerpt

<h4>Background and Aim</h4> The capabilities of large language models (LLMs) to self-assess their own confidence in answering questions in the biomedical realm remain underexplored. This study evaluates the confidence levels of 12 LLMs across five medical specialties to assess their ability to accurately judge their responses. <h4>Methods</h4> We used 1,965 multiple-choice questions assessing clinical knowledge...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
b2cb4e00-ebcf-5753-b82e-8090e7123233
DOI
10.1101/2024.08.11.24311810
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Benchmarking the Confidence of Large Language Models in Clinical QuestionsDOI 10.1101/2024.08.11.24311810
Select a neighboring publication to make it the new centre.