Back to search

Article

Across Generations, Sizes, and Types, Large Language Models Poorly Report Self-Confidence in Gastroenterology Clinical Reasoning Tasks

2025-06-04

Abstract excerpt

<title>Abstract</title> <p>This study evaluated confidence calibration across 48 large language models (LLM) using 300 gastroenterology board exam style questions. Regardless of response accuracy, all models demonstrated poor certainty estimation. Even the best-calibrated systems (o1 preview, GPT-4o, Claude-3.5-Sonnet) showed substantial overconfidence (Brier scores 0.15-0.2, AUROC ~0.6). Most concerning, models...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
16dffec3-5bf2-5b66-aba2-c1b0919d3c14
DOI
10.21203/rs.3.rs-6725427/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Across Generations, Sizes, and Types, Large Language Models Poorly Report Self-Confidence in Gastroenterology Clinical Reasoning TasksDOI 10.21203/rs.3.rs-6725427/v1
Select a neighboring publication to make it the new centre.