Back to search

Article

Multidimensional Evaluation of Large Language Models on the AAP In-Service Examination: Assessing Accuracy, Calibration, and Citation Reliability

2025-10-17

Abstract excerpt

<h4>Background</h4> Large language models (LLMs) have demonstrated rapid advancements in natural language understanding and generation, prompting their integration into biomedical research, clinical practice, and professional education. However, systematic evaluation of LLMs in specialty-specific domains such as dentistry and periodontology remain limited, particularly regarding multidimensional performance metri...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
5c8ae831-764a-5013-9a0e-378730109190
DOI
10.1101/2025.10.14.25338040
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Multidimensional Evaluation of Large Language Models on the AAP In-Service Examination: Assessing Accuracy, Calibration, and Citation ReliabilityDOI 10.1101/2025.10.14.25338040
Select a neighboring publication to make it the new centre.