Back to search

Article

General purpose large language models match human performance on gastroenterology board exam self-assessments

2023-09-25

Abstract excerpt

<h4>Introduction</h4> While general-purpose large language models(LLMs) were able to pass USMLE-style examinations, their ability to perform in a specialized context, like gastroenterology, is unclear. In this study, we assessed the performance of three widely available LLMs: PaLM-2, GPT-3.5, and GPT-4 on the most recent ACG self-assessment(2022), utilizing both a basic and a prompt-engineered technique. <h4>Metho...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
f643a893-8ed4-520f-8c05-8e585de28b88
DOI
10.1101/2023.09.21.23295918
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
General purpose large language models match human performance on gastroenterology board exam self-assessmentsDOI 10.1101/2023.09.21.23295918
Select a neighboring publication to make it the new centre.