Back to search

Article

Benchmarking the Performance of Large Language Models in Uveitis: A Comparative Analysis of ChatGPT-3.5, ChatGPT-4.0, Google Gemini, and Anthropic Claude3

2024-06-25

Abstract excerpt

<title>Abstract</title> <p>BACKGROUND/OBJECTIVE This study aimed to evaluate the accuracy, comprehensiveness, and readability of responses generated by various Large Language Models (LLMs) (ChatGPT-3.5, Gemini, Claude 3, and GPT-4.0) in the clinical context of uveitis, utilizing a meticulous grading methodology. METHODS Twenty-seven clinical uveitis questions were presented individually to four Large Language M...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
23578c49-5c96-553e-b066-25358fbbfedd
DOI
10.21203/rs.3.rs-4237467/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Benchmarking the Performance of Large Language Models in Uveitis: A Comparative Analysis of ChatGPT-3.5, ChatGPT-4.0, Google Gemini, and Anthropic Claude3DOI 10.21203/rs.3.rs-4237467/v1
Select a neighboring publication to make it the new centre.