Article
Comparison of Large Language Models’ Performance on Neurosurgical Board Examination Questions
2025-02-24
Abstract excerpt
<h4>Background</h4> Multiple-choice board examinations are a primary objective measure of competency in medicine. Large language models (LLMs) have demonstrated rapid improvements in performance on medical board examinations in the past two years. We evaluated five leading LLMs on neurosurgical board exam questions. <h4>Methods</h4> We evaluated five LLMs (OpenAI o1, OpenEvidence, Claude 3.5 Sonnet, Gemini 2.0, an...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- cb1f3341-1b4e-5246-818e-2956371c3c79
- DOI
- 10.1101/2025.02.20.25322623
