Back to search

Article

Comparison of Large Language Models’ Performance on Neurosurgical Board Examination Questions

2025-02-24

Abstract excerpt

<h4>Background</h4> Multiple-choice board examinations are a primary objective measure of competency in medicine. Large language models (LLMs) have demonstrated rapid improvements in performance on medical board examinations in the past two years. We evaluated five leading LLMs on neurosurgical board exam questions. <h4>Methods</h4> We evaluated five LLMs (OpenAI o1, OpenEvidence, Claude 3.5 Sonnet, Gemini 2.0, an...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
cb1f3341-1b4e-5246-818e-2956371c3c79
DOI
10.1101/2025.02.20.25322623
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Comparison of Large Language Models’ Performance on Neurosurgical Board Examination QuestionsDOI 10.1101/2025.02.20.25322623
Select a neighboring publication to make it the new centre.