Back to search

Article

Textbook-Level Medical Knowledge in Large Language Models: A Comparative Evaluation Using the Japanese National Medical Examination

2025-09-12

Abstract excerpt

<h4>Study aims and objectives</h4> This study aimed to evaluate the performance of four reasoning-enhanced large language models (LLMs)—GPT-5, Grok-4, Claude Opus 4.1, and Gemini 2.5 Pro—on the Japanese National Medical Examination (JNME). <h4>Methods</h4> We evaluated LLM performance using the 2019 and 2025 JNME (n = 793). Questions were entered into each model with chain-of-thought prompting enabled. Accuracy...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
b43d3508-d8c7-5f78-920f-4790189f7640
DOI
10.1101/2025.09.10.25335398
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Textbook-Level Medical Knowledge in Large Language Models: A Comparative Evaluation Using the Japanese National Medical ExaminationDOI 10.1101/2025.09.10.25335398
Select a neighboring publication to make it the new centre.