Article
Performance of reasoning large language models on nephrology multiple-choice questions
2025-12-08
Abstract excerpt
<h4>Aim</h4> Performance of large language models in medicine is improving, yet it remains unclear how the advantage of reasoning models depends on task characteristics in nephrology. <h4>Methods</h4> We evaluated four large language models in two families—OpenAI (GPT-5 reasoning, GPT-4o baseline) and Google (Gemini 2.5 Pro reasoning, Gemini 2.0 Flash baseline)—on 209 self-assessment questions for nephrology boa...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 7f210e7c-eea1-5542-b48f-2003b2c47302
- DOI
- 10.64898/2025.12.03.25341427
