Back to search

Article

Performance of reasoning large language models on nephrology multiple-choice questions

2025-12-08

Abstract excerpt

<h4>Aim</h4> Performance of large language models in medicine is improving, yet it remains unclear how the advantage of reasoning models depends on task characteristics in nephrology. <h4>Methods</h4> We evaluated four large language models in two families—OpenAI (GPT-5 reasoning, GPT-4o baseline) and Google (Gemini 2.5 Pro reasoning, Gemini 2.0 Flash baseline)—on 209 self-assessment questions for nephrology boa...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
7f210e7c-eea1-5542-b48f-2003b2c47302
DOI
10.64898/2025.12.03.25341427
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Performance of reasoning large language models on nephrology multiple-choice questionsDOI 10.64898/2025.12.03.25341427
Select a neighboring publication to make it the new centre.