Back to search

Article

Evaluating a Large Reasoning Model’s Performance on Open-Ended Medical Scenarios

2025-04-30

Abstract excerpt

Large language models (LLMs) have emerged as a dominant form of generative artificial intelligence (GenAI) in multiple domains. In early 2025, DeepSeek R1 was released, which is a new large reasoning model (LRM) that includes CoT (CoT) reasoning, Mixture of Experts (MoE), and reinforcement learning. As these technologies continue to improve, evaluating the accuracy and reliability of LLMs and LRMs in medicine rema...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
af6e78c5-0fc2-5ab7-a6cd-65900ac729ed
DOI
10.1101/2025.04.29.25326666
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Evaluating a Large Reasoning Model’s Performance on Open-Ended Medical ScenariosDOI 10.1101/2025.04.29.25326666
Select a neighboring publication to make it the new centre.