Back to search

Article

Citation Hallucination Determines Success: An Empirical Comparison of Six Medical AI Research Systems

2026-04-04

Abstract excerpt

Large language model (LLM) systems can now generate complete research manuscripts, yet their reliability in clinical medicine — where citation accuracy and reporting standards carry direct consequences — has not been systematically assessed. We introduce MedResearchBench, a benchmark of three clinical epidemiology tasks built on NHANES data, and use it to evaluate six AI research systems across six quality dimensi...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
2570ae72-0eef-5c38-9b98-83437de6730f
DOI
10.64898/2026.04.02.26350091
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Citation Hallucination Determines Success: An Empirical Comparison of Six Medical AI Research SystemsDOI 10.64898/2026.04.02.26350091
Select a neighboring publication to make it the new centre.