Article
Citation Hallucination Determines Success: An Empirical Comparison of Six Medical AI Research Systems
2026-04-04
Abstract excerpt
Large language model (LLM) systems can now generate complete research manuscripts, yet their reliability in clinical medicine — where citation accuracy and reporting standards carry direct consequences — has not been systematically assessed. We introduce MedResearchBench, a benchmark of three clinical epidemiology tasks built on NHANES data, and use it to evaluate six AI research systems across six quality dimensi...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 2570ae72-0eef-5c38-9b98-83437de6730f
- DOI
- 10.64898/2026.04.02.26350091
