Article
ClaroAI-Bench: Evaluating Agentic Scientific Reproducibility on Real Biomedical Papers
2026-05-12
Abstract excerpt
We introduce ClaroAI-Bench , an evaluation suite for measuring AI agents’ ability to reproduce computational findings from published biomedical research. The benchmark comprises 35 real NIH-funded papers spanning five modalities (genomics, imaging, clinical/EHR, epidemiology, wet-lab) scored on a five-dimension rubric: data findability (D1), data accessibility (D2), code availability (D3), environment reconstruct...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 16b1a3c3-4c1a-552e-987b-a16879ec2847
- DOI
- 10.64898/2026.05.08.723611
