Back to search

Article

ClaroAI-Bench: Evaluating Agentic Scientific Reproducibility on Real Biomedical Papers

2026-05-12

Abstract excerpt

We introduce ClaroAI-Bench , an evaluation suite for measuring AI agents’ ability to reproduce computational findings from published biomedical research. The benchmark comprises 35 real NIH-funded papers spanning five modalities (genomics, imaging, clinical/EHR, epidemiology, wet-lab) scored on a five-dimension rubric: data findability (D1), data accessibility (D2), code availability (D3), environment reconstruct...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
16b1a3c3-4c1a-552e-987b-a16879ec2847
DOI
10.64898/2026.05.08.723611
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
ClaroAI-Bench: Evaluating Agentic Scientific Reproducibility on Real Biomedical PapersDOI 10.64898/2026.05.08.723611
Select a neighboring publication to make it the new centre.