Article
BiomniBench: Process-level Evaluation of LLM Agents for Real-world Biomedical Research
2026-05-14
Abstract excerpt
LLM agents now perform real biomedical research, but evaluating them rigorously is hard. Outcome-only benchmarks fail in two ways. First, a correct final answer can come from memorization, reward hacking, or wrong reasoning that produces the right number by chance. Second, valid alternative analyses are marked wrong simply because they differ from the reference. We introduce BiomniBench, a process-level evaluation...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- c682b60e-292c-5c57-8f9a-ebc91296b43d
- DOI
- 10.64898/2026.05.12.724604
