Article
BioML-bench: Evaluation of AI Agents for End-to-End Biomedical ML
2025-09-04
Abstract excerpt
Large language model (LLM) agents hold promise for accelerating biomedical research and development (R&D). Several biomedical agents have recently been proposed, but their evaluation has largely been restricted to question answering (e.g., LAB-Bench) or narrow bioinformatics tasks. Presently, there remains a lack of benchmarks evaluating agent capability in multi-step data analysis workflows or in solving the mach...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 8d150b02-2a53-5dad-8787-84efa307379d
- DOI
- 10.1101/2025.09.01.673319
