Article
Benchmarking and behavioral characterization of LLM agents for protein design
2026-05-08
Abstract excerpt
Large language models (LLMs) are increasingly deployed as agents for scientific discovery, but standardized frameworks for evaluating their performance and behavior in scientific workflows are lacking. Protein design provides a demanding test case because modern workflows combine stochastic generative models, structure prediction systems, and physics-based evaluation tools that require extensive candidate explorat...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 1840f4ca-f7e4-5816-8221-e9a0e6363b21
- DOI
- 10.64898/2026.05.06.723381
