Article
Auditing frontier general-purpose large language models in biomedical tasks: reasoning gains, extraction limits, and benchmark reliability
2026-02-18
Abstract excerpt
<title>Abstract</title> <p>As large language models approach clinical deployment, their deployment-relevant reliability and the validity of the benchmarks used to assess it remain insufficiently examined. Here, we present a unified, reproducible, and human-centric audit of frontier general-purpose language models using representative biomedical text-mining tasks and nine biomedical question-answering benchmarks s...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 4d61cb3a-ece4-53c3-a6e8-d6aadd85a514
- DOI
- 10.21203/rs.3.rs-8605899/v1
