Back to search

Article

Auditing frontier general-purpose large language models in biomedical tasks: reasoning gains, extraction limits, and benchmark reliability

2026-02-18

Abstract excerpt

<title>Abstract</title> <p>As large language models approach clinical deployment, their deployment-relevant reliability and the validity of the benchmarks used to assess it remain insufficiently examined. Here, we present a unified, reproducible, and human-centric audit of frontier general-purpose language models using representative biomedical text-mining tasks and nine biomedical question-answering benchmarks s...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
4d61cb3a-ece4-53c3-a6e8-d6aadd85a514
DOI
10.21203/rs.3.rs-8605899/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Auditing frontier general-purpose large language models in biomedical tasks: reasoning gains, extraction limits, and benchmark reliabilityDOI 10.21203/rs.3.rs-8605899/v1
Select a neighboring publication to make it the new centre.