Back to search

Article

Aggregate benchmark scores obscure patient safety implications of errors across frontier language models

2026-03-20

Abstract excerpt

Frontier language models are widely used for health-related queries, yet aggregate benchmark scores do not capture safety implications of errors. We applied the recent Nature Medicine triage benchmark across nine frontier models, comparing directional error profiles, contextual bias, and crisis calibration. In-range accuracy ranged from 75.0% to 87.7%, obscuring clinically meaningful error differences. Looking at...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
45784e96-f1dd-5727-b68e-af461ad8a57b
DOI
10.64898/2026.03.18.26348695
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Aggregate benchmark scores obscure patient safety implications of errors across frontier language modelsDOI 10.64898/2026.03.18.26348695
Select a neighboring publication to make it the new centre.