Article
Aggregate benchmark scores obscure patient safety implications of errors across frontier language models
2026-03-20
Abstract excerpt
Frontier language models are widely used for health-related queries, yet aggregate benchmark scores do not capture safety implications of errors. We applied the recent Nature Medicine triage benchmark across nine frontier models, comparing directional error profiles, contextual bias, and crisis calibration. In-range accuracy ranged from 75.0% to 87.7%, obscuring clinically meaningful error differences. Looking at...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 45784e96-f1dd-5727-b68e-af461ad8a57b
- DOI
- 10.64898/2026.03.18.26348695
