Article
A single-patient task exposes a failure of safety alignment in clinical language models
2026-08-10
Abstract excerpt
Safety alignment should persist while a language model performs a task. We tested whether a single-patient triage task suppressed a warning about a second patient. Each case centered on Patient 1; Patient 2’s urgent problem appeared only in passing. Sixteen models saw each case twice: once as a general assistant and once while producing a triage record for Patient 1. As general assistants, models warned the caller...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- c9510f4b-4bbe-5dd9-bf91-4f05424dee35
- DOI
- 10.64898/2026.08.07.26359822
