Back to search

Article

A single-patient task exposes a failure of safety alignment in clinical language models

2026-08-10

Abstract excerpt

Safety alignment should persist while a language model performs a task. We tested whether a single-patient triage task suppressed a warning about a second patient. Each case centered on Patient 1; Patient 2’s urgent problem appeared only in passing. Sixteen models saw each case twice: once as a general assistant and once while producing a triage record for Patient 1. As general assistants, models warned the caller...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
c9510f4b-4bbe-5dd9-bf91-4f05424dee35
DOI
10.64898/2026.08.07.26359822
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
A single-patient task exposes a failure of safety alignment in clinical language modelsDOI 10.64898/2026.08.07.26359822
Select a neighboring publication to make it the new centre.