Article
AI Safety Training Can be Clinically Harmful
2026-06-12
Abstract excerpt
<title>Abstract</title> <p>Large language models are being deployed as mental health support agents at scale, yet only 16\% of LLM-based chatbot interventions have undergone rigorous clinical efficacy testing \citep{hua2025charting}. We argue that the dominant alignment paradigm, RLHF-based safety training, can itself become a source of clinical harm when models are placed in protocol-driven psychotherapeutic set...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- ed499ba0-009d-529e-9c49-ed9c0b65fd8f
- DOI
- 10.21203/rs.3.rs-9568253/v1
