Back to search

Article

AI Safety Training Can be Clinically Harmful

2026-06-12

Abstract excerpt

<title>Abstract</title> <p>Large language models are being deployed as mental health support agents at scale, yet only 16\% of LLM-based chatbot interventions have undergone rigorous clinical efficacy testing \citep{hua2025charting}. We argue that the dominant alignment paradigm, RLHF-based safety training, can itself become a source of clinical harm when models are placed in protocol-driven psychotherapeutic set...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
ed499ba0-009d-529e-9c49-ed9c0b65fd8f
DOI
10.21203/rs.3.rs-9568253/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
AI Safety Training Can be Clinically HarmfulDOI 10.21203/rs.3.rs-9568253/v1
Select a neighboring publication to make it the new centre.