Back to search

Article

Harm-Conditioned Computational Friction as a Diagnostic of Alignment Robustness: A Critical Review and Evaluation Framework

2025-12-11

Abstract excerpt

<title>Abstract</title> <p>Robust alignment requires that AI models maintain safety-relevant behavior under distribution shift, adversarial prompting, and optimization pressure. Current evaluation methods often rely on surface compliance metrics—such as refusal rates or policy-template adherence—that may fail to detect fragile safety generalization, reward hacking, or prompt-contingent refusal policies. This pape...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
0295915d-62b9-5d9a-9666-d33b83ec1f33
DOI
10.21203/rs.3.rs-8327468/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Harm-Conditioned Computational Friction as a Diagnostic of Alignment Robustness: A Critical Review and Evaluation FrameworkDOI 10.21203/rs.3.rs-8327468/v1
Select a neighboring publication to make it the new centre.