Article
Harm-Conditioned Computational Friction as a Contrastive Diagnostic of Safety Representations in Language Models: A Proof-of-Concept and Control Audit
2026-04-30
Abstract excerpt
Safety evaluation of large language models is commonly centered on behavioral outcomes, such as whether a model refuses or complies with harmful requests. Such evaluations are necessary but incomplete: two systems can produce similar refusal behavior while relying on different internal mechanisms, differing in brittleness, generalization, and susceptibility to adversarial prompting. This paper proposes _Harm-Condi...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- d2586caa-8c2f-5f4c-9a4b-5cf74c1ed565
- DOI
- 10.32388/ee11wc.2
