Back to search

Article

Harm-Conditioned Computational Friction as a Contrastive Diagnostic of Safety Representations in Language Models: A Proof-of-Concept and Control Audit

2026-04-30

Abstract excerpt

Safety evaluation of large language models is commonly centered on behavioral outcomes, such as whether a model refuses or complies with harmful requests. Such evaluations are necessary but incomplete: two systems can produce similar refusal behavior while relying on different internal mechanisms, differing in brittleness, generalization, and susceptibility to adversarial prompting. This paper proposes _Harm-Condi...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
d2586caa-8c2f-5f4c-9a4b-5cf74c1ed565
DOI
10.32388/ee11wc.2
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Harm-Conditioned Computational Friction as a Contrastive Diagnostic of Safety Representations in Language Models: A Proof-of-Concept and Control AuditDOI 10.32388/ee11wc.2
Select a neighboring publication to make it the new centre.