Back to search

Article

Normalization Is Necessary but Insufficient: A Cross-Class Audit of Glyph-Level Obfuscation and Attributive Reframing Against LLM Content-Safety Guards

2026-08-19

Abstract excerpt

Content-safety guard models are deployed around chat systems to catch harassment and abuse. We ask how robust they are when a message is rewritten so a human still reads it but the surface tokens change. Using a harm-free protocol with mild, generic, non-targeted insult-level items, we audit two real production guard models reached through a local gateway. We define GLYPH, ten glyph-level obfuscation classes in fo...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
759b267e-b293-5a7c-9dbf-22e957cad2a2
DOI
10.20944/preprints202608.1281.v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Normalization Is Necessary but Insufficient: A Cross-Class Audit of Glyph-Level Obfuscation and Attributive Reframing Against LLM Content-Safety GuardsDOI 10.20944/preprints202608.1281.v1
Select a neighboring publication to make it the new centre.