Article
Normalization Is Necessary but Insufficient: A Cross-Class Audit of Glyph-Level Obfuscation and Attributive Reframing Against LLM Content-Safety Guards
2026-08-19
Abstract excerpt
Content-safety guard models are deployed around chat systems to catch harassment and abuse. We ask how robust they are when a message is rewritten so a human still reads it but the surface tokens change. Using a harm-free protocol with mild, generic, non-targeted insult-level items, we audit two real production guard models reached through a local gateway. We define GLYPH, ten glyph-level obfuscation classes in fo...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 759b267e-b293-5a7c-9dbf-22e957cad2a2
- DOI
- 10.20944/preprints202608.1281.v1
