Article
Finding Rare Racial Stereotypes in LLM Responses via Adversarial Anomaly Detection
2026-05-18
Abstract excerpt
Large language models (LLMs) are increasingly aligned with antiracist norms through instruction tuning and reinforcement learning from human feedback (RLHF). However, overtly racist outputs have become rare, while more subtle, context-dependent racial stereotypes and assimilationist assumptions persist — appearing infrequently, often under specific linguistic or topical triggers. These rare stereotypical responses...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- e4a8565c-25eb-58c0-9b21-43ec145fd2c4
- DOI
- 10.14293/pr2199.003531.v2
